William J.B. Mattingly
banner
wjbmattingly.bsky.social
William J.B. Mattingly
@wjbmattingly.bsky.social
Digital Nomad · Historian · Data Scientist · NLP · Machine Learning

Cultural Heritage Data Scientist at Yale
Former Postdoc in the Smithsonian
Maintainer of Python Tutorials for Digital Humanities

https://linktr.ee/wjbmattingly
Thanks so much!
October 6, 2026 at 6:11 PM
Thanks for the question! I responded in the first post. I'm terrible at responding to the correct threads lol
October 6, 2026 at 6:10 PM
would absolutely have substitutions for cum and eum because it isn't a language model, rather a traditional deterministic neural net. So in a sense, this small language model behaves /almost/ like a page-level Kraken model, while being an LLM.
October 6, 2026 at 6:10 PM
This makes the model's output mimic Comma, which is a dataset generated from a deterministic Kraken model. In other words, the LLM's previous training that may or may not have including Matthew would be more representative of the manuscript culture of Comma and the behavior of Kraken which (cont.)
October 6, 2026 at 6:10 PM
Great point! This is a full-finetune, not just a LoRA adapter. I don't have solid evidence to support this, though it wouldn't be hard to do, but it is pretty clear that a lot of the model's previous training was pushed out of it. We used ~40,000 image/text pairs from the Comma dataset. (cont.)
October 6, 2026 at 6:10 PM
So for this example, it was the same model. I just changed the temperature. For the other examples, I would train the model with different seeds and/or train the model with different subsets of the data. Each method works, but just slightly differently.
October 6, 2026 at 2:45 PM
Thanks for sharing!! =) One of the things that you can do to improve this further is train the same model on different subsets of the data or train with different seeds. We're doing that right now and it's showing an improvement of this earlier approach.
October 6, 2026 at 1:50 PM
Absolutely. It will let you essentially connect named entities (when possible) to linked open data (if available). This means that an individual who appears across a newspaper would have the same identifier.
October 5, 2026 at 12:56 PM
Thanks! And thanks for the great suggestion! That's something we are testing. The tricky bit is aligning the 3 models. We have 4 models trained on the same data, that may be a good way to do it.
September 25, 2026 at 9:49 AM
Very cool!! thanks for sharing!
September 16, 2026 at 5:34 PM
=) me too!
September 16, 2026 at 2:15 PM
Thanks for your interest! I will very soon!! I'm in the middle of a conference right now, but will try and share a blog post in a week or two.
September 9, 2026 at 8:34 AM
These models follow the CATMuS guidelines
September 8, 2026 at 10:18 AM
I should emphasize that these models were taught to not transcribe drop-capitals and marginalia, but they ignore watermarks. Image was from a post by @aaronm.bsky.social
September 8, 2026 at 10:17 AM
Since each primary source material is extracted from these articles, you can also track how a specific source is used across secondary lit.
July 21, 2026 at 7:08 PM