William J.B. Mattingly
banner
wjbmattingly.bsky.social
William J.B. Mattingly
@wjbmattingly.bsky.social
Digital Nomad · Historian · Data Scientist · NLP · Machine Learning

Cultural Heritage Data Scientist at Yale
Former Postdoc in the Smithsonian
Maintainer of Python Tutorials for Digital Humanities

https://linktr.ee/wjbmattingly
Reposted by William J.B. Mattingly
We still have some files to QC, but overall, we can announce it: we have finalized a first version of DTS-ready TLG CD, based on Diogenes earlier work, called the CLLG Freed Corpus.

Git: gitlab.inria.fr/almanach/cll...
Static reading environment: freed-corpus-bae861.gitlabpages.inria.fr/index.html
CLLG Freed corpus
freed-corpus-bae861.gitlabpages.inria.fr
October 7, 2026 at 1:20 PM
Reposted by William J.B. Mattingly
✨Version bump✨

latincy-vocab v0.4.0
Latin vocabulary list builder

Improved handling of missing glosses, better POS-based lemma resolution, and more

github.com/latincy/lati...
GitHub - latincy/latincy-vocab: Latin vocabulary list builder powered by the LatinCy spaCy ecosystem
Latin vocabulary list builder powered by the LatinCy spaCy ecosystem - latincy/latincy-vocab
github.com
October 6, 2026 at 3:02 PM
Finetuned GliNER Decision is now beating Jev in entity linking. In our tests, we start to see it enter parity with around 600 examples, but we pushed it a bit further with around 3,000 examples. We'll be sharing the model and the data soon!
October 1, 2026 at 7:43 PM
woot woot! First finetune of GliNER-Decide!
September 25, 2026 at 1:16 PM
Here's my blog on how to leverage the stochastic nature of VLMs to identify areas of potential divergence (most likely areas for errors).

Read more here! wjbmattingly.com/blog/where-t...
September 24, 2026 at 3:58 PM
One of the fun things about LLMs is we can use their greatest weakness in OCR/HTR (their non-determinism) to be their greatest strength in identifying potential errors. The same model running over the same image multiple times shows where divergence (errors) appear. Blog and app coming soon!
September 16, 2026 at 10:55 AM
Qwen 3.5 0.8B finetune comma base model now has LoRA adapters that can do bbox alignment in the same pass as HTR.
September 9, 2026 at 7:34 AM
Still finetuning these models. Page-level VLM Qwen 3.5 0.8B finetunes for medieval Latin manuscripts. All models were trained on 10k images. I am now training further on an additional 20k images from the comma dataset. You can test the models on @hf.co : huggingface.co/spaces/wjbma...
September 8, 2026 at 10:13 AM
I've added 210 articles to this local database and the knowledge graph is starting to get extensive. KGs get more useful as they get larger. Of these 210 articles, there are 30 references to Alcuin corresponding with (only type of many relationships in the KG) someone with citations and excerpts.
July 22, 2026 at 3:42 PM
Prototype in place for going from a collection of scholarly articles to knowledge graph. Also, there's a component in place that automatically links to linked open data, like Wikidata. All of this knowledge is from only 5 articles.
July 21, 2026 at 7:08 PM
Daniel does a lot of cool stuff, but this is awesome!
Open ASR models with speaker diarization are now fast and cheap: I diarized 174 hours of Apollo 11 mission audio (the real July 1969 NASA tapes) for $9.46 with a 0.9B open model. Search it, hear any moment on the original tape:
huggingface.co/spaces/davan...
July 10, 2026 at 3:15 PM
Sneak peak at what I've been working on =). Zero-shot with Gemini 3.5 Flash. Object detection for lines and images and transcription at the line level all in a single prompt. I'll be talking about this in Vienna in September. Fable 5 made the viewer.
July 10, 2026 at 2:49 PM
We found a 398-node geographic cycle in Yale's LUX database that poisoned 458,214 downstream records. By using Tarjan's SCC algorithm and cleaning the data with Gemini Flash Lite, we fixed the issue for about $6.50.
June 15, 2026 at 6:44 PM
Reposted by William J.B. Mattingly
Recently published: article on the research that went into the 2.0 version of the Shakespeare and Company Project datasets, and the potential of the updates and additional data on member addresses and authors.
doi.org/10.22148/jca...

One of the first few in the newly launched JCA!
<em>Shakespeare and Company Project</em> Data Sets, Version 2.0
The Shakespeare and Company Project data sets provide a detailed portrait of Shakespeare and Company, Sylvia Beach’s bookshop and lending library in interwar Paris. This article outlines the research,...
doi.org
May 18, 2026 at 4:50 PM
Reposted by William J.B. Mattingly
New #DigitalHumanities journal
Awesome! The first issue of "Digital Humanities Intersections" has been published!

https://dhi.iiti.ac.in/index.php/dhjournal/index

DHI is a new "open-access, peer-reviewed journal committed to advancing multilingual and interdisciplinary research in Digital Humanities."

Check it out!
May 18, 2026 at 12:50 PM
We tested OpenAI’s Whisper on 1,847 Holocaust testimonies from the Fortunoff at Yale. It achieved 85% accuracy, but the errors reveal how AI struggles in this domain. Full findings: https://wjbmattingly.com/blog/transcribing-holocaust-testimonies-in-yale-s-fortunoff-archive
May 12, 2026 at 5:08 PM
Mallorca is so photogenic. It’s hard to take a bad photo.
May 10, 2026 at 1:25 PM
Reposted by William J.B. Mattingly
‘Exploring Digital Cultural Heritage’ is now out in the wild! You can download an open-access copy (PDF) or read the free digital edition on Manifold uolpress.co.uk/book/explori.... Loved working on this with Eirini Goudarouli, @amsichani.bsky.social & the @uolpress.bsky.social team.
May 7, 2026 at 11:32 AM
Off to Mallorca to present on our work at Fortunoff! More coming soon!
May 6, 2026 at 11:39 PM
We parsed 3.6 million historical names with 96% accuracy by moving away from expensive frontier models to fine-tuned Qwen 3.5 models. We also found that switching our output format from JSON to YAML dramatically increased stability.
May 4, 2026 at 1:44 PM
Haven't gotten to use graph theory in a long time. The last few days have been a lot of fun! I just finished writing a blog about how we are using graph theory to improve how places are represented in LUX at Yale. More coming soon!
April 30, 2026 at 5:32 PM
I’ve replaced Conda with uv by @crmarsh.com for my Python projects. The slow dependency resolution and cross-platform friction between Windows and Linux was the catalyst. uv is faster, leaner, and more reliable for team workflows. Full breakdown here: https://wjbmattingly.com/blog/bye-conda-hello-uv
April 27, 2026 at 3:48 PM
I spent the weekend migrating my old PythonHumanities website from WordPress to a custom design (thanks, Claude) that allows for more flexibility, including live coding exercises in each page. I also updated things to make getting started with Python easier by using uv. More soon!
April 27, 2026 at 11:08 AM
Reposted by William J.B. Mattingly
✨ LatinCy v3.9 sm/md/lg/trf pipelines for SpaCy available ✨

- Improved tokenization and u/v norm
- New custom Latin-specific XPOS tags
- Better, more consistent lemma/morph coverage

huggingface.co/latincy/la_c...

#digiclass #nlproc
March 26, 2026 at 2:43 PM