Sebastian Majstorovic
banner
storytracer.com
Sebastian Majstorovic
@storytracer.com
Managing Director of datovis.com. Open Data Consultant for @eleutherai.bsky.social. Technical Director of @datarescueproject.org. Personal website: storytracer.com.
Reposted by Sebastian Majstorovic
Made some improvements to the OCR scripts onboarding in my uv-scripts collection.

First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models.
huggingface.co/datasets/uv-...
September 16, 2026 at 3:09 PM
Reposted by Sebastian Majstorovic
Need a historical illustration?

Search 1.49 million images from British Library books and Britannica (1500s–1920s).

Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.

huggingface.co/spaces/davan...
September 11, 2026 at 1:42 PM
Reposted by Sebastian Majstorovic
OCR for Japanese manga, Swedish handwriting or Arabic print?

There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.

I’ve gathered 41 models into four collections, with short notes to help you choose:

huggingface.co/collections/...
OCR on the Hub - a davanstrien Collection
Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.
huggingface.co
September 7, 2026 at 3:43 PM
Reposted by Sebastian Majstorovic
A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger.

On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.

huggingface.co/spaces/fineb...
September 3, 2026 at 2:53 PM
Reposted by Sebastian Majstorovic
Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved!

One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
August 20, 2026 at 10:42 AM
Reposted by Sebastian Majstorovic
Interested in participating in the 2027 Central European History Convention July 15-17, 2027? Submit a paper proposal! Deadline September 13, 2026. Check here for all the different options of how to apply to be part of the program! ceh-c.univie.ac.at/news/call-fo... Please share far and wide!
Call for 2027
Call for 2027
ceh-c.univie.ac.at
August 12, 2026 at 2:18 PM
Reposted by Sebastian Majstorovic
FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment.

Step one: work out which modern OCR models are actually good enough.
August 10, 2026 at 12:18 PM
Are open OCR models good enough to unlock historical knowledge? @danielvanstrien.bsky.social and I built a leaderboard using six expert-transcribed volumes from the @biodivlibrary.bsky.social sky.social. Meet FineBooks, from @hf.co and @eleutherai.bsky.social! huggingface.co/blog/fineboo...
FineBooks: are open OCR models good enough to unlock historical knowledge?
A Blog post by FineBooks on Hugging Face
huggingface.co
August 10, 2026 at 12:03 PM
Reposted by Sebastian Majstorovic
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big...

You can also do semantic search against the images here: huggingface.co/spaces/davan...
August 7, 2026 at 5:18 PM
Heon Cha Haus was a great tip for Seoul @tedunderwood.com. I had their signature iced tea and an extremely tasty Red Bean & Cream Rice Cake. Not to mention the wonderful service! #DH2026
July 26, 2026 at 8:32 AM
Enjoying the weekend in Seoul before #DH2026 kicks off next week in Daejeon @dh2026daejeon.bsky.social.
July 25, 2026 at 2:11 PM
Reposted by Sebastian Majstorovic
I saw this and had to share.
July 10, 2026 at 11:01 PM
Reposted by Sebastian Majstorovic
Reading the Archive by Machine:
An OCR Benchmark for Historians, 1612–1921

Here is version 1 of a working paper on the new OCR tools that are transforming digital history.

working-papers-in-critical-search.github.io/paper-004-oc...
Reading the Archive by Machine – Working Papers in Critical Search
A benchmark of six OCR systems (Tesseract, olmOCR 2, Chandra 2, Infinity Parser 2, GLM-OCR, and Gemini 3.5 Flash) on human-transcribed archival documents spanning 1612–1921: early-modern print, ninete...
working-papers-in-critical-search.github.io
July 10, 2026 at 4:21 PM
Reposted by Sebastian Majstorovic
We are so very grateful for this honor! Thank you to our supporters and everyone who voted. Congratulations to everyone who was nominated -- we feel very privileged to be counted among your incredible work.
Congratulations to the 2026 Data Integrity Award winners
- @datarescueproject.org
- Gina Plata-Nino of @fracposts.bsky.social
- @mapresearch.bsky.social & Williams Institute
- @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal
Details in 🧵
Congratulations Fireworks Display
ALT: Congratulations Fireworks Display
static.klipy.com
July 8, 2026 at 4:06 PM
Reposted by Sebastian Majstorovic
Congratulations to the 2026 Data Integrity Award winners
- @datarescueproject.org
- Gina Plata-Nino of @fracposts.bsky.social
- @mapresearch.bsky.social & Williams Institute
- @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal
Details in 🧵
Congratulations Fireworks Display
ALT: Congratulations Fireworks Display
static.klipy.com
July 8, 2026 at 3:08 AM
Reposted by Sebastian Majstorovic
Read the feasibility study for the European Books Data Commons and discover the next steps!

Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tY
Feasibility study paves the way for a European Books Data Commons
The National Library of the Netherlands (KB) and the Europeana Foundation have just published a feasibility study for the European Books Data Commons (EBDC).
www.dataspace-culturalheritage.eu
July 8, 2026 at 8:23 AM
Reposted by Sebastian Majstorovic
@storytracer.com shows how volunteers or #GLAMS workers can be responsible actors in the current political climate by protecting datasets from war, political censorship and defunding! #glamlabsfutures
June 25, 2026 at 3:39 PM
Reposted by Sebastian Majstorovic
I think VLM-based OCR might finally be close to working on historic newspapers!

Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow.

Surya OCR 2 (a 650M model!) does a very good job!
June 22, 2026 at 3:16 PM
Reposted by Sebastian Majstorovic
Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via
@europeana.bsky.social newspapers.

Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
June 23, 2026 at 8:36 AM
Reposted by Sebastian Majstorovic
Read our newly published paper on AI 'The case for Public AI: making it happen with cultural heritage.'

400+ professionals, one shared position - cultural heritage should shape AI, not just feed it.

👉 bit.ly/4xLdyU3

#PublicAI #AI #ArtificialIntelligence #CulturalHeritage #DataSpace
Making Public AI reality: How cultural heritage can lead the way
After months of collaborative development and iteration with the data space community and the Europeana Initiative, we are excited to share the newly published paper, ‘The case for Public AI: making i...
www.dataspace-culturalheritage.eu
June 23, 2026 at 6:13 AM
Reposted by Sebastian Majstorovic
Unfortunate to see this outcome but the fight for the panels at Washington’s house will continue in the community!

This is our history and we want to #SaveOurSigns
The three-judge panel unanimously agreed to toss out an injunction that ordered the National Park Service to restore interpretive panels telling the story of nine people enslaved at the site.
Slavery exhibits at President’s House can be replaced by Trump administration, Third Circuit rules
www.inquirer.com
June 19, 2026 at 10:27 AM
Reposted by Sebastian Majstorovic
In the 90 minutes after @theguardian.com article went live, BHL received more than US$2,500 in donations.

Thank you to everyone helping keep biodiversity knowledge free and open. Please read, share, and help us build momentum.
A bonanza for fans of the natural world: the digital library sharing 64m pages of scientific knowledge with everyone
The Biodiversity Heritage Library is an invaluable online archive of historic texts on species living and lost supplied by the world’s leading museums and universities. Now its future is in doubt
www.theguardian.com
June 18, 2026 at 1:05 PM
Reposted by Sebastian Majstorovic
A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controversial directive last year.
Judge orders Trump administration to restore signs changed at national parks | CNN Politics
A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controv...
cnn.it
June 13, 2026 at 7:00 PM
Reposted by Sebastian Majstorovic
OCR for Ancient Greek is constrained by a lack of open training data. But the deeper challenge tackled by 2 papers from Inria is doing OCR and document structure recovery together: section hierarchies, milestone numbering, marginal references.

Paper 1: arxiv.org/html/2603.02...
May 29, 2026 at 9:37 AM