First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models.
huggingface.co/datasets/uv-...
First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models.
huggingface.co/datasets/uv-...
Search 1.49 million images from British Library books and Britannica (1500s–1920s).
Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.
huggingface.co/spaces/davan...
Search 1.49 million images from British Library books and Britannica (1500s–1920s).
Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.
huggingface.co/spaces/davan...
There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.
I’ve gathered 41 models into four collections, with short notes to help you choose:
huggingface.co/collections/...
There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.
I’ve gathered 41 models into four collections, with short notes to help you choose:
huggingface.co/collections/...
On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.
huggingface.co/spaces/fineb...
On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.
huggingface.co/spaces/fineb...
One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
Step one: work out which modern OCR models are actually good enough.
Step one: work out which modern OCR models are actually good enough.
You can also do semantic search against the images here: huggingface.co/spaces/davan...
You can also do semantic search against the images here: huggingface.co/spaces/davan...
An OCR Benchmark for Historians, 1612–1921
Here is version 1 of a working paper on the new OCR tools that are transforming digital history.
working-papers-in-critical-search.github.io/paper-004-oc...
An OCR Benchmark for Historians, 1612–1921
Here is version 1 of a working paper on the new OCR tools that are transforming digital history.
working-papers-in-critical-search.github.io/paper-004-oc...
- @datarescueproject.org
- Gina Plata-Nino of @fracposts.bsky.social
- @mapresearch.bsky.social & Williams Institute
- @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal
Details in 🧵
- @datarescueproject.org
- Gina Plata-Nino of @fracposts.bsky.social
- @mapresearch.bsky.social & Williams Institute
- @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal
Details in 🧵
- @datarescueproject.org
- Gina Plata-Nino of @fracposts.bsky.social
- @mapresearch.bsky.social & Williams Institute
- @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal
Details in 🧵
Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tY
Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tY
Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow.
Surya OCR 2 (a 650M model!) does a very good job!
Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow.
Surya OCR 2 (a 650M model!) does a very good job!
@europeana.bsky.social newspapers.
Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
@europeana.bsky.social newspapers.
Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
400+ professionals, one shared position - cultural heritage should shape AI, not just feed it.
👉 bit.ly/4xLdyU3
#PublicAI #AI #ArtificialIntelligence #CulturalHeritage #DataSpace
400+ professionals, one shared position - cultural heritage should shape AI, not just feed it.
👉 bit.ly/4xLdyU3
#PublicAI #AI #ArtificialIntelligence #CulturalHeritage #DataSpace
This is our history and we want to #SaveOurSigns
This is our history and we want to #SaveOurSigns
Thank you to everyone helping keep biodiversity knowledge free and open. Please read, share, and help us build momentum.
Thank you to everyone helping keep biodiversity knowledge free and open. Please read, share, and help us build momentum.
Paper 1: arxiv.org/html/2603.02...
Paper 1: arxiv.org/html/2603.02...