First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models.
huggingface.co/datasets/uv-...
First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models.
huggingface.co/datasets/uv-...
Search 1.49 million images from British Library books and Britannica (1500s–1920s).
Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.
huggingface.co/spaces/davan...
Search 1.49 million images from British Library books and Britannica (1500s–1920s).
Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.
huggingface.co/spaces/davan...
Search 1.49 million images from British Library books and Britannica (1500s–1920s).
Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.
huggingface.co/spaces/davan...
Search 1.49 million images from British Library books and Britannica (1500s–1920s).
Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.
huggingface.co/spaces/davan...
Unsurprisingly, it doesn’t do brilliantly overall: 36.8%
But for a ~16M-parameter recogniser, I think 74.4% on long/tiny text and 57.9% on multi-column pages are pretty interesting.
On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.
huggingface.co/spaces/fineb...
Unsurprisingly, it doesn’t do brilliantly overall: 36.8%
But for a ~16M-parameter recogniser, I think 74.4% on long/tiny text and 57.9% on multi-column pages are pretty interesting.
huggingface.co/deepseek-ai/...
- Asymmetric Causal-Encoder-Decoder: 550B MoE, input 8B / output 16B
- Native vision merged into one endpoint
- KV cache crushed: ~1/4 the HBM vs last one, 437× smaller than their first model
huggingface.co/deepseek-ai/...
- Asymmetric Causal-Encoder-Decoder: 550B MoE, input 8B / output 16B
- Native vision merged into one endpoint
- KV cache crushed: ~1/4 the HBM vs last one, 437× smaller than their first model
There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.
I’ve gathered 41 models into four collections, with short notes to help you choose:
huggingface.co/collections/...
There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.
I’ve gathered 41 models into four collections, with short notes to help you choose:
huggingface.co/collections/...
On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.
huggingface.co/spaces/fineb...
On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.
huggingface.co/spaces/fineb...
As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find.
Thread 🧵
As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find.
Thread 🧵
huggingface.co/datasets/big...
huggingface.co/datasets/big...
One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
Step one: work out which modern OCR models are actually good enough.
Step one: work out which modern OCR models are actually good enough.
You can also do semantic search against the images here: huggingface.co/spaces/davan...
You can also do semantic search against the images here: huggingface.co/spaces/davan...
Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30.
huggingface.co/datasets/hug...
Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30.
huggingface.co/datasets/hug...
So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments.
Search "typing on a computer keyboard", land on the second it happens.
So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments.
Search "typing on a computer keyboard", land on the second it happens.
Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise
~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM).
huggingface.co/datasets/uv-...
Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise
~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM).
huggingface.co/datasets/uv-...
Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.
huggingface.co/datasets/uv-scripts/ocr
Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.
huggingface.co/datasets/uv-scripts/ocr
Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.
huggingface.co/datasets/uv-scripts/ocr
Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.
huggingface.co/datasets/uv-scripts/ocr
An OCR Benchmark for Historians, 1612–1921
Here is version 1 of a working paper on the new OCR tools that are transforming digital history.
working-papers-in-critical-search.github.io/paper-004-oc...
An OCR Benchmark for Historians, 1612–1921
Here is version 1 of a working paper on the new OCR tools that are transforming digital history.
working-papers-in-critical-search.github.io/paper-004-oc...
huggingface.co/spaces/davan...
huggingface.co/spaces/davan...
huggingface.co/spaces/davan...
The ranking flips depending on what you actually want.
The ranking flips depending on what you actually want.
They're searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces...
Now there's public data: each agent's share of Hub traffic, updated monthly 👇
They're searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces...
Now there's public data: each agent's share of Hub traffic, updated monthly 👇