#Grobid
this reference checker is built on top of lots of excellent open source work, like GROBID, so the plan is to tidy it up and make it freely available at some point (when I have time to write up some documentation!)
August 31, 2026 at 10:19 PM
To what extent do researchers funded by Dutch Research Council NWO and ZonMw share the research data and code underlying their publications?

Today we published an analysis based on 10.000+ papers using the open source tool Grobid: www.nwo.nl/en/news/shar...

All underlying data openly available!
February 10, 2025 at 9:10 PM
At #TPDL2026 in Faro ☀️:

🔹 Today, Michael Paris presents our Common Crawl paper on what web crawlers actually see
🔹 Thu afternoon I demo the INRIA DataLake: HAL papers → structured knowledge + software mentions

Come say hi! #DigitalLibraries #CommonCrawlFoundation #Grobid
September 23, 2026 at 6:36 AM
Luca Foppiano, Sana Khamassi, Vipul Gupta
Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion
https://arxiv.org/abs/2609.26381
September 24, 2026 at 4:29 AM
It's not markdown, but there's Grobid to parse PDF to structured text. github.com/grobidOrg/gr...
GitHub - grobidOrg/grobid: A machine learning software for extracting information from scholarly documents
A machine learning software for extracting information from scholarly documents - grobidOrg/grobid
github.com
March 31, 2026 at 10:17 AM
Peut-on parler de GROBID ? 😊
almanach.inria.fr/software_and...
June 23, 2025 at 6:28 PM
Luca Foppiano, Sana Khamassi, Vipul Gupta: Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion https://arxiv.org/abs/2609.26381 https://arxiv.org/pdf/2609.26381 https://arxiv.org/html/2609.26381
September 23, 2026 at 6:40 AM
The 2022 anomaly is probably an artifact of changing how they noted the badges in the PDFs. I'll look into that tomorrow. I'm using grobid to parse the original papers, and an under-development R package called papercheck to process and search the grobid XMLs.
April 2, 2025 at 12:07 AM
yeah, I've been battling that for the longest time, it's not perfect, but it's doing better than existing deterministic tools (GROBID, etc.). I'll be continuing to improve it.

PDF extraction is also in ScienceArena, I'm comparing DocPluck to GROBID and LLMs, it's doing ok: sciencearena.vercel.app
August 4, 2026 at 12:11 PM
Finally, papers on chive.pub now show which other papers cite them, as a graph you can move around in. We extract references from each PDF with grobid (github.com/grobidOrg/grobid) and match them against what's indexed.
September 8, 2026 at 3:40 PM
"FWCI: ... guarantees average FWCI = 1.0 within any cohort." Good that they fixed this, was noticing this bug. Also INTERESTING
"Content API. Direct access to 60M+ open-access PDFs and parsed GROBID TEI XML at predictable URLs like content.openalex.org/works/{id}"
April 27, 2026 at 3:54 PM
5/ Infrastructure: JDK 21, Gradle 9, TensorFlow 2.17 (Python 3.10–3.11), pdfalto 0.6.0, wapiti 1.5.1, virtualenv/conda support for DeLFT.
6/ Full release notes → github.com/kermitt2/...

#GROBID #OpenSource #NLP #ScholarlyInfrastructure
6/6
Release 0.9.0 · grobidOrg/grobid
What's Changed Added Conflict of interest and author contributions statement extraction in header and segmentation models #1319 Extract figures, tables and equations from back/annex sections #...
github.com
April 14, 2026 at 1:36 PM
I can confirm that extracting references from papers is not without error. We’ve been using grobid to parse PDFs for metacheck, and it frequently merges references, especially in 2-column formats. Then a check for the existence of the reference fails. But also the data in crossref is not infallible.
May 16, 2026 at 9:13 AM
Thank to every contributor and user 🙏 for the support!

github.com/grobidOrg...
2/2
GitHub - grobidOrg/grobid: A machine learning software for extracting information from scholarly documents
A machine learning software for extracting information from scholarly documents - grobidOrg/grobid
github.com
July 22, 2026 at 2:25 PM
I’ll add you to the list! We’re especially keen to agree on a common format for extracting metadata and full text from PDFs of research papers. Grobid does an OK job at this, but we’re developing something better and want the metadata format to be as widely useful as possible.
September 2, 2026 at 7:08 PM
Python client release note: github.com/grobidOrg...

And.. Grobid 0.9.1 release note (in case you missed it) github.com/grobidOrg...

🖖
2/2
Release v0.2.0 · grobidOrg/grobid-client-python
What's Changed Add type hints and py.typed marker (PEP 561) by @lfoppiano in #112 Fixed offsets crash on Markdown transformation: removing the html.unescape and repetitive… by @Sanakhamassi in...
github.com
August 23, 2026 at 8:41 AM
Weeknotes time! Spent a few days sorting out petabytes of our uni infrastructure, ranging from the ever-ballooning TESSERA and to our growing living evidence database of papers. Also some really fun RA jobs just opened for Deep Learning with SDMs, if youre job hunting 🌍 anil.recoil.org/notes/2026w31
.plan-26-31: Sorting out Tessera and Evidence TAP infrastructure
A petabyte of TESSERA embeddings moves to Source Cooperative, and Taposaur's GROBID metadata index and capability-based downloader take shape for Evidence TAP, while Eio gets some native Windows support.
anil.recoil.org
August 3, 2026 at 8:14 PM
Length,string,semantic similarity?Preprints: ArXiv PDF. Full text: CrossRef API q w/ DOI. All w/ GROBID to TEI XML. https://twitter.com/hvdsomp/status/745405554055614464
November 12, 2024 at 5:21 PM


Wanna join the community?
- Announcement mailing list: grobid.netlify.app/mail
- Discord: grobid.netlify.app/d...

... and don't forget to star us on Github

#GROBID #NLP #TextMining #PDF #OpenScience #TDM #MachineLearning
3/3
August 7, 2026 at 2:33 PM
Since the PDF to grobid conversion only needs to be done once per paper and then everything else can work off the created XML, I agree accuracy is more important than speed for this one.
July 1, 2025 at 8:06 AM
Open science can also mean to make bibliographic metadata openly available. In a new open-bibliometrics post, Paul Donner (DZHW) shows that GROBID can extract many references often with surprisingly good results — but better, domain-specific training data will be needed. tinyurl.com/ys46cram
November 18, 2025 at 8:36 AM
The reference stuff needs a fair amount of work, as it inherits any inaccuracies in grobid’s categorisation of references and citations. We’re looking at using grobid with citation consolidation for crossref to improve this, but then pdf to xml conversion is a lot slower.
June 28, 2025 at 4:27 PM
On the 26-27 November we held the #Grobid Camp at the Centre de #Inria Paris.
The goal was to have a meeting with the major players in the French community which spaces from government institutes, to companies and large scale projects.
1/4
December 8, 2025 at 12:59 PM
- 🔤 Improved recognition of non-standard fonts
- 🛠️ Various bug fixes and security vulnerabilities addressed

github.com/kermitt2/...
🔽
Release 0.8.2 · kermitt2/grobid
What's Changed Added New model specialization/variants (flavors) mechanism #1151 Specialization/variant process for a lightweight processing that covers other types of scientific articles that...
github.com
May 18, 2025 at 8:27 AM
"SoFAIR will extend the capabilities of widely used open scholarly infrastructures (CORE, Software Heritage, HAL) and tools (GROBID) operated by the consortium partners, delivering and deploying an effective solution for the management of the research software lifecycle"

Making Software FAIR […]
Original post on code4lib.social
code4lib.social
January 24, 2025 at 12:16 PM