#OLMoTrace
No, I didn’t know either … arxiv.org/abs/2504.07096
March 26, 2026 at 7:28 AM
Has anyone played around with OlmoTrace? Folks I'm fascinated by it. What an amazing tool, for research and teaching.
March 25, 2026 at 11:35 PM
If you get bored you can play with OLMoTrace.
Couple of minutes should illustrate how it's not referencing training data, but literally doing searches.
Going beyond open data – increasing transparency and trust in language models with OLMoTrace | Ai2
Ai2, a non-profit research institute founded by Paul Allen, is committed to breakthrough AI to solve the world’s biggest problems.
allenai.org
September 16, 2026 at 5:08 PM
olmoTrace for connecting model generations to training data won Best Paper for System Demonstrations at #ACL2025!
Today we're unveiling OLMoTrace, a tool that enables everyone to understand the outputs of LLMs by connecting to their training data.

We do this on unprecedented scale and in real time: finding matching text between model outputs and 4 trillion training tokens within seconds. ✨
For years it’s been an open question — how much is a language model learning and synthesizing information, and how much is it just memorizing and reciting?

Introducing OLMoTrace, a new feature in the Ai2 Playground that begins to shed some light. 🔦
July 31, 2025 at 8:33 PM
OLMoTrace, from A2I, is really useful-- see the source training documents that contain strings from the output of the LLM. Doing it with generated stories is especially enlightening (even when we know it's often regurgitating). playground.allenai.org
April 10, 2025 at 8:05 AM
[immediately opens browser and searches for OlmoTrace...]
March 26, 2026 at 2:22 AM
For years it’s been an open question — how much is a language model learning and synthesizing information, and how much is it just memorizing and reciting?

Introducing OLMoTrace, a new feature in the Ai2 Playground that begins to shed some light. 🔦
April 9, 2025 at 1:16 PM
Learn more about how OLMoTrace works on our blog: allenai.org/blog/olmotrace
Going beyond open data – increasing transparency and trust in language models with OLMoTrace | Ai2
OLMoTrace lets you trace the outputs of language models back to their full, multi-trillion-token training data in real time.
allenai.org
April 9, 2025 at 1:16 PM
Today we're unveiling OLMoTrace, a tool that enables everyone to understand the outputs of LLMs by connecting to their training data.

We do this on unprecedented scale and in real time: finding matching text between model outputs and 4 trillion training tokens within seconds. ✨
For years it’s been an open question — how much is a language model learning and synthesizing information, and how much is it just memorizing and reciting?

Introducing OLMoTrace, a new feature in the Ai2 Playground that begins to shed some light. 🔦
April 9, 2025 at 1:37 PM
@simonwillison.net have you looked at OLMoTrace yet?

allenai.org/blog/olmotrace

It’s gotten surprisingly little coverage, yet it highlights an incredible feature that the closed models by definition cannot do, namely reverse-lookups of original-source training materials.

bsky.app/profile/erle...
@anthropic.com would be really cool to see an experimental build of Claude that is limited to the same open datasets as OLMo by @ai2.bsky.social

Seeing that version of Claude benchmarked against its commercial counterpart would no doubt be enlightening.
May 20, 2025 at 9:45 AM
✍️ Read the blog about OLMoTrace: buff.ly/ACEl3v5
📝 Review the paper: buff.ly/sWHVmI3
💻 Download the code: buff.ly/W3dsBHt
Going beyond open data – increasing transparency and trust in language models with OLMoTrace | Ai2
OLMoTrace lets you trace the outputs of language models back to their full, multi-trillion-token training data in real time.
allenai.org
June 30, 2025 at 5:37 PM
You're having a weird conversation with yourself that doesn't exist here

Olmotrace has nothing to do with agents or plagiarism
September 16, 2026 at 9:01 PM
Last week we released OLMoTrace as part of #GoogleCloudNext
April 14, 2025 at 7:31 PM
OLMoTrace connects phrases or even whole sentences in the language model’s output back to verbatim matches in its training data. It does this by searching billions of documents and trillions of tokens in real time and highlighting where it finds compelling matches.
April 9, 2025 at 1:16 PM
Ai2 launched a new tool where your responses from OLMo get mapped back to related training data. We're using this actively to improve our post-training data and hope many others will use it for understanding and transparency around leading language models!
Some musings:
Looking at the training data
On building tools where truly open-source models can shrine (OLMo 2 32B Instruct, for today). OLMoTrace lets you poke around.
buff.ly
April 9, 2025 at 8:12 PM
Demystifying AI Decision-Making: Ai2 Unveils Open-Source Tool to Trace LLM Outputs to Training Data

venturebeat.com/ai/whats-ins...
What’s inside the LLM? Ai2 OLMoTrace will ‘trace’ the source
Ai2's new open-source OLMoTrace tool allows enterprises to directly trace LLM outputs back to original training data, bringing transparency to AI decision-making and addressing trust barriers.
venturebeat.com
April 11, 2025 at 1:46 AM
Try OLMoTrace in the Ai2 Playground today: playground.allenai.org
playground.allenai.org
April 9, 2025 at 1:16 PM
OLMoTrace is useful for fact checking✅, understanding hallucinations🎃, tracing reasoning capabilities🧠, or just generally helping you see where an LLMs response may have come from.
April 9, 2025 at 1:16 PM
That can easily be done. It's already being done by the people working at the Paul Allen Institute. Does anyone actually believe that Elon Musk, Mark Zuckerberg and Sam Altman will agree to do this? Open source is fighting for AI democracy.
allenai.org/blog/olmotrace
Going beyond open data – increasing transparency and trust in language models with OLMoTrace | Ai2
OLMoTrace lets you trace the outputs of language models back to their full, multi-trillion-token training data in real time.
allenai.org
July 26, 2025 at 9:15 PM
Yup this is a massive loss. OLMo (+entire ecosystem, OLMoTrace, the NeurIPS tutorial on the LM pipeline,...) was incredibly valuable and now I feel I took it for granted all this time. Not even counting all the great research coming out of AllenAI.
March 24, 2026 at 9:12 PM
How can we better understand how models make predictions and which components of a training dataset are shaping their behaviors? In April we introduced OLMoTrace, a feature that lets you trace the outputs of language models back to their full training data in real time. 🧵
June 30, 2025 at 5:37 PM
Since then, we’ve improved on it with OLMo 1.7, OLMo 2, and now OLMo 2 32B. We’ve added additional transparency through OLMoTrace, and efficiency with OLMoE — and there’s more to come.
May 6, 2025 at 8:55 PM
Using the OlmoTrace LLM, trained without violating copyrights, one can see exactly which part of the training data led to particular pieces of output, according to the AI2 folks. Interesting if that could be retrofit to the commercial LLMs, to see how much straight-up plagiarism there is.
April 14, 2025 at 12:22 AM