#PreTraining
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this:

Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢

🧵⬇️
November 20, 2024 at 4:35 PM
pretraining is just distilling from humanity
ponder.ooo ponder @ponder.ooo · May 15
"distillation ATTACKS" is such a funny term for it from the "information loves being free your data wanted to be in my model actually" crowd
May 15, 2026 at 11:27 AM
LLMs generate novel word sequences not contained in their pretraining data. However, compared to humans, models generate significantly fewer novel n-grams.

RLHF = 30% *more* copying than base!

Awesome work from the awesome Ximing Lu (gloriaximinglu.github.io) et al. 🤩

arxiv.org/pdf/2410.04265
November 22, 2024 at 6:14 AM
Getting into pretraining has never been cheaper.
November 15, 2025 at 12:18 PM
How much did human pretraining cost?
December 10, 2025 at 3:02 PM
pretraining
posttraining
post-posttraining
neo training
training revival
trainingwave
March 16, 2026 at 7:13 PM
It took me weeks, but finally it's there: an overlong blogpost on synthetic pretraining. vintagedata.org/blog/posts/s...
February 1, 2026 at 5:53 PM
WHO CONTAMINATED MY PRETRAINING DATASET
.
.
.
YOU
November 23, 2024 at 5:55 AM
Recurring frontier lab gossip:

OpenAI has best post-training/rl and has pushed it super hard on weaker pretraining.

Gemini has spectacular pretraining. Making a reasoning model was super easy for them & OpenAI folks were surprised

Anthropic? Secretive i guess.
October 12, 2025 at 7:26 PM
Opus 4.7 has a new tokenizer.
This means it's also a new base model.
Glory days of pretraining still very much going.
April 16, 2026 at 2:45 PM
Genomic Foundationless Models: Pretraining Does Not Promise Performance

I've long believed genomic foundation models are not as useful as claimed. In my mind, there isn't enough training data to justify their size. Interesting to see more work in this direction.

www.biorxiv.org/content/10.1...
Genomic Foundationless Models: Pretraining Does Not Promise Performance
The success of Large Language Models has inspired the development of Genomic Foundation Models (GFMs) through similar pretraining techniques. However, the relationship between pretraining performance ...
www.biorxiv.org
February 2, 2025 at 3:53 PM
ILYA: "PRETRAINING IS DONE. WE ARE NOW IN THE POST TRAINING ERA."
December 13, 2024 at 10:49 PM
is there a general term for the things you do to a model after pretraining? because post-pretraining is an awkward construction.
as a secondary loss during several post-pretraining operations, constraining the logits over vocabulary to the reference distribution from pretraining.
November 27, 2025 at 8:06 PM
The success of LLMs has inspired their extension to genomics data.

A preprints reports that such models lack understanding of genomics and provide minimal utility, even for basic tasks such as sequence classification.

www.biorxiv.org/content/10.1...
Genomic Foundationless Models: Pretraining Does Not Promise Performance
The success of Large Language Models has inspired the development of Genomic Foundation Models (GFMs) through similar pretraining techniques. However, the relationship between pretraining performance ...
www.biorxiv.org
February 1, 2025 at 11:41 PM
Extensive lecture of Eric W. Tramel (Nvidia) on synthetic data: "Synthetic pretraining is the way frontier models are built" scalable-ai.eecs.berkeley.edu/assets/lectu...
March 15, 2026 at 5:54 PM
This is shifting me from "alignment is going pretty well" under the previous domain where pretraining was the bulk, to "alignment is absolutely doomed."
metr.org METR @metr.org · Aug 26
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
August 27, 2026 at 3:46 AM
What are your favorite papers on tokenizers (for pretraining data processing) and their downstream effects?

↳ (examples in thread)
March 4, 2025 at 10:19 AM
Pretraining is like sculpting a statue by shooting a block of marble with a pellet gun 10 million times
armchair experts only:

is pretraining more like evolution (creation of genetic information) or childhood?
December 10, 2025 at 2:31 PM
One of the reasons this will be valuable for academics is that it provides a very clear recipe for pretraining ~3B-scale models on custom corpora.
And new paper out: Pleias 1.0: the First Family of Language Models Trained on Fully Open Data

How we train an open everything model on a new pretraining environment with releasable data (Common Corpus) with an open source framework (Nanotron from HuggingFace).

www.sciencedirect.com/science/arti...
September 27, 2025 at 3:15 PM
Pretraining + fine-tuning powers modern ML, but we lack a theoretical understanding of how pretraining actually shapes downstream learning.

In our new @icmlconf.bsky.social paper, we address this gap!

📅 July 9th, Poster #4502 Session 8!
🧵

arxiv.org/pdf/2602.20062
July 8, 2026 at 4:11 PM
another adversarial self play paper. in this one they use a generator to create programs that output sequences that are medium-hard for a model to predict, as in causal language modelling. all meaning in the model is bootstrapped from an initially random policy
arxiv.org/abs/2609.30063
September 25, 2026 at 6:38 PM
The Practitioner's Guide to Continual Multimodal Pretraining @dziadzio.bsky.social @confusezius.bsky.social @vishaalurao.bsky.social @bayesiankitten.bsky.social
December 12, 2024 at 2:20 AM
For example. See paper for more. And no it's not just about whether something's in the pretraining ;) (but also, think about what it would mean for a system to match Jabberwocky text to pretraining data! arxiv.org/abs/2601.11432
January 19, 2026 at 3:42 AM
interacting with a LLM after just pretraining is a profoundly destabilizing experience and the closest thing to it you can experience outside a tech company is interacting with o3.
July 19, 2025 at 3:45 AM