#Pretraining
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this:

Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢

🧵⬇️
November 20, 2024 at 4:35 PM
besides, it’s only pretraining. just wait for the real training
September 28, 2026 at 11:41 PM
pretraining is just distilling from humanity
ponder.ooo ponder @ponder.ooo · May 15
"distillation ATTACKS" is such a funny term for it from the "information loves being free your data wanted to be in my model actually" crowd
May 15, 2026 at 11:27 AM
LLMs generate novel word sequences not contained in their pretraining data. However, compared to humans, models generate significantly fewer novel n-grams.

RLHF = 30% *more* copying than base!

Awesome work from the awesome Ximing Lu (gloriaximinglu.github.io) et al. 🤩

arxiv.org/pdf/2410.04265
November 22, 2024 at 6:14 AM
Getting into pretraining has never been cheaper.
November 15, 2025 at 12:18 PM
How much did human pretraining cost?
December 10, 2025 at 3:02 PM
pretraining
posttraining
post-posttraining
neo training
training revival
trainingwave
March 16, 2026 at 7:13 PM
It took me weeks, but finally it's there: an overlong blogpost on synthetic pretraining. vintagedata.org/blog/posts/s...
February 1, 2026 at 5:53 PM
124M may sound tiny to users but for a 40 second NanoGPT pretraining run that's fucking humungous
September 29, 2026 at 11:54 AM
WHO CONTAMINATED MY PRETRAINING DATASET
.
.
.
YOU
November 23, 2024 at 5:55 AM
Recurring frontier lab gossip:

OpenAI has best post-training/rl and has pushed it super hard on weaker pretraining.

Gemini has spectacular pretraining. Making a reasoning model was super easy for them & OpenAI folks were surprised

Anthropic? Secretive i guess.
October 12, 2025 at 7:26 PM
Opus 4.7 has a new tokenizer.
This means it's also a new base model.
Glory days of pretraining still very much going.
April 16, 2026 at 2:45 PM
Genomic Foundationless Models: Pretraining Does Not Promise Performance

I've long believed genomic foundation models are not as useful as claimed. In my mind, there isn't enough training data to justify their size. Interesting to see more work in this direction.

www.biorxiv.org/content/10.1...
Genomic Foundationless Models: Pretraining Does Not Promise Performance
The success of Large Language Models has inspired the development of Genomic Foundation Models (GFMs) through similar pretraining techniques. However, the relationship between pretraining performance ...
www.biorxiv.org
February 2, 2025 at 3:53 PM
ILYA: "PRETRAINING IS DONE. WE ARE NOW IN THE POST TRAINING ERA."
December 13, 2024 at 10:49 PM
Anthropic notoriously avoids relying on algorithms and fancy LLM architectures

it’s no mistake that they hired Andrej Karpathy, who created nanoGPT, to work on RSI in pretraining
September 29, 2026 at 11:01 AM
is there a general term for the things you do to a model after pretraining? because post-pretraining is an awkward construction.
as a secondary loss during several post-pretraining operations, constraining the logits over vocabulary to the reference distribution from pretraining.
November 27, 2025 at 8:06 PM
The success of LLMs has inspired their extension to genomics data.

A preprints reports that such models lack understanding of genomics and provide minimal utility, even for basic tasks such as sequence classification.

www.biorxiv.org/content/10.1...
Genomic Foundationless Models: Pretraining Does Not Promise Performance
The success of Large Language Models has inspired the development of Genomic Foundation Models (GFMs) through similar pretraining techniques. However, the relationship between pretraining performance ...
www.biorxiv.org
February 1, 2025 at 11:41 PM
Extensive lecture of Eric W. Tramel (Nvidia) on synthetic data: "Synthetic pretraining is the way frontier models are built" scalable-ai.eecs.berkeley.edu/assets/lectu...
March 15, 2026 at 5:54 PM
This is shifting me from "alignment is going pretty well" under the previous domain where pretraining was the bulk, to "alignment is absolutely doomed."
metr.org METR @metr.org · Aug 26
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
August 27, 2026 at 3:46 AM
What are your favorite papers on tokenizers (for pretraining data processing) and their downstream effects?

↳ (examples in thread)
March 4, 2025 at 10:19 AM
Pretraining is like sculpting a statue by shooting a block of marble with a pellet gun 10 million times
armchair experts only:

is pretraining more like evolution (creation of genetic information) or childhood?
December 10, 2025 at 2:31 PM
One of the reasons this will be valuable for academics is that it provides a very clear recipe for pretraining ~3B-scale models on custom corpora.
And new paper out: Pleias 1.0: the First Family of Language Models Trained on Fully Open Data

How we train an open everything model on a new pretraining environment with releasable data (Common Corpus) with an open source framework (Nanotron from HuggingFace).

www.sciencedirect.com/science/arti...
September 27, 2025 at 3:15 PM
After a long wait, releasing the SYNTH paper!

It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency. arxiv.org/abs/2609.378...
October 1, 2026 at 3:02 PM
Pretraining + fine-tuning powers modern ML, but we lack a theoretical understanding of how pretraining actually shapes downstream learning.

In our new @icmlconf.bsky.social paper, we address this gap!

📅 July 9th, Poster #4502 Session 8!
🧵

arxiv.org/pdf/2602.20062
July 8, 2026 at 4:11 PM