Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢
🧵⬇️
Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢
🧵⬇️
RLHF = 30% *more* copying than base!
Awesome work from the awesome Ximing Lu (gloriaximinglu.github.io) et al. 🤩
arxiv.org/pdf/2410.04265
RLHF = 30% *more* copying than base!
Awesome work from the awesome Ximing Lu (gloriaximinglu.github.io) et al. 🤩
arxiv.org/pdf/2410.04265
posttraining
post-posttraining
neo training
training revival
trainingwave
posttraining
post-posttraining
neo training
training revival
trainingwave
OpenAI has best post-training/rl and has pushed it super hard on weaker pretraining.
Gemini has spectacular pretraining. Making a reasoning model was super easy for them & OpenAI folks were surprised
Anthropic? Secretive i guess.
OpenAI has best post-training/rl and has pushed it super hard on weaker pretraining.
Gemini has spectacular pretraining. Making a reasoning model was super easy for them & OpenAI folks were surprised
Anthropic? Secretive i guess.
This means it's also a new base model.
Glory days of pretraining still very much going.
This means it's also a new base model.
Glory days of pretraining still very much going.
I've long believed genomic foundation models are not as useful as claimed. In my mind, there isn't enough training data to justify their size. Interesting to see more work in this direction.
www.biorxiv.org/content/10.1...
I've long believed genomic foundation models are not as useful as claimed. In my mind, there isn't enough training data to justify their size. Interesting to see more work in this direction.
www.biorxiv.org/content/10.1...
A preprints reports that such models lack understanding of genomics and provide minimal utility, even for basic tasks such as sequence classification.
www.biorxiv.org/content/10.1...
A preprints reports that such models lack understanding of genomics and provide minimal utility, even for basic tasks such as sequence classification.
www.biorxiv.org/content/10.1...
↳ (examples in thread)
↳ (examples in thread)
is pretraining more like evolution (creation of genetic information) or childhood?
How we train an open everything model on a new pretraining environment with releasable data (Common Corpus) with an open source framework (Nanotron from HuggingFace).
www.sciencedirect.com/science/arti...
In our new @icmlconf.bsky.social paper, we address this gap!
📅 July 9th, Poster #4502 Session 8!
🧵
arxiv.org/pdf/2602.20062
In our new @icmlconf.bsky.social paper, we address this gap!
📅 July 9th, Poster #4502 Session 8!
🧵
arxiv.org/pdf/2602.20062
arxiv.org/abs/2609.30063
arxiv.org/abs/2609.30063