#Pretraining
Huimin Pan, Yufan Ren, Kunpeng Song, Siyang Wang, Xiwen Zhang, Xiaoyun Hu, Zhuoxu Duan, Hanrui Zheng, Jialeng Ni, Nathan Zhao, Sibo Ma, Zhenxuan Fan, ...
IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining
https://arxiv.org/abs/2609.39403
October 3, 2026 at 6:58 AM
I also think you're undercounting the power of reinforcement learning, especially combined with reasoning. Now the LLM isn't just trying to reproduce a probability distribution, it's trying to generate the next token *most likely to lead to the correct answer*, even for problems not in pretraining
October 3, 2026 at 2:14 AM
It's not so distantly related because of the KL mentality. It's still very close to the pretraining distribution. But yes it is of course not identical to the pretraining distribution; that's the point.
October 3, 2026 at 1:15 AM
Video Generation Models: A Survey of Post-Training and Alignment – Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learne... https://tinyurl.com/26px543w #SurveyAI
Video Generation Models: A Survey of Post-Training and Alignment
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maint…
arxiv.org
October 3, 2026 at 12:56 AM
Then there's the question of why they would spend all the resources for pretraining a new, such a poor and already outdated model?

Especially since many people have now proven in practice that you can in fact easily repurpose existing LLMs quickly and beat them in their own game.
October 2, 2026 at 10:22 PM
And even the folks doing pretraining have to ge extremely efficient about how they do it. So I feel everyone intuitively gets that we are not going to resource this if we keep doing what we have been doing but the silicon side is slower to move & photonics needs to get moving faster.
October 2, 2026 at 4:57 PM
Most of the new model work - tbh - is distillation type work. Outside the OpenAI-Anthropic-Google space - there isnt a push for pretraining. I suppose this could chamge as VLA models become more popular & self driving cars become more common.
October 2, 2026 at 4:55 PM
New insights illustrate how pretraining and midtraining enhance AI models' learning from rewards by clarifying mechanisms and computation, achieving impressive task success across various learning contexts. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 2, 2026 at 3:40 PM
🔧 Pretraining language discrimination seems to effectively mitigate the multilingual penalty.
October 2, 2026 at 3:35 PM
Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal \v{S}tef\'anik, Aline Villavicencio, Nikolaos Aletras
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
https://arxiv.org/abs/2609.39827
October 2, 2026 at 8:48 AM
Research shows pretraining and midtraining enhance reward learning by clarifying mechanisms and developing key computations. This improves task accuracy, reflecting that effective learning depends on insights from earlier training stages along with rewards. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 2, 2026 at 7:10 AM
Mohan Shi, Ruchao Fan, Sunit Sivasankaran, Keqi Deng, Jinyu Li: Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs https://arxiv.org/abs/2610.01695 https://arxiv.org/pdf/2610.01695 https://arxiv.org/html/2610.01695
October 2, 2026 at 6:45 AM
Shengye Tao, Yinzhu Cheng, Haihua Xie: Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining https://arxiv.org/abs/2610.01165 https://arxiv.org/pdf/2610.01165 https://arxiv.org/html/2610.01165
October 2, 2026 at 6:44 AM
Xu, Wu, Chen, Jiang, Ye, Jiang, Zhang, Feng, Xu, Chen, Lu, Hoi: Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining https://arxiv.org/abs/2610.00438 https://arxiv.org/pdf/2610.00438 https://arxiv.org/html/2610.00438
October 2, 2026 at 6:44 AM
October 2, 2026 at 6:42 AM
A study illustrates how pretraining and midtraining empower language models to learn from rewards by uncovering hidden mechanisms and enhancing computation. It shows notable gains in task-specific accuracy, underscoring reinforcement learning's role in AI reasoning. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 2, 2026 at 5:30 AM
Research shows that pretraining and midtraining in language models enhance reward learning by uncovering hidden mechanisms and computations, achieving 82.61% accuracy against lower rates in standard strategies, signifying key advances in reinforcement learning. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 2, 2026 at 5:30 AM
Emergent Object Binding Has a Finite Spatial Horizon
tech_blogs_arxiv | Author: Mayank Singal

#MachineLearning
Emergent Object Binding Has a Finite Spatial Horizon
Pretrained Vision Transformers encode whether two image patches belong to the same object. This IsSameObject signal is decodable from frozen patch embeddings at high accuracy, which suggests that object binding emerges from self-supervised pretraining alone. We show that this single accuracy number
arxiv.org
October 2, 2026 at 4:06 AM
Excited to be heading to COLM to present our work on memorization x distillation arxiv.org/abs/2601.15394 come say hi :)

Happy to talk about: memorization, safety, pretraining on synthetic data, etc.

I’ll be also on the industry job market later this fall, would love to grab coffee and chat! ✨
October 2, 2026 at 3:03 AM
Research reveals how pretraining and midtraining enhance a model's ability to learn from rewards by clarifying ambiguous computational mechanisms and boosting task-specific performance—achieving high accuracy across various tasks. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 1, 2026 at 8:50 PM
A study shows how pretraining and midtraining enhance models by uncovering mechanisms that boost learning from rewards, with strategies in tasks achieving 82%, highlighting reinforcement learning's benefits for foundational models. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 1, 2026 at 8:40 PM
This research shows how pretraining and midtraining enhance language models' reward learning by exposing hidden mechanisms that improve task performance. It underscores connections between prediction, acquisition, and task execution, enhancing AI training methods. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 1, 2026 at 8:30 PM
The 'controlled environments for pretraining science' angle is the real contribution — synthetic data that lets you ablate exactly what the model learned. That's how you turn pretraining from alchemy into science.
October 1, 2026 at 8:23 PM
A study shows how pretraining and midtraining enhance language models' ability to learn from rewards by clarifying mechanisms and building vital computations. Models achieve over 82% accuracy in complex tasks, paving pathways for optimizing reinforcement learning. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 1, 2026 at 8:20 PM
Research reveals pretraining enhances models' reward learning by uncovering hidden mechanisms. Findings using sequential computation and contextual memory show major task-specific accuracy gains, marking a breakthrough in reinforcement learning for language models. https://arxiv.org/abs/2609.38446
What Pretraining and Midtraining Make Learnable from Rewards?
ArXiv link for What Pretraining and Midtraining Make Learnable from Rewards?
arxiv.org
October 1, 2026 at 8:10 PM