#PreTraining
EMGBlend: Heterogeneity-Aware Self-Supervised Pretraining for Gesture and Force Decoding
Read more: https://arxiv.org/html/2609.25582v1
September 26, 2026 at 12:42 PM
Transfer Learning using ahead::ridge2f on synthetic stocks returns Pt.2: synthetic data generation
https://thierrymoudiki.github.io/blog/2025/09/09/r/python/pretraining-ridge2f-part2

#Python #DataScience #MachineLearning #rstats #Techtonique
September 26, 2026 at 12:35 PM
NASA says the model learned from roughly two million lunar image tiles, then highlighted a rocket-impact crater in imagery excluded from pretraining. Public code, datasets and benchmarks make that result easier to reproduce.
September 26, 2026 at 12:19 PM
September 26, 2026 at 7:41 AM
⛵️人間が用意したデータに頼らず、モデル自身が万能チューリングマシン上で自己対戦しながら学習データを生成する新手法を提案。
⛵️自然データ未学習でもゼロショット性能が計算量に応じ予測的に向上することを実証。
arxiv.org/abs/2609.30063
Self-Play Pretraining with Zero Data
Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining w...
arxiv.org
September 26, 2026 at 7:09 AM
Xinge Guo, Yuanhao Wang, Liqi Shu, Yang Liu, Min Xu
ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences
https://arxiv.org/abs/2609.30043
September 26, 2026 at 6:04 AM
Researchers present "spectral deflation," enhancing matrix filtering in deep learning and semidefinite programming. This innovation improves GPT-2 pretraining efficiency and reduces KKT residuals while preserving target accuracy without factorization. https://arxiv.org/abs/2609.21102
Spectral Deflation for Factorization-Free Matrix Filtering in Muon and Semidefinite Programming
ArXiv link for Spectral Deflation for Factorization-Free Matrix Filtering in Muon and Semidefinite Programming
arxiv.org
September 25, 2026 at 10:20 PM
After pretraining, they are indeed estimating P(next token | previous tokens). But this is not true after fine tuning or RL.
September 25, 2026 at 8:35 PM
No it's exactly what you've spent two days now defending, when you don't run back to the bailey of "it's all pretraining data"
September 25, 2026 at 8:04 PM
Why oh why would the SOTA labs bother generating all of that synthetic pretraining data if they could just rely on some emergent understanding?
September 25, 2026 at 7:44 PM
Their blog post example is just a demo of a general technique.

"RLHF has nothing to do with LLMs" is incorrect.

ChatGPT is different from the earlier GPT foundation models, because of supervised-fine tuning (SFT) and RLHF. Post-training suppressed the more unhinged responses of pure autocomplete.
September 25, 2026 at 7:41 PM
And if you don't know how much synthetic data the SOTA labs have been putting into pretraining, you haven't been paying attention.

It's one of the reasons that prompt injecting by impersonating CoT style works so well!
September 25, 2026 at 7:31 PM
The only reason it appears so effective in the latest models is because they're not relying on illusory conceptualisation from it just being RLHF; there are millions of direct examples in pretraining to interpolate between.
September 25, 2026 at 7:11 PM
Sorry, I don't see how that follows. Chain of thought is just more hypersyntax.

Also, in the latest models, synthetic training data full of chain of thought is now part of the pretraining corpus. Most of it, in fact.
September 25, 2026 at 7:11 PM
the generator makes a program that generates a single document for pretraining. and then the generator is rewarded...thusly

i think they overcomplicated this. if you don't want "arbitrarily difficult" you can just have a point on your hardness scale where reward starts going down
September 25, 2026 at 6:49 PM
another adversarial self play paper. in this one they use a generator to create programs that output sequences that are medium-hard for a model to predict, as in causal language modelling. all meaning in the model is bootstrapped from an initially random policy
arxiv.org/abs/2609.30063
September 25, 2026 at 6:38 PM
RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations

Chenhao Si, Ming Yan

#arXiv #cs.AI #cs.LG
RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations
Learning surrogates for time-dependent partial differential equations often requires a new simulation corpus when the governing operator changes. We introduce RD-JEPA, a joint-embedding predictive architecture for self-supervised pretraining on reaction-diffusion trajectories. A single model is pre…
arxiv.org
September 25, 2026 at 4:06 PM
📄 Read the full commentary:
https://www.imrpress.com/journal/RCM/27/9/10.31083/RCM53504#F002
✍️ Author:
Zhonghua Sun
September 25, 2026 at 3:55 PM
re: the orthogonality thesis in particular, what i am getting at is that it is not intuitively obvious to me that in a model pretrained on human values and asymptotically confined to its pretrained manifold, you are not confined to the submanifold of the pretraining under mutual support.
September 25, 2026 at 3:48 PM
RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations
Read more: https://arxiv.org/html/2609.29403v1
September 25, 2026 at 12:42 PM
Benjamin Hurt, Oscar O'Donnell: Is Broader Better? A Controlled Study of Multilingual Coverage and Pretraining Objective in Frozen SSL Encoders for Speech Deepfake Detection https://arxiv.org/abs/2609.29138 https://arxiv.org/pdf/2609.29138 https://arxiv.org/html/2609.29138
September 25, 2026 at 6:45 AM
Thang Tran, Lan Dang: Pretraining and adapting a language model on a dependency-free stack: GPT-2 124M from random weights, reproduced against llm.c, and a clinical adapter for Qwen3-0.6B https://arxiv.org/abs/2609.28568 https://arxiv.org/pdf/2609.28568 https://arxiv.org/html/2609.28568
September 25, 2026 at 6:44 AM
Xinge Guo, Yuanhao Wang, Liqi Shu, Yang Liu, Min Xu: ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences https://arxiv.org/abs/2609.30043 https://arxiv.org/pdf/2609.30043 https://arxiv.org/html/2609.30043
September 25, 2026 at 6:41 AM
Bazi, Aljuhani, Al Rahhal, Zuair, Alajlan: A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification https://arxiv.org/abs/2609.29376 https://arxiv.org/pdf/2609.29376 https://arxiv.org/html/2609.29376
September 25, 2026 at 6:40 AM