Damien Teney
banner
damienteney.bsky.social
Damien Teney
@damienteney.bsky.social
Research Scientist @ Idiap Research Institute. @idiap.bsky.social
Adjunct lecturer @ Australian Institute for ML. @aimlofficial.bsky.social
Occasionally cycling across continents.
https://www.damienteney.info
Reviewer #2 striking again?
February 21, 2026 at 1:15 PM
This all looks very promising, and there's a lot more to explore! Paper and code ⬇️
Procedural Pretraining: Warming Up Language Models with Abstract Data
www.arxiv.org/abs/2601.21725
github.com/zlshinnick/p...
Procedural Pretraining: Warming Up Language Models with Abstract Data
Pretraining directly on web-scale corpora is the de facto paradigm for building language models. We study an alternative setting where the model is initially exposed to abstract structured data, as a ...
www.arxiv.org
February 20, 2026 at 12:39 PM
🧩 Multiple types of procedural data can be combined.
We get further gains by mixing either
• multiple types of data, or
• weights of models individually warmed-up on different types of data.
February 20, 2026 at 12:39 PM
⚙️ MLPs vs. attention: where is the information located?
We try resetting selected weights to random, before standard pretraining. Surprisingly, we obtain further gains, but they're domain-specific:
• warmed-up MLPs benefit natural language
• warmed-up attention helps code/math
February 20, 2026 at 12:39 PM
📈 Benefits on subsequent standard pretraining.
By front-loading as little as 0.1% procedural data, models achieve significantly better pretraining performance on language, code, and math. They use up to 45% less semantic data to reach a baseline perplexity.
February 20, 2026 at 12:39 PM
🔍 Different procedural data = different benefits.
We first determine the effect of different types of procedural data with algorithmic diagnostic tasks. The benefits range from long-context recall to arithmetic, depending on the type of procedural data.
February 20, 2026 at 12:39 PM
💡 Humans learn better when starting with simple structure and logic rather than memorizing a massive set of facts. By analogy, we use abstract, structured data to build a scaffold in language models, free of semantic biases.
February 20, 2026 at 12:39 PM
Indeed the effect in *late* layers was very surprising!
My optimist interpretation is that the procedural pretraining creates circuits for computations general enough to serve as a useful scaffold for visual tasks. This would explain why they help and why they don't wash out with more training.
December 10, 2025 at 8:29 PM
Sounds 😋 What's the objective function? simplicity/low cost/?
December 10, 2025 at 8:22 PM
In summary, a lightweight generic warm-up improves accuracy and data efficiency, with effects distinct from ImageNet pretraining.
Lots of exciting open questions! 🔍
-Other types of procedural data?
-Other downstream tasks?
-Closed-form instantiation?
arxiv.org/abs/2511.13945
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
Transformers show remarkable versatility across domains, suggesting the existence of inductive biases beneficial across modalities. In this work, we explore a new way to instil such generic biases in ...
arxiv.org
December 10, 2025 at 1:23 PM
🔍𝐖𝐡𝐞𝐫𝐞 𝐢𝐬 𝐭𝐡𝐢𝐬 𝐤𝐧𝐨𝐰𝐥𝐞𝐝𝐠𝐞 𝐬𝐭𝐨𝐫𝐞𝐝?
Ablations show that the knowledge mostly locates in *late* layers: the opposite of normal visual pretraining which shapes early layers. Procedural data seems to provide a qualitatively unique training signal!
December 10, 2025 at 1:23 PM
🧠𝐖𝐡𝐚𝐭 𝐤𝐢𝐧𝐝 𝐨𝐟 𝐝𝐚𝐭𝐚 𝐰𝐨𝐫𝐤𝐬?
Formal languages with hierarchical structure seem best. If we shuffle the training tokens (eliminating nested structures), the gains disappear, showing that the benefits are not due to surface-level frequencies.
December 10, 2025 at 1:23 PM
📉𝐏𝐫𝐨𝐜𝐞𝐝𝐮𝐫𝐚𝐥 𝐝𝐚𝐭𝐚 𝐜𝐚𝐧 𝐫𝐞𝐩𝐥𝐚𝐜𝐞 𝐫𝐞𝐚𝐥 𝐢𝐦𝐚𝐠𝐞𝐬
Allocating just 1% of the ImageNet pretraining budget to the procedural warmup lets the ViT match the baseline accuracy with 28% fewer images!
December 10, 2025 at 1:23 PM
📈𝐀 𝐝𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐨𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐭𝐫𝐚𝐣𝐞𝐜𝐭𝐨𝐫𝐲
Our warmed-up models don't just get a head-start, they train differently. On ImageNet (below), they follow a distinct training trajectory and converge to a better accuracy.
December 10, 2025 at 1:23 PM
🔥Our procedural data has no semantic or visual meaning: it simply forces the model to discover generic structure in the data. As initialisation for standard image-based training, it
-boosts accuracy,
-improves data efficiency,
-complements ImageNet pretraining.
December 10, 2025 at 1:23 PM
💡Prior work has already shown that LLMs acquire useful knowledge when pretrained on formal languages. To test this on ViTs, we devise a procedural warm-up: pretraining for next-token prediction on symbolic sequences, bypassing the visual patch embedding.
December 10, 2025 at 1:23 PM
In summary, a lightweight generic warm-up improves accuracy and data efficiency, with effects distinct from ImageNet pretraining.
Lots of exciting open questions! 🔍
- Other types of procedural data?
- Other downstream tasks?
- Closed-form instantiation?
arxiv.org/abs/2511.13945
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
Transformers show remarkable versatility across domains, suggesting the existence of inductive biases beneficial across modalities. In this work, we explore a new way to instil such generic biases in ...
arxiv.org
December 10, 2025 at 11:08 AM
🔍𝐖𝐡𝐞𝐫𝐞 𝐢𝐬 𝐭𝐡𝐢𝐬 𝐤𝐧𝐨𝐰𝐥𝐞𝐝𝐠𝐞 𝐬𝐭𝐨𝐫𝐞𝐝?
Ablations show that the knowledge mostly locates in *late* layers: the opposite of normal visual pretraining which shapes early layers. Procedural data seems to provide a qualitatively unique training signal!
December 10, 2025 at 11:06 AM
🧠𝐖𝐡𝐚𝐭 𝐤𝐢𝐧𝐝 𝐨𝐟 𝐝𝐚𝐭𝐚 𝐰𝐨𝐫𝐤𝐬?
Formal languages with hierarchical structure seem best. If we shuffle the training tokens (eliminating nested structures), the gains disappear, showing that the benefits are not due to surface-level frequencies.
December 10, 2025 at 11:06 AM
📉𝐏𝐫𝐨𝐜𝐞𝐝𝐮𝐫𝐚𝐥 𝐝𝐚𝐭𝐚 𝐜𝐚𝐧 𝐫𝐞𝐩𝐥𝐚𝐜𝐞 𝐫𝐞𝐚𝐥 𝐢𝐦𝐚𝐠𝐞𝐬
Allocating just 1% of the ImageNet pretraining budget to the procedural warmup lets the ViT match the baseline accuracy with 28% fewer images!
December 10, 2025 at 11:06 AM
📈𝐀 𝐝𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐨𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐭𝐫𝐚𝐣𝐞𝐜𝐭𝐨𝐫𝐲
Our warmed-up models don't just get a head-start, they train differently. On ImageNet (below), they follow a distinct training trajectory and converge to a better accuracy.
December 10, 2025 at 11:06 AM
🔥Our procedural data has no semantic or visual meaning: it simply forces the model to discover generic structure in the data. As initialisation for standard image-based training, it
-boosts accuracy,
-improves data efficiency,
-complements ImageNet pretraining.
December 10, 2025 at 11:06 AM
Academic Strava?🤓 It feels like an underrepresented group in my Strava feed!
July 26, 2025 at 10:10 AM