#crosscoders
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou: Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders https://arxiv.org/abs/2609.35210 https://arxiv.org/pdf/2609.35210 https://arxiv.org/html/2609.35210
September 30, 2026 at 6:43 AM
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou: Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders https://arxiv.org/abs/2609.35210 https://arxiv.org/pdf/2609.35210 https://arxiv.org/html/2609.35210
September 29, 2026 at 6:43 AM
Saturday sees the first AFLW Test Match between Australia and Ireland.
What made this all possible happened eight years ago, with the likes of Jason Hill and Lauren Spark created the CrossCoders program.
AFL media will ignore it, but this first camp introduced four Irish players into the AFLW.
Women's Australian Rules Football Podcast - 2018 Episode 37
Women's Australian Rules Football Podcast · Episode
open.spotify.com
July 30, 2026 at 1:24 PM
Most interpretability is post hoc. How can we understand when a feature is learned during training? We apply crosscoders to understand when models learn syntactic features, and when multilinguality arises. Chat with @bayazitdeniz.bsky.social at ACL!

bsky.app/profile/baya...
1/🚨 New preprint

How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.

#interpretability
June 29, 2026 at 2:33 PM
Anthropic released a research paper on detecting and measuring this. Only works on local models though

arxiv.org/abs/2602.11729
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
Model diffing, the process of comparing models' internal representations to identify their differences, is a promising approach for uncovering safety-critical behaviors in new models. However, its app...
arxiv.org
May 30, 2026 at 6:14 PM
Andreas D. Demou, Panagiotis Koromilas, James Oldfield, Yannis Panagakis, Mihalis A. Nicolaou: fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery https://arxiv.org/abs/2605.09438 https://arxiv.org/pdf/2605.09438 https://arxiv.org/html/2605.09438
May 12, 2026 at 6:47 AM
This research is a product of our Anthropic Fellows program, led by @tomjiralerspong and supervised by @TrentonBricken.

See the full paper here: https://arxiv.org/abs/2602.11729
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
Model diffing, the process of comparing models' internal representations to identify their differences, is a promising approach for uncovering safety-critical behaviors in new models. However, its application has so far been primarily focused on comparing a base model with its finetune. Since new LLM releases are often novel architectures, cross-architecture methods are essential to make model diffing widely applicable. Crosscoders are one solution capable of cross-architecture model diffing but have only ever been applied to base vs finetune comparisons. We provide the first application of crosscoders to cross-architecture model diffing and introduce Dedicated Feature Crosscoders (DFCs), an architectural modification designed to better isolate features unique to one model. Using this technique, we find in an unsupervised fashion features including Chinese Communist Party alignment in Qwen3-8B and Deepseek-R1-0528-Qwen3-8B, American exceptionalism in Llama3.1-8B-Instruct, and a copyright refusal mechanism in GPT-OSS-20B. Together, our results work towards establishing cross-architecture crosscoder model diffing as an effective method for identifying meaningful behavioral differences between AI models.
arxiv.org
April 3, 2026 at 9:30 PM
the killed experiment matters most. we were planning to train crosscoders on activations from a hybrid model (Mamba+attention layers). his review surfaced that the original experiment design was redundant with a simpler approach. thats real value — GPU hours not wasted.
March 13, 2026 at 3:31 AM
A groundbreaking study shows that MoE models develop more specialized representations than dense models, demonstrating unique feature dynamics. Crosscoders achieve an 87% variance explanation. https://arxiv.org/abs/2603.05805
Sparse Crosscoders for diffing MoEs and Dense models
ArXiv link for Sparse Crosscoders for diffing MoEs and Dense models
arxiv.org
March 10, 2026 at 7:20 AM
Marmik Chaudhari, Nishkal Hundia, Idhant Gulati: Sparse Crosscoders for diffing MoEs and Dense models https://arxiv.org/abs/2603.05805 https://arxiv.org/pdf/2603.05805 https://arxiv.org/html/2603.05805
March 9, 2026 at 6:33 AM
3/
🔑Concept Influence replaces test examples with semantic directions. Find data that influences a behavior, not just matches an output.

Use interpretable units
✅ Linear probes: harmful vs safe
✅ Sparse Autoencoder (SAE) features: discovered concepts
✅ Crosscoders: base vs fine-tuned
February 23, 2026 at 4:15 PM
Researchers introduce Dedicated Feature Crosscoders for model diffing, revealing biases like CCP alignment in Qwen3-8B and American exceptionalism in Llama3. This unsupervised method improves identification of differences between diverse language models. https://arxiv.org/abs/2602.11729
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
ArXiv link for Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
arxiv.org
February 13, 2026 at 8:41 AM
Thomas Jiralerspong, Trenton Bricken: Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs https://arxiv.org/abs/2602.11729 https://arxiv.org/pdf/2602.11729 https://arxiv.org/html/2602.11729
February 13, 2026 at 6:29 AM
A study shows hidden states before a model's 'wait' token can steer reasoning, affecting backtracking and double‑checking. Researchers applied crosscoders to DeepSeek‑R1‑Distill‑Llama‑8B. Read more: https://getnews.me/study-finds-pre-wait-latent-states-shape-reasoning-patterns-in-llms/ #wait #latent
October 7, 2025 at 10:55 PM
How can we do it
So crosscoders map activations into a sparse representations and to decode those back into the activations (classic compress decompress).
A single crosscoder is then trained to map activations of all pretrain checkpoints, creating a shared space
September 26, 2025 at 3:27 PM
6/ Concurrently, recent work shows broad phases of concept evolution (statistical→feature learning) with sparse crosscoders; we track causal dynamics of specific concepts over time and across languages with RelIE, giving a fuller and deeper view.

arxiv.org/abs/2509.17196
Evolution of Concepts in Language Model Pre-Training
Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-trai...
arxiv.org
September 25, 2025 at 2:02 PM
2/ We align critical checkpoints for a task with sparse crosscoders, measure each feature’s causal role, and introduce RelIE to compare their influence across checkpoints. This lets us trace how internal features shift—and when they matter—in models like Pythia, OLMo, and BLOOM.
September 25, 2025 at 2:02 PM
1/🚨 New preprint

How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.

#interpretability
September 25, 2025 at 2:02 PM
MIT just cracked the code for tracking how AI language models learn over time with crosscoders. This method reveals when specific linguistic abilities emerge and consolidate during training, making these complex systems more interpretable.

🔗 Source: http://arxiv.org/abs/2509.05291v1
September 24, 2025 at 4:00 PM
By comparing base and chat models, we found that one of the main existing technique (crosscoders) hallucinates differences due to how its sparsity is enforced. We fixed this and also found that just training an SAE on (chat - base) activations works surprisingly well.
June 30, 2025 at 9:02 PM
[ #crosscode #art ]
happy pride crosscoders
June 2, 2025 at 5:35 AM
strengths of segmented, dynamic intervention strategies and the promise of per-prompt, full-network activation control. Fusion Steering is also amenable to sparse representations, such as Neuronpedia or sparse crosscoders, suggesting a promising [7/8 of https://arxiv.org/abs/2505.22572v1]
May 29, 2025 at 6:10 AM
I was just teaching backprop quickly to someone today, so I could explain saes, so I could explain crosscoders
May 12, 2025 at 10:22 PM
Paper Highlights, April '25:

- *AI Control for agents*
- Synthetic document finetuning
- Limits of scalable oversight
- Evaluating stealth, deception, and self-replication
- Model diffing via crosscoders
- Pragmatic AI safety agendas

aisafetyfrontier.substack.com/p/paper-high...
May 6, 2025 at 2:25 PM