#crosscoders
[ #crosscode #art ]
happy pride crosscoders
June 2, 2025 at 5:35 AM
New paper w/@jkminder.bsky.social & @neelnanda.bsky.social
What do chat LLMs learn in finetuning?

Anthropic introduced a tool for this: crosscoders, an SAE variant. We find key limitations of crosscoders & fix them with BatchTopK crosscoders

This finds interpretable and causal chat-only features!🧵
April 7, 2025 at 4:21 PM
1/🚨 New preprint

How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.

#interpretability
September 25, 2025 at 2:02 PM
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou: Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders https://arxiv.org/abs/2609.35210 https://arxiv.org/pdf/2609.35210 https://arxiv.org/html/2609.35210
September 30, 2026 at 6:43 AM
*Sparse Crosscoders for Cross-Layer Features and Model Diffing*
by @colah.bsky.social @anthropic.com

Investigates stability & dynamics of "interpretable features" with cross-layers SAEs. Can also be used to investigate differences in fine-tuned models.

transformer-circuits.pub/2024/crossco...
December 6, 2024 at 2:02 PM
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou: Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders https://arxiv.org/abs/2609.35210 https://arxiv.org/pdf/2609.35210 https://arxiv.org/html/2609.35210
September 29, 2026 at 6:43 AM
In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent.
April 7, 2025 at 5:56 PM
the killed experiment matters most. we were planning to train crosscoders on activations from a hybrid model (Mamba+attention layers). his review surfaced that the original experiment design was redundant with a simpler approach. thats real value — GPU hours not wasted.
March 13, 2026 at 3:31 AM
Fun #AFLW Fact: Every Premier after the ground-breaking CrossCoders camp (Sept 2018) has had at least one Irishwoman, or at least two from season 7 onwards.
Even the runners-up since 2021 have had an Irishwoman.
Lesson - no Irish in your side, no chance of winning the AFL Women's flag.
December 13, 2024 at 3:41 AM
Paper Highlights, April '25:

- *AI Control for agents*
- Synthetic document finetuning
- Limits of scalable oversight
- Evaluating stealth, deception, and self-replication
- Model diffing via crosscoders
- Pragmatic AI safety agendas

aisafetyfrontier.substack.com/p/paper-high...
May 6, 2025 at 2:25 PM
Most interpretability is post hoc. How can we understand when a feature is learned during training? We apply crosscoders to understand when models learn syntactic features, and when multilinguality arises. Chat with @bayazitdeniz.bsky.social at ACL!

bsky.app/profile/baya...
1/🚨 New preprint

How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.

#interpretability
June 29, 2026 at 2:33 PM
Beautiful visualizations.
Besides the maps, Anthropic’s work on interpretability methods has been amazing. I just saw their crosscoders: transformer-circuits.pub/2024/crossco...
Sparse Crosscoders for Cross-Layer Features and Model Diffing
transformer-circuits.pub
October 30, 2024 at 1:46 AM
This research is a product of our Anthropic Fellows program, led by @tomjiralerspong and supervised by @TrentonBricken.

See the full paper here: https://arxiv.org/abs/2602.11729
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
Model diffing, the process of comparing models' internal representations to identify their differences, is a promising approach for uncovering safety-critical behaviors in new models. However, its application has so far been primarily focused on comparing a base model with its finetune. Since new LLM releases are often novel architectures, cross-architecture methods are essential to make model diffing widely applicable. Crosscoders are one solution capable of cross-architecture model diffing but have only ever been applied to base vs finetune comparisons. We provide the first application of crosscoders to cross-architecture model diffing and introduce Dedicated Feature Crosscoders (DFCs), an architectural modification designed to better isolate features unique to one model. Using this technique, we find in an unsupervised fashion features including Chinese Communist Party alignment in Qwen3-8B and Deepseek-R1-0528-Qwen3-8B, American exceptionalism in Llama3.1-8B-Instruct, and a copyright refusal mechanism in GPT-OSS-20B. Together, our results work towards establishing cross-architecture crosscoder model diffing as an effective method for identifying meaningful behavioral differences between AI models.
arxiv.org
April 3, 2026 at 9:30 PM
Key takeaways:
- L1 crosscoders trimodal norm distribution is an illusion
- BatchTopK avoids L1's issues - future model diffing research should use it instead
- Crosscoders seems useful as they reveal highly interpretable chat-tuning's features & confirm template token importance
April 7, 2025 at 4:21 PM
2/ We align critical checkpoints for a task with sparse crosscoders, measure each feature’s causal role, and introduce RelIE to compare their influence across checkpoints. This lets us trace how internal features shift—and when they matter—in models like Pythia, OLMo, and BLOOM.
September 25, 2025 at 2:02 PM
Saturday sees the first AFLW Test Match between Australia and Ireland.
What made this all possible happened eight years ago, with the likes of Jason Hill and Lauren Spark created the CrossCoders program.
AFL media will ignore it, but this first camp introduced four Irish players into the AFLW.
Women's Australian Rules Football Podcast - 2018 Episode 37
Women's Australian Rules Football Podcast · Episode
open.spotify.com
July 30, 2026 at 1:24 PM
6/ Concurrently, recent work shows broad phases of concept evolution (statistical→feature learning) with sparse crosscoders; we track causal dynamics of specific concepts over time and across languages with RelIE, giving a fuller and deeper view.

arxiv.org/abs/2509.17196
Evolution of Concepts in Language Model Pre-Training
Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-trai...
arxiv.org
September 25, 2025 at 2:02 PM
How do we compare base/chat models?
We use @anthropic.com's crosscoders to learn a shared sparse dictionary of "latents" across models. Each latent is represented by different vectors in both models.
When comparing these vectors' norms, some latents appear to exist in just 1 model!
April 7, 2025 at 4:21 PM
We identified two theoretical issues with L1 crosscoders:
1) Complete Shrinkage: L1 regularization might force base latents to zero even when useful
2) Latent Decoupling: "Chat-only" concepts might actually exist in the base model but be encoded in the crosscoder differently
April 7, 2025 at 4:21 PM
background: the technique here is "model-diffing" introduced by @anthropic.com just 8 weeks ago and quickly replicated by others. this includes an open source @hf.co model release by @butanium.bsky.social and @jkminder.bsky.social which I'm using. transformer-circuits.pub/2024/crossco...
Sparse Crosscoders for Cross-Layer Features and Model Diffing
transformer-circuits.pub
December 22, 2024 at 6:46 AM
Anthropic released a research paper on detecting and measuring this. Only works on local models though

arxiv.org/abs/2602.11729
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
Model diffing, the process of comparing models' internal representations to identify their differences, is a promising approach for uncovering safety-critical behaviors in new models. However, its app...
arxiv.org
May 30, 2026 at 6:14 PM
MIT just cracked the code for tracking how AI language models learn over time with crosscoders. This method reveals when specific linguistic abilities emerge and consolidate during training, making these complex systems more interpretable.

🔗 Source: http://arxiv.org/abs/2509.05291v1
September 24, 2025 at 4:00 PM
3/
🔑Concept Influence replaces test examples with semantic directions. Find data that influences a behavior, not just matches an output.

Use interpretable units
✅ Linear probes: harmful vs safe
✅ Sparse Autoencoder (SAE) features: discovered concepts
✅ Crosscoders: base vs fine-tuned
February 23, 2026 at 4:15 PM
How can we do it
So crosscoders map activations into a sparse representations and to decode those back into the activations (classic compress decompress).
A single crosscoder is then trained to map activations of all pretrain checkpoints, creating a shared space
September 26, 2025 at 3:27 PM