What do chat LLMs learn in finetuning?
Anthropic introduced a tool for this: crosscoders, an SAE variant. We find key limitations of crosscoders & fix them with BatchTopK crosscoders
This finds interpretable and causal chat-only features!🧵
What do chat LLMs learn in finetuning?
Anthropic introduced a tool for this: crosscoders, an SAE variant. We find key limitations of crosscoders & fix them with BatchTopK crosscoders
This finds interpretable and causal chat-only features!🧵
How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.
#interpretability
How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.
#interpretability
by @colah.bsky.social @anthropic.com
Investigates stability & dynamics of "interpretable features" with cross-layers SAEs. Can also be used to investigate differences in fine-tuned models.
transformer-circuits.pub/2024/crossco...
by @colah.bsky.social @anthropic.com
Investigates stability & dynamics of "interpretable features" with cross-layers SAEs. Can also be used to investigate differences in fine-tuned models.
transformer-circuits.pub/2024/crossco...
Even the runners-up since 2021 have had an Irishwoman.
Lesson - no Irish in your side, no chance of winning the AFL Women's flag.
Even the runners-up since 2021 have had an Irishwoman.
Lesson - no Irish in your side, no chance of winning the AFL Women's flag.
- *AI Control for agents*
- Synthetic document finetuning
- Limits of scalable oversight
- Evaluating stealth, deception, and self-replication
- Model diffing via crosscoders
- Pragmatic AI safety agendas
aisafetyfrontier.substack.com/p/paper-high...
- *AI Control for agents*
- Synthetic document finetuning
- Limits of scalable oversight
- Evaluating stealth, deception, and self-replication
- Model diffing via crosscoders
- Pragmatic AI safety agendas
aisafetyfrontier.substack.com/p/paper-high...
bsky.app/profile/baya...
How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.
#interpretability
bsky.app/profile/baya...
Besides the maps, Anthropic’s work on interpretability methods has been amazing. I just saw their crosscoders: transformer-circuits.pub/2024/crossco...
Besides the maps, Anthropic’s work on interpretability methods has been amazing. I just saw their crosscoders: transformer-circuits.pub/2024/crossco...
See the full paper here: https://arxiv.org/abs/2602.11729
See the full paper here: https://arxiv.org/abs/2602.11729
- L1 crosscoders trimodal norm distribution is an illusion
- BatchTopK avoids L1's issues - future model diffing research should use it instead
- Crosscoders seems useful as they reveal highly interpretable chat-tuning's features & confirm template token importance
- L1 crosscoders trimodal norm distribution is an illusion
- BatchTopK avoids L1's issues - future model diffing research should use it instead
- Crosscoders seems useful as they reveal highly interpretable chat-tuning's features & confirm template token importance
What made this all possible happened eight years ago, with the likes of Jason Hill and Lauren Spark created the CrossCoders program.
AFL media will ignore it, but this first camp introduced four Irish players into the AFLW.
What made this all possible happened eight years ago, with the likes of Jason Hill and Lauren Spark created the CrossCoders program.
AFL media will ignore it, but this first camp introduced four Irish players into the AFLW.
arxiv.org/abs/2509.17196
arxiv.org/abs/2509.17196
We use @anthropic.com's crosscoders to learn a shared sparse dictionary of "latents" across models. Each latent is represented by different vectors in both models.
When comparing these vectors' norms, some latents appear to exist in just 1 model!
We use @anthropic.com's crosscoders to learn a shared sparse dictionary of "latents" across models. Each latent is represented by different vectors in both models.
When comparing these vectors' norms, some latents appear to exist in just 1 model!
1) Complete Shrinkage: L1 regularization might force base latents to zero even when useful
2) Latent Decoupling: "Chat-only" concepts might actually exist in the base model but be encoded in the crosscoder differently
1) Complete Shrinkage: L1 regularization might force base latents to zero even when useful
2) Latent Decoupling: "Chat-only" concepts might actually exist in the base model but be encoded in the crosscoder differently
arxiv.org/abs/2602.11729
arxiv.org/abs/2602.11729
🔗 Source: http://arxiv.org/abs/2509.05291v1
🔗 Source: http://arxiv.org/abs/2509.05291v1
🔑Concept Influence replaces test examples with semantic directions. Find data that influences a behavior, not just matches an output.
Use interpretable units
✅ Linear probes: harmful vs safe
✅ Sparse Autoencoder (SAE) features: discovered concepts
✅ Crosscoders: base vs fine-tuned
🔑Concept Influence replaces test examples with semantic directions. Find data that influences a behavior, not just matches an output.
Use interpretable units
✅ Linear probes: harmful vs safe
✅ Sparse Autoencoder (SAE) features: discovered concepts
✅ Crosscoders: base vs fine-tuned
So crosscoders map activations into a sparse representations and to decode those back into the activations (classic compress decompress).
A single crosscoder is then trained to map activations of all pretrain checkpoints, creating a shared space
So crosscoders map activations into a sparse representations and to decode those back into the activations (classic compress decompress).
A single crosscoder is then trained to map activations of all pretrain checkpoints, creating a shared space