What made this all possible happened eight years ago, with the likes of Jason Hill and Lauren Spark created the CrossCoders program.
AFL media will ignore it, but this first camp introduced four Irish players into the AFLW.
What made this all possible happened eight years ago, with the likes of Jason Hill and Lauren Spark created the CrossCoders program.
AFL media will ignore it, but this first camp introduced four Irish players into the AFLW.
bsky.app/profile/baya...
How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.
#interpretability
bsky.app/profile/baya...
arxiv.org/abs/2602.11729
arxiv.org/abs/2602.11729
See the full paper here: https://arxiv.org/abs/2602.11729
See the full paper here: https://arxiv.org/abs/2602.11729
🔑Concept Influence replaces test examples with semantic directions. Find data that influences a behavior, not just matches an output.
Use interpretable units
✅ Linear probes: harmful vs safe
✅ Sparse Autoencoder (SAE) features: discovered concepts
✅ Crosscoders: base vs fine-tuned
🔑Concept Influence replaces test examples with semantic directions. Find data that influences a behavior, not just matches an output.
Use interpretable units
✅ Linear probes: harmful vs safe
✅ Sparse Autoencoder (SAE) features: discovered concepts
✅ Crosscoders: base vs fine-tuned
So crosscoders map activations into a sparse representations and to decode those back into the activations (classic compress decompress).
A single crosscoder is then trained to map activations of all pretrain checkpoints, creating a shared space
So crosscoders map activations into a sparse representations and to decode those back into the activations (classic compress decompress).
A single crosscoder is then trained to map activations of all pretrain checkpoints, creating a shared space
arxiv.org/abs/2509.17196
arxiv.org/abs/2509.17196
How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.
#interpretability
How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks.
#interpretability
🔗 Source: http://arxiv.org/abs/2509.05291v1
🔗 Source: http://arxiv.org/abs/2509.05291v1
- *AI Control for agents*
- Synthetic document finetuning
- Limits of scalable oversight
- Evaluating stealth, deception, and self-replication
- Model diffing via crosscoders
- Pragmatic AI safety agendas
aisafetyfrontier.substack.com/p/paper-high...
- *AI Control for agents*
- Synthetic document finetuning
- Limits of scalable oversight
- Evaluating stealth, deception, and self-replication
- Model diffing via crosscoders
- Pragmatic AI safety agendas
aisafetyfrontier.substack.com/p/paper-high...