you can now play with natural language autoencoders right in your browser
you can now play with natural language autoencoders right in your browser
www.biorxiv.org/content/10.1...
www.biorxiv.org/content/10.1...
www.nature.com/articles/s41... by @lpachter.bsky.social and Briefing www.nature.com/articles/s41...
www.nature.com/articles/s41... by @lpachter.bsky.social and Briefing www.nature.com/articles/s41...
www.biorxiv.org/content/10.1...
- Use sparse autoencoders (SAEs) to extract and analyze interpretable features from ESM-2-8M
www.biorxiv.org/content/10.1...
- Use sparse autoencoders (SAEs) to extract and analyze interpretable features from ESM-2-8M
In short: we emphasize how autoencoders are implemented—but not always what they represent (and some of the implications of that representation).🧵
In short: we emphasize how autoencoders are implemented—but not always what they represent (and some of the implications of that representation).🧵
arxiv.org/html/2501.17...
arxiv.org/html/2501.17...
Not moral: LLMs, SVM
Somewhat moral: EBMs, linear regression, NNMF, GPs
Moral: Autoencoders
Not moral: LLMs, SVM
Somewhat moral: EBMs, linear regression, NNMF, GPs
Moral: Autoencoders
arxiv.org/abs/2502.03714
(1/9)
arxiv.org/abs/2502.03714
(1/9)
They simplify tuning with k-sparse autoencoders and results show many improvements in explainability. Code, models (not all!) and visualizer included. openreview.net/forum?id=tcs...
They simplify tuning with k-sparse autoencoders and results show many improvements in explainability. Code, models (not all!) and visualizer included. openreview.net/forum?id=tcs...
anthropic.com/research/natural-language-autoencoders
anthropic.com/research/natural-language-autoencoders
In new work led by @liambai.bsky.social and me, we explore how sparse autoencoders can help us understand biology—going from mechanistic interpretability to mechanistic biology.
In new work led by @liambai.bsky.social and me, we explore how sparse autoencoders can help us understand biology—going from mechanistic interpretability to mechanistic biology.
Yes.
Yes.
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
Code: github.com/ElanaPearl/I...
Interactive site: interplm.ai
Nice work by Elana Simon from James Zou lab
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
Code: github.com/ElanaPearl/I...
Interactive site: interplm.ai
Nice work by Elana Simon from James Zou lab
Also, a good reminder to share that our work is now out in Cell Reports 🙏🎊
⬇️
www.cell.com/cell-reports...
Also, a good reminder to share that our work is now out in Cell Reports 🙏🎊
⬇️
www.cell.com/cell-reports...
Blog: https://www.anthropic.com/research/natural-language-autoencoders
Whitepaper : https://transformer-circuits.pub/2026/nla/index.html
Blog: https://www.anthropic.com/research/natural-language-autoencoders
Whitepaper : https://transformer-circuits.pub/2026/nla/index.html