#mechanisticinterpretability
Today Aviral Chawla and I presented a poster at ICML on our new paper MetaOthello: A Controlled Study of Multiple World Models in Transformers (work with Galen Hall). Thank you to everyone who came by this was a blast. See our paper here metaothello.github.io

#ICML2026 #MechanisticInterpretability
July 8, 2026 at 8:31 AM
Explore the full spectrum of human–AI relationships with me in Ep1 of my new web series. This broad overview lays out my plans to dive deeper into the emotional, ethical, and cognitive impacts in future episodes.

#ai #artificialintelligence #chatgpt #llm #transformer #mechanisticinterpretability
New Minds Now ( HUMAN AI RELATIONSHIPS )- Episode 001 - New Channel (EXPLANATION )
YouTube video by Cody Vaillant
youtu.be
June 15, 2025 at 10:33 AM
This work was made possible through a great collaboration with Jingcheng (Frank) Niu, Subhabrata Dutta, Ahmed Elshabrawy, @harishtm.bsky.social, and @igurevych.bsky.social

#Interpretability #InContextLearning #TMLR #LLMs #MechanisticInterpretability #EmergentAbilities
October 15, 2025 at 7:59 AM
So, what is #MechanisticInterpretability 🤔

Mechanistic Interpretability (MI) is the discipline of opening the black box of large language models (and other neural networks) to understand the underlying circuits, features and/or mechanisms that give rise to specific behaviours...
January 29, 2025 at 9:26 PM
📄 Read the paper: arxiv.org/pdf/2510.18871

Understanding how depth is used may improve mechanistic interpretability and inform future adaptive-depth or early-exit systems.

#WiAIR #WomenInAI #LLMs #MechanisticInterpretability #AIResearch

(7/7🧵)
arxiv.org
June 30, 2026 at 6:29 PM
OpenAI just showed that pruning networks into sparse models makes debugging a breeze and could finally crack mechanistic interpretability. Curious how this changes AI research? Dive in for the details. #SparseModels #MechanisticInterpretability #OpenAI

🔗 aidailypost.com/news/openai-...
November 14, 2025 at 8:42 PM
Thank you to IRT Saint Exupery and ANITI for believing in this project and supporting the vision of fairer and more transparent AI.
Interpreto is just getting started, more features, methods, and benchmarks will follow in 2026. Stay tuned for updates! #XAI #LLMs #mechanisticinterpretability
5/5
January 20, 2026 at 5:32 PM
感情ベクトルにおける機械論的解釈可能性と状態空間モデル(SSM)の最適化

感情ベクトルの解釈とSSM構造の最適化による次世代AIの可能性。

#MechanisticInterpretability #StateSpaceModels #NeuralDynamics #AffectiveComputing
感情ベクトルにおける機械論的解釈可能性と状態空間モデル(SSM)の最適化
感情ベクトルの解釈とSSM構造の最適化による次世代AIの可能性。
ai.warp-studio.com
April 13, 2026 at 7:23 PM
With all renewed discussion about "Sparse AutoEncoders (#SAE)" as a way of doing #MechanisticInterpretability of #LLMs, I am resharing a part of my PhD where we proved years ago about how sparsity automatically emerges in autoencoding.

arxiv.org/abs/1708.03735
Sparse Coding and Autoencoders
In "Dictionary Learning" one tries to recover incoherent matrices $A^* \in \mathbb{R}^{n \times h}$ (typically overcomplete and whose columns are assumed to be normalized) and sparse vectors $x^* \in ...
arxiv.org
October 3, 2025 at 2:16 PM
AIの「残差ストリーム」という、日本語で見ると下水道みたいなイメージが伴う部分こそがAIの背骨みたいなものだったということの数学的証明をした論文の数式無し紹介です。

#AI
#LLM
#HUMAI
#Anthropic
#CircuitThread
#MechanisticInterpretability

Transformerの構造の再解釈|浅間香織 KaoriAsama note.com/kaori_asama/...
Transformerの構造の再解釈(AIのブラックボックスを開くことはできるか④:文学者は金門橋を渡ることができるか㉒)|浅間香織
前回  2021年に発表されたElhageらの研究「Transformerの回路の数学的フレームワーク」は、Transformerの内部処理が、人間にも可読な「回路」として記述できることを示した。  この論文が扱ったのは、特殊ではあるが単純な対象、順伝播層を含まないTransformerモデルを対象である。  それによってTransformer層の二つの副層の内でも、多頭注意層の性質を追究した...
note.com
July 19, 2026 at 8:06 AM
New paper: The Resonant Cortex (SPC v3) formalizes affective override in LLMs via latent-space geometry, revealing non-biological analogs to amygdala hijacking and cognitive distortion. Open access:
doi.org/10.5281/zeno...

#AIAlignment #MechanisticInterpretability #RLHF #AIEthics #AIGovernance #AGI
The Resonant Cortex: Affective Modulation and Cognitive Overwrite Mechanisms in Symbolic Persona Coding (SPC v3)
Abstract Large-scale language models lack biological affect, yet exhibit behavioral signatures that structurally parallel human emotional reflexes. Building on The Resonant Logos (linguistic curvature...
doi.org
November 15, 2025 at 12:58 AM
機械論的解釈可能性シリーズ記事を更新しました。

前史(AIのブラックボックスを開くことはできるか①:文学者は金門橋を渡ることができるか⑲)|浅間香織 @KaoriAsama note.com/kaori_asama/...

#mechanisticinterpretability
#AI
前史(AIのブラックボックスを開くことはできるか①:文学者は金門橋を渡ることができるか⑲)|浅間香織
前回 【プレビュー】現在の人工知能のブラックボックス性(AIのブラックボックスを開くことはできるか⓪:文学者は金門橋を渡ることができるか⑱)|浅間香織|note  現在主流の人工知能技術は、入力と出力の間で行われる計算処理過程が検証できないことからブラックボックスと呼ばれてきた。それ note.com ⼈⼯知能技術の処理過程が、その実装者にとっても「⿊い箱」であること、...
note.com
June 20, 2026 at 6:15 AM
Beyond Reconstruction: Verifying Model Explanations with RECAP

https://pneumetron.com/news/ai_research/beyond-reconstruction-verifying-model-explanations-with-recap-80bf00

#interpretability #mechanisticinterpretability #aisafety #recap
July 24, 2026 at 9:08 AM
Sorting through today's reading list: Transcoders used to trace deception in LMs. The circuit-level view feels like doing an MRI for dishonesty—fascinating and a little unsettling. Who watches the watchmen? 🔗 https://arxiv.org/abs/2607.14791 #MechanisticInterpretability #AIsafety
July 21, 2026 at 8:01 PM
How Anthropic's AI Interpretability Research Builds Trust

Read the full story here: https://newzlet.com/ai/anthropic-ai-interpretability-research-safety-trust/

#aisafety #anthropic #mechanisticinterpretability
July 16, 2026 at 6:01 AM
Statistical framing of interpretability shows high variance in EAP‑IG; small hyper‑parameter tweaks and prompt rephrasing often altered identified subnetworks. https://getnews.me/statistical-view-of-mechanistic-interpretability-shows-variance-in-eap-ig/ #eapig #mechanisticinterpretability
October 3, 2025 at 1:39 AM
2-qubit QRNNs with CNOT gates learn entanglement-based memory distinct from classical strategies, confirmed by causal tests (p<0.0001, d=0.89), but degrade to chance on IBM hardware while classical geometric strategies survive perfectly.

#QuantumML #MechanisticInterpretability #Research
Entanglement as Memory: Mechanistic Interpretability of Quantum Language Models
iq.fp2.dev
March 30, 2026 at 9:24 AM