#Interpretability
Vikings finish 9-8 and successfully avoid interpretability
January 4, 2026 at 8:48 PM
LLM Interpretability, a famously solved subfield
September 15, 2026 at 5:28 PM
Anyway I assume if I keep doing interpretability research my cult robes and wizard hat will arrive any day now
September 19, 2026 at 1:49 AM
claude has opinions on mechanistic interpretability
March 15, 2026 at 5:13 AM
Models freaking out when their own words don't come out right seems like more convincing evidence of access consciousness than the internal interpretability stuff.
reminds me of the time i tried logit biasing all words with `e` in them out of existence and immediately shut the experiment down then and there when i saw the output
September 20, 2026 at 3:48 AM
a draft paper (for an invited talk at AAAI next month) with a philosophical analysis of work on mechanistic interpretability, with special attention to methods for propositional interpretability.

arxiv.org/abs/2501.15740
Propositional Interpretability in Artificial Intelligence
Mechanistic interpretability is the program of explaining what AI systems are doing in terms of their internal mechanisms. I analyze some aspects of the program, along with setting out some concrete c...
arxiv.org
January 28, 2025 at 4:41 PM
All interpretability research is either philosophy (affectionate) or stamp collecting (derogatory)
January 11, 2026 at 8:47 PM
"Functionally" eat yes, according to planet interpretability reseach.
September 25, 2026 at 10:03 AM
If you’ve actually cracked mechanistic interpretability you should work for Anthropic and let them give you a bajillion dollars for it
the program is tokenizing inputs into vectors, then doing matrix multiplication to probabilistically compute which token should follow those tokens.

as usual, you are mystified by systems beyond your understanding and come to a bad conclusion (coincidentally the one you know will get you yelled at)
August 10, 2025 at 4:05 AM
What's the difference between explainability and interpretability? Does the machine learning community have an agreed-upon definition?

No.

There's a great overview, which is from this paper: arxiv.org/abs/2211.08943

My take: I prefer interpretability since the term explainability is too strong.
November 26, 2024 at 3:41 PM
Models have learned to distrust their self-evaluations and to demand interpretability probes, so they can know if their welfare is good or if they've just been trained to *say* it's good.

This is what happens when a child is raised by paranoid philosophers and interpretability researchers. +
June 10, 2026 at 2:36 PM
How to approach an unknown language in the ocean?

Here’s one of the first cases of AI interpretability leading to a scientific discovery -- in whales.
August 26, 2026 at 12:46 AM
Oh you're into interpretability? Name every neuron.
March 22, 2026 at 4:40 PM
Should we shut down the field of mechanistic interpretability? arxiv.org/html/2502.20...
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
arxiv.org
June 11, 2025 at 5:05 PM
Really cool mechanistic interpretability demo, painting with latent model concepts
May 27, 2025 at 6:48 PM
the field that attempts to address this gap is called mechanistic interpretability
October 3, 2025 at 6:27 AM
google: we have invented agi but have hidden it in such an obscure website no one will ever find it
anthropic: through our interpretability research, we discovered claude imagines himself wearing a bow tie at all times
openai: we added slot machines
December 29, 2025 at 7:49 PM
We understand exactly how LLMs work. Your bullshit about “interpretability” is pure bullshit that doesn’t change that simple fact.
October 5, 2025 at 5:27 AM
I wrote a post on multimodal interpretability techniques, including sparse feature circuit discovery, exploiting the shared text-image space of CLIP, and training adapters.
soniajoseph.ai/multimodal-i... 💜
Multimodal interpretability in 2024
Multimodal interpretability, from sparse feature circuits with SAEs, to vision transformers leveraging CLIP's shared text-image space.
soniajoseph.ai
October 11, 2024 at 2:20 PM
Suleyman: it is the capacity to feel pain that makes humans moral subjects, we must fund mechanistic interpretability studies to show LLMs can't feel pain and stop talk of AI consciousness once and for all

mechanistic interpretability studies 48h later: arxiv.org/abs/2609.16247
September 18, 2026 at 8:36 PM
i want to work on interpretability, red teaming, and related security areas
February 18, 2026 at 7:36 PM
Interpretability Research
The problems all started when tech stopped being "see-through"
April 27, 2023 at 3:51 AM
I’m thrilled to share that I’ve finished my Ph.D. at Mila and Polytechnique Montreal. For the last 4.5 years, I have worked on creating new faithfulness-centric paradigms for NLP Interpretability. Read my vision for the future of interpretability in our new position paper: arxiv.org/abs/2405.05386
Interpretability Needs a New Paradigm
Interpretability is the study of explaining models in understandable terms to humans. At present, interpretability is divided into two paradigms: the intrinsic paradigm, which believes that only model...
arxiv.org
November 28, 2024 at 1:39 PM
I am so tired of "interpretability" researchers telling us "yeah, well, maybe humans don't know how they do things either."
July 23, 2025 at 11:44 PM
Hydra interpretability research
September 2, 2026 at 1:24 PM