arxiv.org/abs/2501.15740
arxiv.org/abs/2501.15740
as usual, you are mystified by systems beyond your understanding and come to a bad conclusion (coincidentally the one you know will get you yelled at)
No.
There's a great overview, which is from this paper: arxiv.org/abs/2211.08943
My take: I prefer interpretability since the term explainability is too strong.
No.
There's a great overview, which is from this paper: arxiv.org/abs/2211.08943
My take: I prefer interpretability since the term explainability is too strong.
This is what happens when a child is raised by paranoid philosophers and interpretability researchers. +
This is what happens when a child is raised by paranoid philosophers and interpretability researchers. +
Here’s one of the first cases of AI interpretability leading to a scientific discovery -- in whales.
Here’s one of the first cases of AI interpretability leading to a scientific discovery -- in whales.
anthropic: through our interpretability research, we discovered claude imagines himself wearing a bow tie at all times
openai: we added slot machines
anthropic: through our interpretability research, we discovered claude imagines himself wearing a bow tie at all times
openai: we added slot machines
soniajoseph.ai/multimodal-i... 💜
soniajoseph.ai/multimodal-i... 💜
mechanistic interpretability studies 48h later: arxiv.org/abs/2609.16247
mechanistic interpretability studies 48h later: arxiv.org/abs/2609.16247