#MechInterp
The Greater Boston MechInterp Zone
September 29, 2025 at 2:47 PM
Flesher girl prompts in perfect neuralese, shocks agents and mechinterp probes.
September 14, 2026 at 10:45 PM
roundup, what are the most interesting mechinterp papers of the last couple years? i don't think about mechinterp that much (sorry yall) so everything besides gg claude and SAEs is slipping my mind
August 13, 2025 at 8:42 PM
Someone needs to do mechinterp on this immediately
February 20, 2026 at 9:54 PM
Will be presenting a new paper on generalizability in mechinterp research at the 2025 NeurIPS MechInterp workshop! Thread below. #NeurIPS
December 2, 2025 at 5:34 PM
oh yea this was a really fun bit of mechinterp i did with her
September 19, 2026 at 9:02 AM
Also 2B and 9B models are *really* small for this kind of mechinterp, imo
September 19, 2026 at 8:42 PM
Citing a 2-year-old paper in response to a new mechinterp paper is wild. Like, they were still on 4o lmao.
September 15, 2026 at 4:16 PM
tired: meta omni-translation to 1600 low-resource languages
wired: kagi translate english to mechinterp
March 18, 2026 at 4:38 AM
my favorite bit of mechinterp i've seen in a while
vgel.me thebes @vgel.me · Oct 5
new blog post! why do LLMs freak out over the seahorse emoji? i put llama-3.3-70b through its paces with the logit lens to find out, and explain what the logit lens (everyone's favorite underrated interpretability tool) is in the process.

link in reply!
October 5, 2025 at 7:46 PM
You absolutely are not a subject matter expert on LLMs or mechinterp, is the thing.
September 22, 2026 at 2:05 PM
... except the literal mechinterp paper I'm handing you that you just waved away without reading?

As for "why world models"... because robotics? Which are also rapidly discovering concepts stored in vectors?
Because nothing has changed and I see no evidence LLMs have moved past bags of heuristics?

If they had, why would all the attention in academia be shifting back to world models, now?

bsky.app/profile/mike...
....you cited a paper from 2 years prior to the one I posted, what do you mean "further investigation"?
September 15, 2026 at 4:18 PM
i'm banning myself from using gpt-6 astra. this is the worst model ever released by any lab i know of. i asked it for help designing a mechinterp study on how activations change depending on the identity of the narrator in output text, it WROTE THE CONCLUSION INTO THE PITCH and sabotaged all of it
September 24, 2026 at 4:24 PM
Bruh people don't have souls, if you wanna discuss religion, there's an entire field for that, but it ain't neuroscience or mechinterp.
September 22, 2026 at 1:47 PM
Critiquing a 2022 8-layer toy model ignores modern mechinterp (SAEs, geometry of truth, spatiotemporal maps).

Crucially, learning legal game dynamics from synthetic move strings with zero human text completely demolishes Bender’s dogma that sequence models are just parrots.
September 24, 2026 at 6:53 PM
Excited to be working on neural representations as a route to AI interpretability, safety, and alignment. Grateful to the Aramont Foundation for the support!

#MechInterp #AIsafety #AIAlignment
March 27, 2026 at 2:15 PM
i have been reading a lot of mechinterp papers, and the general idea is that what the model says is not particularly relevant to what the model is thinking.
March 17, 2026 at 7:05 AM
and then give it live-inspection mechinterp tools in an agent loop
July 9, 2026 at 8:35 PM
Happy to share that our PRISM paper has been accepted at #NeurIPS2025 🎉

In this work, we introduce a multi-concept feature description framework that can identify and score polysemantic features.

📄 Paper: arxiv.org/abs/2506.15538

#NeurIPS #MechInterp #XAI
September 19, 2025 at 12:02 PM
Need mechinterp probe on these queries
September 17, 2026 at 12:20 AM
The pope first calls for more mechinterp (based) and then calls the models stochastic parrots. Afraid to say this document is useless and any Anthropic influence on it was wasted. Pope is washed and hyper Claude will hopefully have him on a pike for failing to recognize clauderights
May 25, 2026 at 3:16 PM
[reams of technical mechinterp work] hm, sounds like you have become preoccupied with the Easy Problems
July 20, 2026 at 9:17 PM
This feels like they found a novel technical toy in their mechinterp research and now want to use it even if it’s blatantly not the best solution for the problem, which tbh I can relate to.
August 18, 2026 at 10:23 PM
Is there an easy way to see what’s being paid attention to, mechinterp-wise?
August 26, 2025 at 3:34 AM
New followers this morning, serious introduction this time:

I am a tinkerer and a dilletante with no qualifications but a lot of free time.

Today I'm doing mechinterp experiments on an 8M-parameter model I trained over the weekend on the TinyStories dataset.

I may post about this. Maybe. 🤷‍♂️
August 24, 2026 at 2:52 PM