No.
There's a great overview, which is from this paper: arxiv.org/abs/2211.08943
My take: I prefer interpretability since the term explainability is too strong.
No.
There's a great overview, which is from this paper: arxiv.org/abs/2211.08943
My take: I prefer interpretability since the term explainability is too strong.
I love the term ‘traceability and explainability gaps’ as a way of saying ‘We have no way of knowing why our LLM bots do things and we’ve suddenly realized that’s a problem’.
I love the term ‘traceability and explainability gaps’ as a way of saying ‘We have no way of knowing why our LLM bots do things and we’ve suddenly realized that’s a problem’.
Reply or DM if you want to be added, and help me reach others!
go.bsky.app/DZv6TSS
Reply or DM if you want to be added, and help me reach others!
go.bsky.app/DZv6TSS
Graze's LLM & feed builder also provides some explainability since you can look at the blocks that the LLM generates for you.
As it turns out, graze.social has the technology. Creating feeds can be as easy as pie. Give it a try!
No AI required. Just I.
Graze's LLM & feed builder also provides some explainability since you can look at the blocks that the LLM generates for you.
Meeting Rm 205-207 @ 9am - amazing talks by @surbhigoel.bsky.social @sanmikoyejo.bsky.social Baharan Mirzasoleiman, Robert Geirhos, @coallaoh.bsky.social + exciting contributed talks!
Details: attrib-workshop.cc
Register now: learning.ecmwf.int/course/view....
Register now: learning.ecmwf.int/course/view....
Model explainability (e.g. Shap values) is becoming increasingly popular to open the "black-box" of complex analytical models. Unfortunately, opening black-boxes often reveals things of limited interest to readers and may even mislead
doi.org/10.1016/S2589-7500(21)00208-9
Model explainability (e.g. Shap values) is becoming increasingly popular to open the "black-box" of complex analytical models. Unfortunately, opening black-boxes often reveals things of limited interest to readers and may even mislead
doi.org/10.1016/S2589-7500(21)00208-9
The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't.
📄Paper: dl.acm.org/doi/10.1145/...
At #FAccT2025? Let's connect if you're interested in improving the usability of explainability methods!
📄Paper: dl.acm.org/doi/10.1145/...
At #FAccT2025? Let's connect if you're interested in improving the usability of explainability methods!
by @norabelrose.bsky.social et al.
An open-source pipeline for finding interpretable features in LLMs with sparse autoencoders and automated explainability methods from @eleutherai.bsky.social.
arxiv.org/abs/2410.13928
by @norabelrose.bsky.social et al.
An open-source pipeline for finding interpretable features in LLMs with sparse autoencoders and automated explainability methods from @eleutherai.bsky.social.
arxiv.org/abs/2410.13928
doi.org/10.1017/cfc....
doi.org/10.1017/cfc....
I'll be recruiting PhD students this upcoming cycle for fall 2026. (And if you're a UMD grad student, sign up for my fall seminar!)