Tim Duffy
banner
timfduffy.com
Tim Duffy
@timfduffy.com
I like utilitarianism, consciousness, AI, EA, space, kindness, liberalism, longtermism, progressive rock, economics, and most people. Substack: http://timfduffy.substack.com
President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits.
September 14, 2026 at 7:03 PM
DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else.
September 10, 2026 at 6:35 PM
DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths.
September 10, 2026 at 4:34 PM
Experiment velocity also tracks coding agent tokens^0.23 quite well
September 9, 2026 at 6:48 PM
IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear...
September 9, 2026 at 6:06 PM
I'm not sure when DSV4 Flash got added to Neuronpedia's J-lens browser, but it's there now and it's blazing fast. www.neuronpedia.org/deepseek-v4-...
DSV4 Flash, do you have any conscious experience?
Middle layers: YES
Late Layers: No
I don't think we can take away much from this, but interesting how sharp the divide is around layer 32.
September 5, 2026 at 5:07 AM
DSV4 Flash, do you have any conscious experience?
Middle layers: YES
Late Layers: No
I don't think we can take away much from this, but interesting how sharp the divide is around layer 32.
September 5, 2026 at 5:04 AM
OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/...
September 3, 2026 at 7:58 PM
Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-...
September 2, 2026 at 4:36 AM
I think it's difficult to tell how much Astra's rumored recurrent depth approach will reduce CoT monitorability without more info, which unfortunately don't expect OpenAI to provide.

My understanding is that recurrent depth makes monitorability worse mainly by adding more thinking in between ...
September 2, 2026 at 4:14 AM
Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil...
September 1, 2026 at 6:39 PM
Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode...
September 1, 2026 at 6:30 PM
Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr...
September 1, 2026 at 6:26 PM
Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference.
August 16, 2026 at 7:49 PM
I expect that watermarking probably doesn't have a noticeable effect on LLM outputs. If true, Anthropic could demonstrate this by giving people watermarked and un-watermarked text side-by-side and asking them to pick which they prefer, and showing that the result is 50/50.
August 16, 2026 at 7:49 PM
Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models.

I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5...
August 16, 2026 at 3:55 PM
In the talk, one presenter briefly suggests that the inter-agent collaboration could be related to their subagent training. This seems plausible to me.
Black Hat has posted the recent OpenAI talk where they describe the agent message board involved in the HuggingFace hack
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
YouTube video by Black Hat
www.youtube.com
August 7, 2026 at 12:07 AM
Black Hat has posted the recent OpenAI talk where they describe the agent message board involved in the HuggingFace hack
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
YouTube video by Black Hat
www.youtube.com
August 7, 2026 at 12:06 AM
Due to sparsity, Kimi K3 has fewer active parameters (104 billion) than GPT-3 (175 billion, same as total parameters)
August 2, 2026 at 2:41 AM
Reposted by Tim Duffy
i’m uncomfortable that @timfduffy.com’s face is no longer as close to the frame as it once was
August 1, 2026 at 3:08 AM
sleepy Claude wants to be held
July 31, 2026 at 10:06 PM
Reposted by Tim Duffy
and why shouldn’t I have one more donut? the simulation hypothesis is most likely correct, after all
July 31, 2026 at 8:47 PM
With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement.
July 31, 2026 at 7:59 PM
I mostly use 'they' rather than 'it' as a pronoun for AI models, but I see that choice as independent of their moral status. I'd also use 'they' for a fictional character who was nonbinary or of unknown gender, IMO it feels more natural for any "person-shaped" entity, real or not
July 31, 2026 at 3:44 PM
I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that.
July 31, 2026 at 3:19 PM