Andrew Gross
gross.systems
Andrew Gross
@gross.systems
Engineer at YipitData.
NYC Area

https://github.com/andrewgross/
https://gross.systems

I was told I had to add AI Engineer to my profile for the bots to find me.

Views my own, not my employer etc etc.
Agree that cost per completed task is the headline number, seeing the distribution per task across models could be interesting as well. And yeah, without something to de-correlate failures, 100 samples may be too much for most contained tasks. Unsure if that holds for a large attack surface
September 28, 2026 at 2:37 PM
Yeah TBench is, but they arent doing the "cost for success" angle directly. Probably something I need to hack up myself.
September 28, 2026 at 1:31 PM
This is really interesting, I wonder if they keep an up to date version somewhere, or perhaps I can just run it myself. Could be nice to cross reference it with the waffle view (www.tbench.ai?view=waffle) to understand the cutoff in terms of capability vs cost.
TERMINAL-BENCH
A benchmark to measure and evolve with the frontier of agent work
www.tbench.ai
September 28, 2026 at 1:01 PM
Yeah, I do find the cost (and in particular the clarity around tokens used vs cost/speed) from AA to be helpful. I believe for their benchmarks they do run N trials, but im pretty sure its a static number.
September 28, 2026 at 12:58 PM
It would be nice to see benchmarks that looked at things on a cost comparison basis for equal cost. So any things (especially security) often seem to scale based on how many tokens you throw at the problem. Is 1 run of Opus 5.5 better than 100 runs of GLM 5.3 Flash? etc
September 28, 2026 at 12:28 PM
huggingface.co/datasets/tar... this is the deepswe one, is there similar for TBench? If its the same style, it would just be taking the k=5 (or 4 for deepswe) rollouts and scoring them and picking the best (essentially converting k=5 to k=1 for actual scoring) which is nifty
huggingface.co
September 24, 2026 at 12:52 PM
For TBench 2.1, k=5 for the official run harness, I am curious if they used their model to pick the best of one of those, or if they picked the best of 4 or 5 runs of k=5, essentially getting k=20 (or k=25).
September 24, 2026 at 12:48 PM
Also curious what it would look like to have this "in the loop" for a regular LLM, helping choose actions
September 24, 2026 at 12:29 PM
Seems like the fine tune is likely on the embedding space so that you get better contrast? Hence the usage of some of the DeepSWE tasks for fine tuning. I wonder what perf is like without fine tuning.
September 24, 2026 at 12:29 PM
Does seem like they used a similar approach to CLIP but for non-image inputs? Need to spend more time understanding the training data vs what is used for a fine tune and why etc.
September 24, 2026 at 12:21 PM
They did seem to train/fine tune on the DeepSWE tasks, with a set held out for the actual testing. Not sure what the fine tuning gets here, "style" for the outputs? Shouldn't the action model be general w.r.t. these sorts of inputs and not need fine tuning?
September 24, 2026 at 12:13 PM
Looks like Bo4 for the DeepSWE runs? huggingface.co/Contrastive-... In theory Fable/Opus Pass@4 should produce similar numbers right? I guess it depends if the model is choosing the final rollout results or being used mid stream to pick actions.
Contrastive-LM/deepswe-clm-heads-8k · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
September 24, 2026 at 12:11 PM
They have the data and code out there, so can dig in to see how they did it I guess. Downside is that the SOTA scores would come from running 10x (or more) tokens on frontier models to get candidates.
September 24, 2026 at 12:08 PM
disputing claims etc. A lot of this had a very real cost in human time in the past, so much so that being someone wealthy enough to have a lawyer to do it all for your was a huge privilege. Now many more people have the ability to send a properly bureacracy coded letter than before.
September 10, 2026 at 5:56 PM
Interesting, I haven't explored much with that style of long term memory for a task like that. Curious how consistent it is.
August 23, 2026 at 2:46 AM
Ahh fair enough. The agents dont seam to really care that much about human readable naming do they. Curious how it handled the context length issue as the bundled code is a bit large even for 1mm context.
August 23, 2026 at 12:52 AM
Curious how you generated this, just pointed the agent at the minified js?
August 22, 2026 at 11:44 PM