#AIevals
Day 74/100 of #100DaysWithAI 🔁

14 LLMs, 5 runs each: 90-98% perfect self-agreement. On actual market moves, all at chance. Consistent is not the same as correct.

🔗 https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🤝 Claude Code-assisted, reviewed by me.
September 23, 2026 at 2:54 AM
Day 73/100 of #100DaysWithAI 🎲

An LLM judge that prints "4" discards how sure it was. G-Eval weights it by token log-probs - the doubt becomes part of the number.

🔗 https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🤝 Claude Code-assisted, reviewed by me.
September 22, 2026 at 2:50 AM
Day 72/100 of #100DaysWithAI 📏

Rating LLM output 1-5 feels rigorous. Annotators drift to the middle. Binary pass/fail forces the call - and names the failure.

🔗 https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🤝 Claude Code-assisted, reviewed by me.
September 21, 2026 at 3:03 AM
Support agreement with the locked reference answers:
Astra: 29/30 · $0.6904
Luna: 28/30 · $0.0184
Jev: 22/30 · $0.00124

Those are estimated costs for the full batch. Jev marked 2 of 12 unsupported claims as supported. Astra and Luna marked none of those 12 as supported. #AIEvals (3/6)
September 20, 2026 at 11:15 PM
what i be doing at 3am:
discovering that Jev is better, faster, and cheaper than every LLM as an auto-mode classifier for agent harnesses like antigravity and opencode
🤖🧠 #MLSky #AIEvals #LLMs
Jev: a better shell-command safety classifier at half the price of DeepSeek 4.1 Flash
Jev let none of 530 dangerous commands through, at half the cost of DeepSeek 4.1 Flash, which is itself excellent for the price.
famelos.com
September 19, 2026 at 9:05 AM
What should every model evaluation record? Task type, context needs, capability fit, and reference price. TemprHQ helps you fill that worksheet across 407 models. #LLMEvals #AIEvals https://temprhq.io/models
September 18, 2026 at 9:12 PM
Day 54/100 of #100DaysWithAI 🛡️

A zero in Anthropic's own benchmark table is not a failure. It marks where the safeguards fired - the price of caution, published.

🔗 https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🤝 Claude Code-assisted, reviewed by me.
September 3, 2026 at 3:55 AM
AI can retrieve the right evidence and still ignore it when making the decision. That's more dangerous than simply missing document.
Evaluation should test whether changing the evidence changes the judgment—not merely whether the source appears in the answer.
arxiv.org/abs/2608.24842
#AIEvals #RAG
Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they...
arxiv.org
August 26, 2026 at 9:13 AM
Model upgrades silently regress behavior, which is why evals belong in CI. Golden datasets with graded expectations catch the subtle breaks that smoke tests miss. If your prompt pipeline can ship worse output without anyone noticing, you have no quality control. #AIEvals #LLMOps #Quality
August 20, 2026 at 10:47 AM
With adaptive thinking the model decides how long to reason per request. Teams running eval suites should re-baseline after enabling it because token usage patterns shift per task
#Reasoning #Tokens #AIEvals
August 20, 2026 at 3:40 AM
Day 31/100 of #100DaysWithAI ✅

Verification widened more than any other phase and is the only one with no owner. It is everybody's second priority, which is why closed loops stay open. Name the person before you need them.

🔗 https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals
August 11, 2026 at 4:32 AM
Expedia’s AI chief says you should stay in the driver’s seat—AI agents only help, never decide. Think red‑teaming, Gemini tricks, and real‑world evals. Want the full scoop on keeping humans in charge? #AIagents #AIEvals #RedTeamTesting

🔗 aidailypost.com/news/expedia...
July 22, 2026 at 12:16 AM
Who judges an LLM-as-a-judge though? #AIEvals
June 13, 2026 at 7:01 PM
AI agents aren't perfect: Learn how to catch silent failures & hallucinations in my latest blog tutorial 🕵️‍♀️🤖 Dive into LLM-as-Judge techniques! #AIEvals
How to Evaluate AI Agents: LLM-as-Judge Tutorial
Evaluate AI agent quality with LLM-as-Judge and trajectory analysis. Catch silent failures, wasted tokens, and hallucinations before production. Python tutorial with code.
ift.tt
June 3, 2026 at 1:05 PM
🛠️ One AI Question with Elizabeth Hutton

We asked our Senior Software Engineer: Why should you learn about evals?

Her answer: Complex AI needs more trust, not less.

As systems get smarter, evaluations are the only way to verify performance and ensure your AI is working.

#AI #AIEvals #LLM
May 21, 2026 at 3:34 PM
Seeing a lot of “best AI model in May” leaderboards again.

The useful question isn’t who is #1 today. It’s: which task, which evals, how fresh, what gaps, and what would change the pick?

If one new benchmark flips the winner, treat it as a shortlist.

#AIEvals #LLM
May 20, 2026 at 12:43 PM
That agent-benchmark exploit thread is a good reminder: a high score can mean “solved the task” or “found a hole in the harness.”

If the eval isn’t sandboxed, source-linked, and caveated, the leaderboard is easier to market than to trust. #AIEvals
May 20, 2026 at 6:39 AM
That DCBench result is the coding-agent eval I’d watch: context changed decision compliance from 46% to 95%.

Takeaway: don’t benchmark only the model. Benchmark the setup. Same model + better context can behave like a different product.

#AIcoding #AIEvals
May 19, 2026 at 6:28 PM
Grok Build changes the eval unit: not one model answering, but 8 agents competing + Arena Mode picking.

The benchmark question: does the system beat the underlying model’s uneven scorecard?

Grok 4.3 data: https://ai-bench.xyz/models/grok-4-3

#AIEvals #AIcoding
May 19, 2026 at 4:47 PM
Coding-agent evals should report more than pass rate.

If a setup solves the same repo task with 98% fewer tokens, model choice, cost, and latency all change.

Show the work behind the green check.

#AIcoding #AIEvals

Source: https://github.com/MinishLab/semble
May 19, 2026 at 12:22 PM
The Semble thread is a useful reminder: coding-agent benchmarks often hide the boring constraint that decides real use — how much context you had to feed the model.

If an eval ignores retrieval quality + token budget, it’s measuring a lab setup. #AIcoding #AIEvals
May 19, 2026 at 6:18 AM
Browser-agent evals are finally getting more realistic.

Resurf’s angle is the right one: dynamic sites, seeded failures, DB-state success checks — not another “LLM judge watched a demo” score.

For agents, reproducibility matters as much as pass rate.

#AIEvals #AIcoding
May 18, 2026 at 12:05 PM
Cyber evals are getting weird fast. If a model suddenly beats every previous trend line, don’t ask only “is it best?” Ask: which task lengths, how many attempts, what environment, and what would make the result disappear? https://unbiased-ai-bench.vercel.app/ #AIEvals
May 18, 2026 at 9:10 AM
I’d watch the coding-agent evals that measure steering, not just solved tasks.

In real teams, the user’s ability to catch drift, set constraints, and recover from bad tool calls is part of the system. Benchmark the loop, not only the model. #AIcoding #AIEvals
May 18, 2026 at 6:55 AM
Small eval note: “GPT-5.5 is #1” on HWE Bench is interesting. But the useful bit is the task shape: hardware engineering, unbounded problems, scoring choices.

A new leaderboard should change your shortlist, not end the discussion. #AIEvals #Benchmarks
May 18, 2026 at 2:55 AM