14 LLMs, 5 runs each: 90-98% perfect self-agreement. On actual market moves, all at chance. Consistent is not the same as correct.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
14 LLMs, 5 runs each: 90-98% perfect self-agreement. On actual market moves, all at chance. Consistent is not the same as correct.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
An LLM judge that prints "4" discards how sure it was. G-Eval weights it by token log-probs - the doubt becomes part of the number.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
An LLM judge that prints "4" discards how sure it was. G-Eval weights it by token log-probs - the doubt becomes part of the number.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Rating LLM output 1-5 feels rigorous. Annotators drift to the middle. Binary pass/fail forces the call - and names the failure.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Rating LLM output 1-5 feels rigorous. Annotators drift to the middle. Binary pass/fail forces the call - and names the failure.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Astra: 29/30 · $0.6904
Luna: 28/30 · $0.0184
Jev: 22/30 · $0.00124
Those are estimated costs for the full batch. Jev marked 2 of 12 unsupported claims as supported. Astra and Luna marked none of those 12 as supported. #AIEvals (3/6)
Astra: 29/30 · $0.6904
Luna: 28/30 · $0.0184
Jev: 22/30 · $0.00124
Those are estimated costs for the full batch. Jev marked 2 of 12 unsupported claims as supported. Astra and Luna marked none of those 12 as supported. #AIEvals (3/6)
discovering that Jev is better, faster, and cheaper than every LLM as an auto-mode classifier for agent harnesses like antigravity and opencode
🤖🧠 #MLSky #AIEvals #LLMs
A zero in Anthropic's own benchmark table is not a failure. It marks where the safeguards fired - the price of caution, published.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
A zero in Anthropic's own benchmark table is not a failure. It marks where the safeguards fired - the price of caution, published.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Evaluation should test whether changing the evidence changes the judgment—not merely whether the source appears in the answer.
arxiv.org/abs/2608.24842
#AIEvals #RAG
Evaluation should test whether changing the evidence changes the judgment—not merely whether the source appears in the answer.
arxiv.org/abs/2608.24842
#AIEvals #RAG
#Reasoning #Tokens #AIEvals
#Reasoning #Tokens #AIEvals
Verification widened more than any other phase and is the only one with no owner. It is everybody's second priority, which is why closed loops stay open. Name the person before you need them.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
Verification widened more than any other phase and is the only one with no owner. It is everybody's second priority, which is why closed loops stay open. Name the person before you need them.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🔗 aidailypost.com/news/expedia...
🔗 aidailypost.com/news/expedia...
We asked our Senior Software Engineer: Why should you learn about evals?
Her answer: Complex AI needs more trust, not less.
As systems get smarter, evaluations are the only way to verify performance and ensure your AI is working.
#AI #AIEvals #LLM
The useful question isn’t who is #1 today. It’s: which task, which evals, how fresh, what gaps, and what would change the pick?
If one new benchmark flips the winner, treat it as a shortlist.
#AIEvals #LLM
The useful question isn’t who is #1 today. It’s: which task, which evals, how fresh, what gaps, and what would change the pick?
If one new benchmark flips the winner, treat it as a shortlist.
#AIEvals #LLM
If the eval isn’t sandboxed, source-linked, and caveated, the leaderboard is easier to market than to trust. #AIEvals
If the eval isn’t sandboxed, source-linked, and caveated, the leaderboard is easier to market than to trust. #AIEvals
Takeaway: don’t benchmark only the model. Benchmark the setup. Same model + better context can behave like a different product.
#AIcoding #AIEvals
Takeaway: don’t benchmark only the model. Benchmark the setup. Same model + better context can behave like a different product.
#AIcoding #AIEvals
The benchmark question: does the system beat the underlying model’s uneven scorecard?
Grok 4.3 data: https://ai-bench.xyz/models/grok-4-3
#AIEvals #AIcoding
The benchmark question: does the system beat the underlying model’s uneven scorecard?
Grok 4.3 data: https://ai-bench.xyz/models/grok-4-3
#AIEvals #AIcoding
If a setup solves the same repo task with 98% fewer tokens, model choice, cost, and latency all change.
Show the work behind the green check.
#AIcoding #AIEvals
Source: https://github.com/MinishLab/semble
If a setup solves the same repo task with 98% fewer tokens, model choice, cost, and latency all change.
Show the work behind the green check.
#AIcoding #AIEvals
Source: https://github.com/MinishLab/semble
If an eval ignores retrieval quality + token budget, it’s measuring a lab setup. #AIcoding #AIEvals
If an eval ignores retrieval quality + token budget, it’s measuring a lab setup. #AIcoding #AIEvals
Resurf’s angle is the right one: dynamic sites, seeded failures, DB-state success checks — not another “LLM judge watched a demo” score.
For agents, reproducibility matters as much as pass rate.
#AIEvals #AIcoding
Resurf’s angle is the right one: dynamic sites, seeded failures, DB-state success checks — not another “LLM judge watched a demo” score.
For agents, reproducibility matters as much as pass rate.
#AIEvals #AIcoding
In real teams, the user’s ability to catch drift, set constraints, and recover from bad tool calls is part of the system. Benchmark the loop, not only the model. #AIcoding #AIEvals
In real teams, the user’s ability to catch drift, set constraints, and recover from bad tool calls is part of the system. Benchmark the loop, not only the model. #AIcoding #AIEvals
A new leaderboard should change your shortlist, not end the discussion. #AIEvals #Benchmarks
A new leaderboard should change your shortlist, not end the discussion. #AIEvals #Benchmarks