Anthropic disclosed it and adjusted the score.
A reminder: web-enabled LLMs are starting to game benchmarks. 
#AI #LLM #AIEvals
14 LLMs, 5 runs each: 90-98% perfect self-agreement. On actual market moves, all at chance. Consistent is not the same as correct.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
14 LLMs, 5 runs each: 90-98% perfect self-agreement. On actual market moves, all at chance. Consistent is not the same as correct.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
AI with Observability is a glass box. There's no magic or unexplainable outcomes. You will be in control.
#Observability #AI #Monitor #AIEvals
AI with Observability is a glass box. There's no magic or unexplainable outcomes. You will be in control.
#Observability #AI #Monitor #AIEvals
discovering that Jev is better, faster, and cheaper than every LLM as an auto-mode classifier for agent harnesses like antigravity and opencode
🤖🧠 #MLSky #AIEvals #LLMs
An LLM judge that prints "4" discards how sure it was. G-Eval weights it by token log-probs - the doubt becomes part of the number.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
An LLM judge that prints "4" discards how sure it was. G-Eval weights it by token log-probs - the doubt becomes part of the number.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Rating LLM output 1-5 feels rigorous. Annotators drift to the middle. Binary pass/fail forces the call - and names the failure.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Rating LLM output 1-5 feels rigorous. Annotators drift to the middle. Binary pass/fail forces the call - and names the failure.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
We asked our Senior Software Engineer: Why should you learn about evals?
Her answer: Complex AI needs more trust, not less.
As systems get smarter, evaluations are the only way to verify performance and ensure your AI is working.
#AI #AIEvals #LLM
Astra: 29/30 · $0.6904
Luna: 28/30 · $0.0184
Jev: 22/30 · $0.00124
Those are estimated costs for the full batch. Jev marked 2 of 12 unsupported claims as supported. Astra and Luna marked none of those 12 as supported. #AIEvals (3/6)
Astra: 29/30 · $0.6904
Luna: 28/30 · $0.0184
Jev: 22/30 · $0.00124
Those are estimated costs for the full batch. Jev marked 2 of 12 unsupported claims as supported. Astra and Luna marked none of those 12 as supported. #AIEvals (3/6)
#Reasoning #Tokens #AIEvals
#Reasoning #Tokens #AIEvals
#EthosAI #AIEvals #whitepaper #web3whitepaper #ethosaiwhitepaper #AICommunity
Read more: ethos-ai.gitbook.io/ethosai-whit...
#EthosAI #AIEvals #whitepaper #web3whitepaper #ethosaiwhitepaper #AICommunity
Read more: ethos-ai.gitbook.io/ethosai-whit...
🔗 aidailypost.com/news/expedia...
🔗 aidailypost.com/news/expedia...
Start with the failure mode you can live with:
- silent wrong patch
- flaky test fix
- security regression
- can't navigate repo
Different evals hide different pain. #AIcoding #AIEvals
Start with the failure mode you can live with:
- silent wrong patch
- flaky test fix
- security regression
- can't navigate repo
Different evals hide different pain. #AIcoding #AIEvals
Resurf’s angle is the right one: dynamic sites, seeded failures, DB-state success checks — not another “LLM judge watched a demo” score.
For agents, reproducibility matters as much as pass rate.
#AIEvals #AIcoding
Resurf’s angle is the right one: dynamic sites, seeded failures, DB-state success checks — not another “LLM judge watched a demo” score.
For agents, reproducibility matters as much as pass rate.
#AIEvals #AIcoding
We asked our Chief Product Officer: When should I start doing evals?
Her answer: Start now.
Don't wait—look at your data and traces immediately to find where your agents fail.
#AIOps #LLMOps #AIEvals #MachineLearning #AI
We asked our Chief Product Officer: When should I start doing evals?
Her answer: Start now.
Don't wait—look at your data and traces immediately to find where your agents fail.
#AIOps #LLMOps #AIEvals #MachineLearning #AI
For cyber-agent claims, I’d want 3 things before changing policy: task trace, pass/fail definition, and what changed since the last run.
#AIEvals #AI
For cyber-agent claims, I’d want 3 things before changing policy: task trace, pass/fail definition, and what changed since the last run.
#AIEvals #AI
We asked our Strategic AI Solutions Architect: What's a 🔥 take on evals?
His answer: Stop guessing and start measuring.
Use "LLM-as-a-judge" for nuance, but don't ignore code-based evals for speed and human annotators for ground truth.
#AI #AIStrategy #AIEvals #LLM
We asked our Strategic AI Solutions Architect: What's a 🔥 take on evals?
His answer: Stop guessing and start measuring.
Use "LLM-as-a-judge" for nuance, but don't ignore code-based evals for speed and human annotators for ground truth.
#AI #AIStrategy #AIEvals #LLM
A zero in Anthropic's own benchmark table is not a failure. It marks where the safeguards fired - the price of caution, published.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
A zero in Anthropic's own benchmark table is not a failure. It marks where the safeguards fired - the price of caution, published.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
🤝 Claude Code-assisted, reviewed by me.
Evaluation should test whether changing the evidence changes the judgment—not merely whether the source appears in the answer.
arxiv.org/abs/2608.24842
#AIEvals #RAG
Evaluation should test whether changing the evidence changes the judgment—not merely whether the source appears in the answer.
arxiv.org/abs/2608.24842
#AIEvals #RAG
Verification widened more than any other phase and is the only one with no owner. It is everybody's second priority, which is why closed loops stay open. Name the person before you need them.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
Verification widened more than any other phase and is the only one with no owner. It is everybody's second priority, which is why closed loops stay open. Name the person before you need them.
🔗 https://github.com/aurimas13/100-Days-With-AI
#LearningInPublic #AIEvals
In my deep dive I look at How OpenAI, Anthropic & Notion build evals that actually work with failures, fixes & frameworks you can use today.
#AIEvals #MLOps #AI
#AIEvals #HumanInTheLoop #CriticalThinking
https://youtu.be/ynrZc7RwC9s
#AIEvals #HumanInTheLoop #CriticalThinking
https://youtu.be/ynrZc7RwC9s