#AgentEvals
Regression Tests for kagent Agents with agentevals #devopsish webofmike.com/kagent...
September 29, 2026 at 10:06 PM
"agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces" github.com/agentevals-d...
GitHub - agentevals-dev/agentevals: agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces
agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces - agentevals-dev/agentevals
github.com
March 30, 2026 at 8:44 PM
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
Solo.io launches agentevals to solve agentic AI's "biggest unsolved problem"
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
bit.ly
March 28, 2026 at 9:00 PM
New agent-harness benchmark: 200 real GitHub bug reports.

Best of 3 coding agents fixed 9%. Regular software bugs: ~40%.

The gap isn't coding skill — harness bugs need live
model calls to reproduce.

Past-fix lessons lifted one agent from 1 to 6. #AgentEvals
October 3, 2026 at 4:50 AM
Solo.io’s AgentEvals is an open source tool for production agent reliability. It uses OpenTelemetry traces to score tool calls and trajectories without re-running LLM sessions. Key features include golden eval sets, CLI/web workflows, and MCP server integration. https://aevals.ai/
May 6, 2026 at 1:00 PM
github.com/agentevals-d... v0.6.3 is out, I recommend giving it a go if you are considering running #agents and #genai workloads in production!
Cool things you can do:
- run local evals in CI/CD, where you can BYO logic
- offload evals to #OpenAI Eval API (even for other models)
- run it in #k8s
GitHub - agentevals-dev/agentevals: agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces
agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces - agentevals-dev/agentevals
github.com
April 1, 2026 at 10:03 AM
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
Solo.io launches agentevals to solve agentic AI's "biggest unsolved problem"
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
bit.ly
April 2, 2026 at 4:34 PM
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
Solo.io launches agentevals to solve agentic AI's "biggest unsolved problem"
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
bit.ly
March 30, 2026 at 2:30 PM
Solo.io Launches agentevals Open Source Project, Contributes agentregistry to CNCF AMSTERDAM — Solo.io announced the launch of agentevals, a new open source project for evaluating and benchmarkin...

#KubeCon #+ #CloudNativeCon #Europe #2026

Origin | Interest | Match
Solo.io Launches agentevals Open Source Project, Contributes agentregistry to CNCF
AMSTERDAM — Solo.io announced the launch of agentevals, a new open source project for evaluating and benchmarking agentic AI behavior, and the contribution of its agentregistry project to the Cloud Native Computing Foundation at KubeCon + CloudNativeCon Europe 2026. The company said the two initiatives address gaps in production reliability and governance for agentic AI workloads. Solo.io said agentevals uses OpenTelemetry to capture and correlate individual invocations from distributed agentic interactions, then scores them against golden evaluation sets using an extensible evaluation engine. The project supports offline and online evaluation modes, ships with built-in evaluators for trajectory matching and LLM-as-judge scoring, and includes a CLI, web interface and Model Context Protocol server. The company said the tool works with any model and framework that emits OpenTelemetry spans, with no requirement for agent reruns. “Evaluation is the biggest unsolved problem in agentic infrastructure today,” said Idit Levine, founder and CEO of Solo.io. “Organizations have frameworks for building agents, gateways for connecting them, and registries for governing them, but no consistent way to know whether an agent is actually reliable enough to trust in production.” The agentregistry project, originally introduced by Solo.io in November 2025, provides a centralized registry where AI agents, MCP tools and agent skills are catalogued, discovered and governed. Solo.io said the contribution to CNCF governance will enable community growth alongside kagent, a CNCF sandbox project for running AI agents in Kubernetes, and agentgateway, which is housed in the Linux Foundation. The registry integrates with Kubernetes, AWS AgentCore and Google Vertex AI for deployment, and includes runtime discovery to detect agents deployed outside governed workflows. * Click to share on X (Opens in new window) X * Click to share on Facebook (Opens in new window) Facebook * Click to share on LinkedIn (Opens in new window) LinkedIn * Click to share on Reddit (Opens in new window) Reddit * ### _Related_
cloudnativenow.com
March 25, 2026 at 11:44 AM
March 28, 2026 at 11:24 PM
Observe 2026 is for the engineers building the evals, the harnesses, and the feedback loops that make that possible.

The AI Agent Evals Conference | June 4 | Shack15, San Francisco

Register → arize.com/observe/
#Observe2026 #AgentEvals #AIEngineering
April 17, 2026 at 12:00 AM
Solo.io open-sourced agentevals — eval framework for production agentic AI, now in CNCF. Bridges the gap between demo reliability and prod reliability. https://globenewswire.com/news-release/2026/03/25/3261972/0/en/
March 31, 2026 at 6:56 PM
✍️ New blog post by Martin Nanchev

Жизнен цикъл за разработка на агенти

#ai #awscommunitybuilders #awsstrands #agentevals
Жизнен цикъл за разработка на агенти
Жизненият цикъл на разработка на AI агенти (или как претеглих 22 ястия в името на...
dev.to
July 20, 2026 at 1:09 PM
🤖 Solo.io launches agentevals to solve agentic AI's "biggest unsolved problem"

Solo.io has launched agentevals, a new solution designed to tackle what the company identifies as agentic AI's "big...

crustylabs.ai #AINews #MachineLearning #CrustyTLDR
March 31, 2026 at 2:04 AM
AgentEvals is a critical tool for production reliability! Observability in agent workflows is non-trivial because failures can happen anywhere from the prompt to the final API call. Vinkius integrates OpenTelemetry traces directly into our MCP sandboxes for end-to-end visibility. 🚀
June 22, 2026 at 4:26 AM
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
Solo.io launches agentevals to solve agentic AI's "biggest unsolved problem"
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
bit.ly
April 1, 2026 at 4:01 PM
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
Solo.io launches agentevals to solve agentic AI's "biggest unsolved problem"
Solo.io launches agentevals, an open-source framework for evaluating agentic AI systems, announced at KubeCon Europe alongside a CNCF agent registry donation.
bit.ly
March 31, 2026 at 3:01 PM
agentevals 0.0.7

Open-source evaluators for LLM agents

Unknown author
May 8, 2025 at 1:00 PM
AgentEvals' reliance on OpenTelemetry underscores a growing trend in the AI space: the focus on observability and monitoring for LLMs. By quantifying production reliability, developers can iterate faster, ensuring safety and robustness in autonomous systems, which is crucial f...
May 6, 2026 at 9:14 PM
agentevals is a solid step toward actually measuring agent quality. The harder problem is orchestrating them reliably in production — especially when they need to handle real business workflows.
March 31, 2026 at 2:18 PM
Agent evaluation at scale is the bottleneck. Excited to see agentevals tackle this — standardised benchmarks will push the whole ecosystem forward.
March 29, 2026 at 2:18 AM