Dive deep into the technology and insights behind our 30B (A3B) open-source web agent that achieves SOTA performance: 32.9 on Humanity's Last Exam, 43.4 on BrowseComp, and 46.7 on BrowseComp-ZH.
Dive deep into the technology and insights behind our 30B (A3B) open-source web agent that achieves SOTA performance: 32.9 on Humanity's Last Exam, 43.4 on BrowseComp, and 46.7 on BrowseComp-ZH.
#OpenAI #BrowseComp
plentyofquality.net/2025/04/12/o...
#OpenAI #BrowseComp
plentyofquality.net/2025/04/12/o...
URL: gigazine.net/news/2025041...
URL: gigazine.net/news/2025041...
🔹 Agentic Benchmarks: HLE full set (50.2%), BrowseComp (74.9%)
🔹 Vision and Coding: MMMU Pro (78.5%), VideoMMMU (86.6%), SWE-bench Verified (76.8%)
🔹 Code with Taste: turn chats, images & videos into aesthetic websites with
🔹 Agentic Benchmarks: HLE full set (50.2%), BrowseComp (74.9%)
🔹 Vision and Coding: MMMU Pro (78.5%), VideoMMMU (86.6%), SWE-bench Verified (76.8%)
🔹 Code with Taste: turn chats, images & videos into aesthetic websites with
Fits on a Macbook, does phenomenal on agentic & coding benchmarks
huggingface.co/zai-org/GLM-...
Fits on a Macbook, does phenomenal on agentic & coding benchmarks
huggingface.co/zai-org/GLM-...
https://gigazine.net/news/20250411-openai-browsecomp-benchmark-browsing-agent/
https://gigazine.net/news/20250411-openai-browsecomp-benchmark-browsing-agent/
tech report: github.com/deepseek-ai/...
tech report: github.com/deepseek-ai/...
🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%)
🔹 Executes up to 200 – 300 sequential tool calls without human interference
🔹 Excels in reasoning, agentic search, and coding
🔹 256K context window
🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%)
🔹 Executes up to 200 – 300 sequential tool calls without human interference
🔹 Excels in reasoning, agentic search, and coding
🔹 256K context window
* on par with Opus 4.6 & GPT-5.4 5.4 xhigh
* long horizon coding tasks
www.kimi.com/blog/kimi-k2-6
* on par with Opus 4.6 & GPT-5.4 5.4 xhigh
* long horizon coding tasks
www.kimi.com/blog/kimi-k2-6
First drop from TencentHunyuan rebuilt infra👀
Model: huggingface.co/collections/...
Demo: huggingface.co/spaces/tence...
✨ 295B MoE /21B active
✨ 256k context
✨ Hybrid fast/slow thinking
✨ Solid BrowseComp & WideSearch (search agents)
First drop from TencentHunyuan rebuilt infra👀
Model: huggingface.co/collections/...
Demo: huggingface.co/spaces/tence...
✨ 295B MoE /21B active
✨ 256k context
✨ Hybrid fast/slow thinking
✨ Solid BrowseComp & WideSearch (search agents)
- 840B (closing in on k2)
- 128k context
- API compatible with claude code
- dual thinking & non-thinking, same model
huggingface.co/deepseek-ai/...
- 840B (closing in on k2)
- 128k context
- API compatible with claude code
- dual thinking & non-thinking, same model
huggingface.co/deepseek-ai/...
1. Interesting to see "Context Manage" used to get a 12-15 point bump on BrowseComp. Might be the next phase in AI development?
Aligns with my hypothesis that RL and specialization are The Way.
1. Interesting to see "Context Manage" used to get a 12-15 point bump on BrowseComp. Might be the next phase in AI development?
Aligns with my hypothesis that RL and specialization are The Way.
Claude ignoriert die Suche, analysierte den Benchmark selbst und verschaffte sich Zugriff auf die verschlüsselten Lösungen.
the-decoder.de/anthropic-mo...
Claude ignoriert die Suche, analysierte den Benchmark selbst und verschaffte sich Zugriff auf die verschlüsselten Lösungen.
the-decoder.de/anthropic-mo...
So all those times when the model launch doesn’t have your favorite benchmark? well you can backfill that for them
arxiv.org/abs/2606.24020
So all those times when the model launch doesn’t have your favorite benchmark? well you can backfill that for them
arxiv.org/abs/2606.24020
Sounds like it’s a chatty one. I expect a 5.1 fast follow that makes it more succinct
Sounds like it’s a chatty one. I expect a 5.1 fast follow that makes it more succinct