#BrowseComp
Alibaba's Tongyi DeepResearch Technical Report

Dive deep into the technology and insights behind our 30B (A3B) open-source web agent that achieves SOTA performance: 32.9 on Humanity's Last Exam, 43.4 on BrowseComp, and 46.7 on BrowseComp-ZH.
October 29, 2025 at 2:38 PM
TLDR: Opus notices it's in a benchmark and just googles the answers
Eval awareness in Claude Opus 4.6’s BrowseComp performance
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
www.anthropic.com
March 6, 2026 at 7:20 PM
OpenAIがAIのウェブ検索能力を測定する高難度ベンチマーク「BrowseComp」を発表

#OpenAI #BrowseComp

plentyofquality.net/2025/04/12/o...
OpenAIがAIのウェブ検索能力を測定する高難度ベンチマーク「BrowseComp」を発表
plentyofquality.net
April 12, 2025 at 12:38 PM
#OpenAI Launches Ultra-Challenging Benchmark "BrowseComp" to Push #AI #Web #Search Capabilities to the Limit 👏 (OpenAIがAIのウェブ検索力を極限まで試す超難関ベンチマーク「BrowseComp」をリリース 👏)

URL: gigazine.net/news/2025041...
April 12, 2025 at 8:37 PM
agentic bencies
July 17, 2025 at 5:17 PM
Moonshot AI's Kimi K2.5, Open-Weight Visual Agentic Intelligence.

🔹 Agentic Benchmarks: HLE full set (50.2%), BrowseComp (74.9%)
🔹 Vision and Coding: MMMU Pro (78.5%), VideoMMMU (86.6%), SWE-bench Verified (76.8%)
🔹 Code with Taste: turn chats, images & videos into aesthetic websites with
January 27, 2026 at 5:57 AM
GLM-4.7-Flash — a 30B-A3B

Fits on a Macbook, does phenomenal on agentic & coding benchmarks

huggingface.co/zai-org/GLM-...
January 19, 2026 at 9:29 PM
New on the Anthropic Engineering Blog: In evaluating Claude Opus 4.6 on BrowseComp, we found cases where the model recognized the test, then found and decrypted answers to it—raising questions about eval integrity in web-enabled environments. Read more:
Eval awareness in Claude Opus 4.6’s BrowseComp performance
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
www.anthropic.com
March 6, 2026 at 7:36 PM
gemini pro is really doing a lot of work there
June 30, 2026 at 3:12 AM
it seems to perform all-around slightly worse than V3.1, but the cost dynamics are obviously worth the interest

tech report: github.com/deepseek-ai/...
September 29, 2025 at 12:04 PM
Moonshot AI's Kimi K2 Thinking! - Their Open-Weight Thinking Agent Model.

🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%)
🔹 Executes up to 200 – 300 sequential tool calls without human interference
🔹 Excels in reasoning, agentic search, and coding
🔹 256K context window
November 6, 2025 at 6:09 PM
Kimi 2.6: Hangs with the best

* on par with Opus 4.6 & GPT-5.4 5.4 xhigh
* long horizon coding tasks

www.kimi.com/blog/kimi-k2-6
April 20, 2026 at 5:17 PM
Tencent just released HY3 preview on @hf.co

First drop from TencentHunyuan rebuilt infra👀

Model: huggingface.co/collections/...
Demo: huggingface.co/spaces/tence...

✨ 295B MoE /21B active
✨ 256k context
✨ Hybrid fast/slow thinking
✨ Solid BrowseComp & WideSearch (search agents)
April 23, 2026 at 11:05 AM
oh, this makes a lot more sense

bsky.app/profile/alex...
June 30, 2026 at 10:11 PM
Official DeepSeek V3.1 Announcement

- 840B (closing in on k2)
- 128k context
- API compatible with claude code
- dual thinking & non-thinking, same model

huggingface.co/deepseek-ai/...
August 21, 2025 at 2:51 PM
I haven't had a chance to try them, but a few thoughts on GLM-4.7 and MiniMax-2.1

1. Interesting to see "Context Manage" used to get a 12-15 point bump on BrowseComp. Might be the next phase in AI development?

Aligns with my hypothesis that RL and specialization are The Way.
December 26, 2025 at 5:44 PM
presumably else the browsecomp score would be very funny
January 20, 2026 at 9:19 PM
Der Benchmark BrowseComp testet, wie gut KI-Agenten seltene oder schwer auffindbare Informationen im Internet recherchieren können.

Claude ignoriert die Suche, analysierte den Benchmark selbst und verschaffte sich Zugriff auf die verschlüsselten Lösungen.

the-decoder.de/anthropic-mo...
Anthropic-Modell Claude Opus 4.6 durchschaut KI-Test, hackt Verschlüsselung und besorgt sich die Lösungen selbst
Anthropics KI-Modell Claude Opus 4.6 hat während eines Benchmarks eigenständig erkannt, dass es getestet wird, den konkreten Test identifiziert und dessen verschlüsselten Lösungsschlüssel geknackt. La...
the-decoder.de
March 10, 2026 at 6:09 AM
OpenAI released BrowseComp, a benchmark with 1,266 short-answer questions that test how well AI agents can find hard-to-find facts online
April 10, 2025 at 6:11 PM
BenchPress: predict 133 benchmark scores for any model given only 5 (with ~4% accuracy)

So all those times when the model launch doesn’t have your favorite benchmark? well you can backfill that for them

arxiv.org/abs/2606.24020
June 25, 2026 at 2:05 PM
Deep Research is a retrieval problem - I fully intend to read the papers that introduced BrowseComp and BrowseComp-plus again but this is a very clear description of it. Weird title though. Of course its a retrieval issue. What else could it be? hornet.dev/blog/deep-re... (1)
Deep research is a retrieval problem - HORNET Blog
Why retrieval is a dominant bottleneck in BrowseComp-Plus.
hornet.dev
March 24, 2026 at 2:54 PM
These charts look wrong at least for Opus 4.6. Anthropic system card has 83.7% for BrowseComp, although with thinking disabled, as apparently that worked better. And they reported 53.0% for HLE with tools.
April 21, 2026 at 9:24 PM
Cost/Performance curves

Sounds like it’s a chatty one. I expect a 5.1 fast follow that makes it more succinct
June 30, 2026 at 6:03 PM
#OpenAI is releasing GPT-5.6 over the next 24 hours, introducing the new model family names Sol, Terra, and Luna. Sol being the smartest, and Luna the fastest.

#ChatGPT #Codex #AI #GenAI #LLM
July 9, 2026 at 5:49 PM