Sonnet still ahead on SWE-bench though, while Gemini takes TerminalBench
Nice to see models getting better at different things
Sonnet still ahead on SWE-bench though, while Gemini takes TerminalBench
Nice to see models getting better at different things
LongCat-Flash-Chat!
▫️ 560B Total Params | 18.6B-31.3B Dynamic Activation
▫️ Trained on 20T Tokens | 100+ tokens/sec Inference
▫️ High Performance: TerminalBench 39.5 | τ²-Bench 67.7
LongCat-Flash-Chat!
▫️ 560B Total Params | 18.6B-31.3B Dynamic Activation
▫️ Trained on 20T Tokens | 100+ tokens/sec Inference
▫️ High Performance: TerminalBench 39.5 | τ²-Bench 67.7
LongCat-Flash-Lite - a model built on this insight.
⚙️ 68.5B Total Params(37.13B non-embedding) | 2.9B~4.5B Active
📊 High Performance: SWE-Bench 54.4 | τ²-Bench 72.8 | TerminalBench 33.75
LongCat-Flash-Lite - a model built on this insight.
⚙️ 68.5B Total Params(37.13B non-embedding) | 2.9B~4.5B Active
📊 High Performance: SWE-Bench 54.4 | τ²-Bench 72.8 | TerminalBench 33.75
qwen.ai/blog?id=qwen...
qwen.ai/blog?id=qwen...
3 model sizes, like Anthropic does. Sol appears to be a Mythos class model at peak reasoning
Also comes with max reasoning effort and a new ultra mode that uses subagents server-side behind the API
openai.com/index/previe...
3 model sizes, like Anthropic does. Sol appears to be a Mythos class model at peak reasoning
Also comes with max reasoning effort and a new ultra mode that uses subagents server-side behind the API
openai.com/index/previe...
- 560B total parameters, 18.6B–31.3B dynamic activation
- Trained on 20T tokens
- 100+ tokens/sec inference
- TerminalBench: 39.5 | τ²-Bench: 67.7
- 560B total parameters, 18.6B–31.3B dynamic activation
- Trained on 20T tokens
- 100+ tokens/sec inference
- TerminalBench: 39.5 | τ²-Bench: 67.7
With over 100 tokens/sec inference and stellar scores on TerminalBench (39.5) and τ²-Bench (67.7), LongCat-Flash-Chat isn’t just big—it’s blazing fast and benchmark-proven.
Huggingface: huggingface.co/meituan-long...
With over 100 tokens/sec inference and stellar scores on TerminalBench (39.5) and τ²-Bench (67.7), LongCat-Flash-Chat isn’t just big—it’s blazing fast and benchmark-proven.
Huggingface: huggingface.co/meituan-long...
🔹 552B-parameter MoE.
🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
https://gigazine.net/news/20260907-github-copilot-hydrafusion/
https://gigazine.net/news/20260907-github-copilot-hydrafusion/
It branches the run into many continuations and sees how they end. It then uses that signal to improve credit assignment in RL training.
2x GRPO's gains on TerminalBench-2.
It branches the run into many continuations and sees how they end. It then uses that signal to improve credit assignment in RL training.
2x GRPO's gains on TerminalBench-2.
https://github.com/dirac-run/dirac
https://news.ycombinator.com/item?id=47920787
https://github.com/dirac-run/dirac
https://news.ycombinator.com/item?id=47920787
gigazine.net/news/2026090...
gigazine.net/news/2026090...
🧵👇#AI
🧵👇#AI
https://gigazine.net/news/20260907-github-copilot-hydrafusion/
https://gigazine.net/news/20260907-github-copilot-hydrafusion/