#SparseAttention
FlashInfer v0.2 by @yzh119.bsky.social

FlashInfer is a library and kernel generator for Large Language Models that provides high-performance implementation of LLM GPU kernels such as FlashAttention, SparseAttention, PageAttention, Sampling, and more.
December 19, 2024 at 9:54 PM
Speed up your LLMs! IndexCache’s sparse attention drops long‑context inference time by 1.82×, blending dense‑sparse tricks inside transformer blocks. Curious how it works? Dive in for the details. #IndexCache #SparseAttention #LongContextAI

🔗 aidailypost.com/news/indexca...
March 27, 2026 at 6:10 PM
September 30, 2025 at 9:02 PM
MiniMax M3 Explained: The Sparse Attention Breakthrough This article was originally published on GetYourDozAi . Key Takeaways MiniMax M3 — the first open-weight model to combine frontier coding, ...

#minimax #ai #machinelearning #sparseattention

Origin | Interest | Match
MiniMax M3 Explained: The Sparse Attention Breakthrough
This article was originally published on GetYourDozAi. * Key Takeaways MiniMax M3 — the...
dev.to
June 24, 2026 at 2:20 AM
#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...
September 11, 2026 at 9:30 AM
#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...
September 11, 2026 at 9:36 AM
DeepSeek unveils V3.2-exp model with sparse attention, slashing AI inference costs by up to 50%. A game-changer for long-context operations! #AI #DeepSeek #SparseAttention #TechInnovation Link: thedailytechfeed.com/deepseeks-v3...
September 30, 2025 at 3:42 PM
MSA brings 100M token context to LLMs, but there's a catch. We break down how sparse attention trades deep reasoning for massive scale, and what it means for AI's future.

https://thepixelspulse.com/posts/msa-memory-sparse-attention-llm-tradeoffs/

#msa #sparseattention #llm
March 24, 2026 at 2:15 PM
LLMs just got a massive context upgrade! MiniMax Sparse Attention cuts per-token compute by 28.4x at 1M tokens & delivers 14.2x faster inference. The future of agentic AI is here. #AI #LLM #SparseAttention

https://www.startuphub.ai/ai-news/ai-research/2026/unlocking-ultra-long-context-for-llms
Unlocking Ultra-Long Context for LLMs
MiniMax Sparse Attention breaks the context window barrier for LLMs, enabling millions of tokens with significant compute reduction and practical speedups.
www.startuphub.ai
June 12, 2026 at 8:02 PM
DeepSeek‑V4.1‑Flash pushes 1M context with KV‑cache tricks, sparse attention & multimodal tokens. Open weights + vLLM turn it into a transformer playground. Curious? Dive into the details! #DeepSeekV4_1_Flash #1MContext #SparseAttention

🔗 aidailypost.com/news/deepsee...
September 10, 2026 at 7:42 AM
[JP] 100万トークンの学習を劇的に高速化!世界初のオープンソース『Flash-MSA』がHopper/Blackwell向けに解禁!
[EN] Dramatically Speeding Up Million-Token Training! The World''s First …

https://ai-minor.com/blog/en/2026-07-13-1783921852603-flash_msa__accelerating_million_token_training_wit

#Flash-MSA #SparseAttention #Blackwell #AI #Tech
Dramatically Speeding Up Million-Token Training! The World''s First Open-Source ''Flash-MSA'' Unleashed for Hopper/Blackwell!
Dramatically Speeding Up Million-Token Training! The World''s First Open-Source ''Flash-MSA'' Unleashed for Hopper/Blackwell!
ai-minor.com
July 13, 2026 at 6:33 AM
A new input‑aware sparse attention method cuts compute by up to 40% and enables real‑time co‑speech video generation with better lip‑sync. Submitted 2 Oct 2025. https://getnews.me/input-aware-sparse-attention-enables-real-time-co-speech-video-generation/ #sparseattention #realtime
October 6, 2025 at 6:25 AM
DeepSeek released the V3.2‑exp model on Monday, using sparse attention to cut API costs for long‑context tasks by half. The model and weights are open‑source on Hugging Face. Read more: https://getnews.me/deepseek-unveils-sparse-attention-model-to-halve-api-costs/ #deepseek #sparseattention
October 2, 2025 at 6:23 PM
ProxyAttn, a training‑free method without additional training using representative heads, claims up to 10.3× faster raw attention and 2.4× speed‑up in LLM pre‑fill. Read more: https://getnews.me/proxyattn-introduces-guided-sparse-attention-using-representative-heads/ #proxyattn #sparseattention
September 30, 2025 at 11:38 PM
FG‑Attn applies fine‑grained sparse attention, yielding an average 1.55× speedup on a single NVIDIA H100 GPU for five‑second 480p clips (up to 1.65×). Read more: https://getnews.me/fg-attn-brings-fine-grained-sparse-attention-to-speed-video-diffusion/ #fgattn #sparseattention
September 24, 2025 at 8:07 AM
DeepSeek V3.2 just dropped a major upgrade—sparse attention for long‑context, solid tool‑use reasoning, and built‑in formatting cues. Open‑source power is finally ready for production. Dive into the details! #DeepSeekV32 #OpenSourceLLM #SparseAttention

🔗 aidailypost.com/news/deepsee...
December 3, 2025 at 12:39 AM