#FlashMLA
Day 1 of DeepSeek open source week: FlashMLA

MLA=Multihead Latent Attention

One of the big innovations that made V3 such a notable model

github.com/deepseek-ai/...
GitHub - deepseek-ai/FlashMLA
Contribute to deepseek-ai/FlashMLA development by creating an account on GitHub.
github.com
February 24, 2025 at 12:25 PM
call that bunnygirl FlashMLA³ the way she Hops¹ 100² times!

¹ "Hops" is a reference to "Hopper"
² "Hopper" and "100" is a reference to the NVIDIA H100 GPU⁴
³ FlashMLA is an operator which is specially designed for NVIDIA Hopper GPUs
⁴ Can also work on other Hopper GPUs (but H100 is most prevalent)
June 23, 2025 at 2:53 AM
DeepSeek's FlashMLA - their efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences and now in production.

✅ BF16 support
✅ Paged KV cache (block size 64)
⚡ 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800

github.com/deepseek-ai/...
GitHub - deepseek-ai/FlashMLA
Contribute to deepseek-ai/FlashMLA development by creating an account on GitHub.
github.com
February 24, 2025 at 2:46 AM
Summary of DeepSeek open source week

This is a fantastic consolidated guide. It goes deep, covers everything, and even has quizzes to test if you understood.

www.pyspur.dev/blog/deepsee...
DeepSeek's open-source week and why it's a big deal
Quick Intro to FlashMLA, DeepEP, DeepGEMM, DualPipe, EPPLB, 3FS and Smallpond
www.pyspur.dev
March 8, 2025 at 11:21 AM
sheesh! what a day for AI. QwQ-Max, sonnet-3.7 AND open source of FlashMLA
February 24, 2025 at 9:31 PM
ThunderMLA: a fused megakernel optimized for variable-prompt decoding! Inspired by DeepSeek's FlashMLA by @spectorb.bsky.social

Blog: hazyresearch.stanford.edu/blog/2025-03...
Code: github.com/HazyResearch...
ThunderMLA: FlashMLA, Faster and Fused-er!
hazyresearch.stanford.edu
March 6, 2025 at 4:09 AM
nice, the vllm project already integrated FlashMLA

win for open source

(graph from DeepSeek’s announcement)

github.com/vllm-project...
February 28, 2025 at 2:36 PM
DeepSeek Open Source FlashMLA – MLA Decoding Kernel for Hopper GPUs Discussion
GitHub - deepseek-ai/FlashMLA
Contribute to deepseek-ai/FlashMLA development by creating an account on GitHub.
github.com
February 24, 2025 at 3:00 AM
On the first day of Open Source Week, @deepseekofficial Deepseek open-sourced: FlashMLA, an efficient MLA decoding kernel optimized for Hopper GPUs, optimized for variable-length sequences, and now in production.

github.com/deepseek-ai/...
GitHub - deepseek-ai/FlashMLA
Contribute to deepseek-ai/FlashMLA development by creating an account on GitHub.
github.com
February 24, 2025 at 2:39 AM
Hot off the presses! github.com/deepseek-ai/...
GitHub - deepseek-ai/FlashMLA
Contribute to deepseek-ai/FlashMLA development by creating an account on GitHub.
github.com
February 24, 2025 at 1:44 AM
Yesterday, DeepSeek released flash-attn-3, sort of. It was actually FlashMLA.

Today, they released DeepEP - the first open-source EP communication library for MoE model training and inference.

✅ Efficient and optimized all-to-all communication
February 25, 2025 at 2:56 AM
DeepSeek Unveils FlashMLA, A Decoding Kernel That’s Make Things Blazingly Fast
DeepSeek Unveils FlashMLA, A Decoding Kernel That’s Make Things Blazingly Fast
cybersecuritynews.com
February 24, 2025 at 8:01 AM
📦 deepseek-ai / FlashMLA
⭐ 12,065 (+25)
🗒 C++

FlashMLA: Efficient Multi-head Latent Attention Kernels
GitHub - deepseek-ai/FlashMLA: FlashMLA: Efficient Multi-head Latent Attention Kernels
FlashMLA: Efficient Multi-head Latent Attention Kernels - deepseek-ai/FlashMLA
github.com
January 22, 2026 at 4:02 PM
FlashMLA

FlashMLA is an efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences serving.

https://github.com/deepseek-ai/FlashMLA
March 5, 2025 at 9:15 PM
DeepSeek Open Source FlashMLA – MLA Decoding Kernel for Hopper GPUs
https://github.com/deepseek-ai/FlashMLA

https://news.ycombinator.com/item?id=43155023
GitHub - deepseek-ai/FlashMLA
Contribute to deepseek-ai/FlashMLA development by creating an account on GitHub.
github.com
February 24, 2025 at 6:00 AM
Together AI introduced ThunderMLA a specialized decode kernel faster than DeepSeek's FlashMLA
hazyresearch.stanford.edu/blog/2025-0...
ThunderMLA: FlashMLA, Faster and Fused-er!
hazyresearch.stanford.edu
March 10, 2025 at 10:25 AM
Original post on medium.com
medium.com
March 1, 2025 at 11:02 AM
DeepSeek Open Source FlashMLA – MLA Decoding Kernel for Hopper GPUs (@github.com)

Main Link | Discussion
February 24, 2025 at 2:43 PM
DeepSeek's open-source week and why it's a big deal

DeepSeek recently open-sourced six powerful software libraries tackling some of the most complex problems in LLM training, inference, and data infrastructure. This is a very nice overview.
DeepSeek's open-source week and why it's a big deal
Quick Intro to FlashMLA, DeepEP, DeepGEMM, DualPipe, EPPLB, 3FS and Smallpond
buff.ly
March 8, 2025 at 7:51 PM
🚀 DeepSeek kicks off Open Source Week with FlashMLA, hitting 3.3k stars in 3 hours! 🎉

Optimized for Nvidia's Hopper GPU, it boosts LLM performance with 93.3% less memory usage. 🔥

Can small models soon gain reasoning abilities? 🤔 #AI #DeepSeek #OpenSource

aidisruption.ai/p/deepseek-r...
DeepSeek Releases FlashMLA, Boosting H800 GPU Performance
DeepSeek launches FlashMLA, an efficient decoding kernel for Nvidia's H800 GPU, boosting AI task performance and lowering training costs with MLA and MoE technologies.
aidisruption.ai
February 24, 2025 at 4:44 AM
Alibaba is attacking Nvidia’s real moat: CUDA. T-Head open-sourced SAIL for its Zhenwu AI chips, with PPU ports of Triton, DeepGEMM, FlashAttention and FlashMLA. Alternative accelerators now have to make migration boring.

github.com/t-head
July 19, 2026 at 11:02 PM
DeepSeek #НеделяОткрытогоИсточника День 1: Выпуск FlashMLA

Большая новость от DeepSeek! Компания официально запустила свой первый репозиторий с открытым исходным кодом, используя ядра CUDA для повышения скорости и эффективности больших языковых моделей (LLM). В основе этого обно…

#ai #deepseek #ml
DeepSeek #OpenSourceWeek Day 1: Release of FlashMLA
www.analyticsvidhya.com
March 7, 2025 at 7:28 PM