MLA=Multihead Latent Attention
One of the big innovations that made V3 such a notable model
github.com/deepseek-ai/...
MLA=Multihead Latent Attention
One of the big innovations that made V3 such a notable model
github.com/deepseek-ai/...
¹ "Hops" is a reference to "Hopper"
² "Hopper" and "100" is a reference to the NVIDIA H100 GPU⁴
³ FlashMLA is an operator which is specially designed for NVIDIA Hopper GPUs
⁴ Can also work on other Hopper GPUs (but H100 is most prevalent)
¹ "Hops" is a reference to "Hopper"
² "Hopper" and "100" is a reference to the NVIDIA H100 GPU⁴
³ FlashMLA is an operator which is specially designed for NVIDIA Hopper GPUs
⁴ Can also work on other Hopper GPUs (but H100 is most prevalent)
✅ BF16 support
✅ Paged KV cache (block size 64)
⚡ 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800
github.com/deepseek-ai/...
✅ BF16 support
✅ Paged KV cache (block size 64)
⚡ 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800
github.com/deepseek-ai/...
This is a fantastic consolidated guide. It goes deep, covers everything, and even has quizzes to test if you understood.
www.pyspur.dev/blog/deepsee...
This is a fantastic consolidated guide. It goes deep, covers everything, and even has quizzes to test if you understood.
www.pyspur.dev/blog/deepsee...
Blog: hazyresearch.stanford.edu/blog/2025-03...
Code: github.com/HazyResearch...
Blog: hazyresearch.stanford.edu/blog/2025-03...
Code: github.com/HazyResearch...
win for open source
(graph from DeepSeek’s announcement)
github.com/vllm-project...
win for open source
(graph from DeepSeek’s announcement)
github.com/vllm-project...
github.com/deepseek-ai/...
github.com/deepseek-ai/...
Today, they released DeepEP - the first open-source EP communication library for MoE model training and inference.
✅ Efficient and optimized all-to-all communication
Today, they released DeepEP - the first open-source EP communication library for MoE model training and inference.
✅ Efficient and optimized all-to-all communication
FlashMLA is an efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences serving.
https://github.com/deepseek-ai/FlashMLA
FlashMLA is an efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences serving.
https://github.com/deepseek-ai/FlashMLA
https://github.com/deepseek-ai/FlashMLA
https://news.ycombinator.com/item?id=43155023
https://github.com/deepseek-ai/FlashMLA
https://news.ycombinator.com/item?id=43155023
hazyresearch.stanford.edu/blog/2025-0...
hazyresearch.stanford.edu/blog/2025-0...
https://medium.com/@pankaj_pandey/deepseek-flashmla-accelerating-transformer-decoding-on-nvidia-hopper-gpus-ddd6dfd82ba3?source=rss------machine_learning-5
#deepseek #python […]
https://medium.com/@pankaj_pandey/deepseek-flashmla-accelerating-transformer-decoding-on-nvidia-hopper-gpus-ddd6dfd82ba3?source=rss------machine_learning-5
#deepseek #python […]
Main Link | Discussion
Main Link | Discussion
DeepSeek recently open-sourced six powerful software libraries tackling some of the most complex problems in LLM training, inference, and data infrastructure. This is a very nice overview.
DeepSeek recently open-sourced six powerful software libraries tackling some of the most complex problems in LLM training, inference, and data infrastructure. This is a very nice overview.
Optimized for Nvidia's Hopper GPU, it boosts LLM performance with 93.3% less memory usage. 🔥
Can small models soon gain reasoning abilities? 🤔 #AI #DeepSeek #OpenSource
aidisruption.ai/p/deepseek-r...
Optimized for Nvidia's Hopper GPU, it boosts LLM performance with 93.3% less memory usage. 🔥
Can small models soon gain reasoning abilities? 🤔 #AI #DeepSeek #OpenSource
aidisruption.ai/p/deepseek-r...
github.com/t-head
github.com/t-head
Большая новость от DeepSeek! Компания официально запустила свой первый репозиторий с открытым исходным кодом, используя ядра CUDA для повышения скорости и эффективности больших языковых моделей (LLM). В основе этого обно…
#ai #deepseek #ml