#DeepSpeed
DeepSpeed-Domino just came out, a communication-free LLM training engine with supporting tensor parallel. We are very glad to be part of this exciting project with DeepSpeed team.

Code: github.com/microsoft/De...
Paper: arxiv.org/abs/2409.15241
DeepSpeed/blogs/deepspeed-domino at master · microsoft/DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective. - microsoft/DeepSpeed
github.com
November 25, 2024 at 9:05 PM
"DeepSpeed" is a palindrome.
May 19, 2025 at 4:28 AM
Not enough mfs here wanting to talk about the up and downsides of using accelerate+trl+deepspeed, a minimal trainer on FSDP2, or pytorch lightning.
August 24, 2025 at 9:53 PM
DeepSpeed Ulysses-Offload

- Unlock the power of long context LLM training and finetuning with our latest system optimizations
- Train LLaMA3-8B on 2M tokens context using 4xA100-80GB
- Achieve over 55% MFU

Blog: github.com/microsoft/De...
December 6, 2024 at 6:43 AM
PyTorch Foundation Welcomes vLLM and DeepSpeed as Hosted Projects: thenewstack.io/pytorch-foun... via @thenewstack.io & @sjvn.bsky.social

The PyTorch Foundation wants to be the home for all #opensource #AI software programs.
PyTorch Foundation Welcomes vLLM and DeepSpeed as Hosted Projects
The PyTorch Foundation wants to be the home for all manner of open-source AI projects.
thenewstack.io
May 7, 2025 at 3:51 PM
Should've bit the upgrade for deepspeed bullet faster; using it as an accelerate plugin actually works pretty well too :)
trying to decide how much work I do building out fsdp/deepspeed for my rtx 4090 so I can do training with efficient cpu offloading versus just using my colab or buying/renting more compute.
August 17, 2025 at 4:08 PM
if you don't follow me on twitter where i've been liveblogging this saga basically i got fed up with deepspeed (an existing framework for doing this) after trying to get it working for two days, gave up this afternoon and just wrote it myself
November 22, 2024 at 7:04 AM
The thing to know about transformers+accelerate+trl+peft+deepspeed is the first 4'll be a bit broken, then you'll flip on deepspeed, & everything will explode.

Just finished harmonizing 'em, and the return is training Pretty Big Model on my rtx4090, but makes me wish I'd just gone FSDP2 tbh.
August 26, 2025 at 7:12 AM
DeepSpeed 이름 진짜 잘 지은 것 같다. 딥 러닝에서 가속을 한다는 솔직한 이름이면서도 회문임..
November 15, 2023 at 4:10 PM
My newsletter is out!

This week's agenda:
🔹 Open Source of the Week - DeepSpeed
🔹New learning resources
🔹Book of the week - Geospatial Data Science Essentials by Milan Janosov

open.substack.com/pub/ramikris...

#ai #datascience
DeepSpeed, Hermes LLM Wiki, GeoAI Essentials | Issue 92
A weekly curated update on data science and engineering topics and resources.
open.substack.com
June 14, 2026 at 2:47 PM
A GRPO demo (gist) code by Will Brown.

It is running smoothly on Qwen-1.5B w/ longer max_completion_length + higher num_generations, but haven't gotten LoRA or grad checkpointing working on multi-gpu w/ deepspeed yet for it.

gist.github.com/willccbb/467...
GRPO Llama-1B
GRPO Llama-1B. GitHub Gist: instantly share code, notes, and snippets.
gist.github.com
January 29, 2025 at 6:43 AM
Generally we let them all interop since one is built on the other and built on the other (so an accelerate DeepSpeed or FSDP config will work always, up through TRL and axolotl even)
November 28, 2024 at 12:33 PM
"Why is it so hard to install PyTorch, or CUDA, or libraries like FlashAttention or DeepSpeed that build against PyTorch and CUDA?"
August 13, 2025 at 6:24 PM
New episode: #TypeChat, #DeepSpeed, #Entra & #Purview updates—and why we love a bit of improv in tech.
Start your week with practical insights and real-world commentary.

youtu.be/y2MeLt-o-D4

#CloudyWithAChance #MicrosoftCloud #Podcast #AI #Cybersecurity
TypeChat, DeepSpeed, Entra & Purview Updates and the Joys of Not Preparing | EP22
YouTube video by Cloudy with a Chance of Insights, The MSFT Podcast
youtu.be
October 20, 2025 at 12:07 PM
flash-attnもvllmもdeepspeedももはや自分でビルドしなくてよくなるらしいです。
February 24, 2026 at 6:05 AM
I'm looking for an intern!

If you are:
* Driven
* Love OSS
* Interested in distributed PyTorch training/FSDPv2/DeepSpeed

Come work with me!

Fully remote, more details to apply in the comments
November 26, 2024 at 4:01 PM
🚀 TRL 0.14 – Featuring GRPO! 🚀

TRL 0.14 brings *GRPO*, the RL algorithm behind 🐳 DeekSeek-R1 .

⚡ Blazing fast generation with vLLM integration.
📉 Optimized training with DeepSpeed ZeRO 1/2/3.
January 30, 2025 at 2:54 PM
DeepNVMe just got faster and more flexible:
✅ Gen5 NVMe support
✅ 20X faster model checkpointing
✅ Cost-efficient SGLang inference via ZeRO-Inference
✅ CPU-only pinned memory support

📘 pytorch.org/blog/deepnvm...
#PyTorch #DeepSpeed #AIInfrastructure
June 17, 2025 at 5:04 PM
Forced into it because torch needs >2.6 for like non-causal masking (???) but as a nice treat for twenty minutes of downloading and massive hazard, I get functional deepspeed. Yoooooo
August 16, 2025 at 8:37 PM
May 6, 2026 at 3:33 PM
📝 Summary:

Higgsfield is an open-source GPU orchestration framework that manages multi-node, fault-tolerant training for trillion-parameter models by coordinating resources, supporting deepspeed/PyTorch sharding, and integrating with GitHub for CI/CD, all aimed at streamlined distributed LLM (1/2)
September 19, 2026 at 1:17 PM
🤖 Microsoft Contributes DeepSpeed to LF AI & Data
lfaidata.foundation/blog/2025/02...

🌉 solo.io Contributes kgateway to CNCF
cncf.io/blog/2025/02...

💰 Kosli Joins FINOS
finos.org/blog/kosli-j...

🔒 How Public Sector Entities Can Improve Supply Chain Security
openssf.org/blog/2025/02...
February 11, 2025 at 4:19 PM
DeepSpeedはなぜ速いのか〜推論編〜
https://zenn.dev/yasu52/articles/2314c8a9f74072
DeepSpeedはなぜ速いのか〜推論編〜
zenn.dev
June 23, 2024 at 10:32 PM
Holy crap, I burnt like 8 hours this week banging my head on trying to fix things, but my big runs were blowing up b/c DeepSpeed 0.16.4 is super borked (fix is disable gradient checkpointing or downgrade to 0.15.0): github.com/deepspeedai/...
[BUG] OOM when train 70B models using deepspeed 0.16.4 · Issue #7116 · deepspeedai/DeepSpeed
We found that using OpenRLHF + DeepSpeed 0.15.0, SFT + Adam Offload can train a 70B model with 8 A100 70G + ZeRO3, whereas DeepSpeed 0.16.4 results in OOM. You can try the script https://github.com...
github.com
March 27, 2025 at 3:21 PM