#SGLang
xAI uses SGLang to inference their LLMs.
September 1, 2025 at 12:19 AM
To teach us how modern LLM inference works, LMSYS gave us a readable version of SGLang where they distilled SGLang from 300K into 5,000 lines. Kept the core design, cut the complexity. Without sacrificing performance.

mini-sglang

github.com/sgl-project/...
December 19, 2025 at 7:40 AM
been getting deep into LLM serving and the most frustrating part is that EVERY model has different preferences about everything

i’d like to say, “oh, we’ll just use sglang everywhere”, but no. one model only supports vLLM, another only hangs with tokasaurus (what?)
October 27, 2025 at 1:29 PM
🔥 Hot Repo! 🔥 (100+ new stars)

📦 sgl-project / sglang
⭐ 8,749 (+130)
🗒 Python

SGLang is a fast serving framework for large language models and vision language models.
GitHub - sgl-project/sglang: SGLang is a fast serving framework for large language models and vision language models.
SGLang is a fast serving framework for large language models and vision language models. - sgl-project/sglang
github.com
February 7, 2025 at 7:01 PM
Intel worked with sglang to get DeepSeek R1 671B running at 15 tok/sec on a Xeon

that’s…usable

lmsys.org/blog/2025-07...
Cost Effective Deployment of DeepSeek R1 with Intel® Xeon® 6 CPU on SGLang | LMSYS Org
<p>The impressive performance of DeepSeek R1 marked a rise of giant Mixture of Experts (MoE) models in Large Language Models (LLM). However, its massive mode...
lmsys.org
July 17, 2025 at 1:52 PM
CVE-2026-93838: SGLang through 0.5.20, unbounded memory alloc in handle_staging_req() via chunk_idx. ZMQ access leads to OOM crash. CVSS 5.9, unpatched. Patch now: https://www.valtersit.com/cve/CVE-2026-93838/ #CVE #infosec #SGLang
CVE-2026-93838 Vulnerability in SGLang
www.valtersit.com
September 27, 2026 at 10:00 PM
GLM 5.2 is here 🔥

huggingface.co/collections/...

✨ 753B - 1M context
✨ MIT license
✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M
✨ AIME 2026: 99.2 (beats GPT-5.5, Claude Opus 4.8)
✨ vLLM / SGLang / Transformers supported
June 16, 2026 at 6:42 PM
vLLM? sglang?
September 1, 2025 at 2:07 PM
i’ve never gotten GPT5 to think this long before
October 27, 2025 at 1:26 PM
RUNPOD: RUNPOD SERVERLESS: UP TO 4X 80GB GPUS PER WORKER, SGLANG QUICK DEPLOY ENDPOINT
September 28, 2026 at 5:45 AM
The SGLang community has reproduced Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning (paper: arxiv.org/abs/2503.09516 and repo: github.com/PeterGriffin...)

github.com/zhaochenyang...
github.com
June 1, 2025 at 2:38 PM
SGLang serves new AI models fast, with major speed gains on new hardware.

- Runs new open models on launch day
- Up to 25x faster inference on GB300
- Also handles video and image generation

Explore it here:
osp.fyi/sglang
September 26, 2026 at 9:30 PM
This isn't new. That blog post links to their 2023 paper that introduced the Outlines library, which has been supported by SGLang I think since pretty much the beginning of the project. vLLM also supports structured outputs via constrained decoding, just not with that particularly library.
Structured Outputs — SGLang
docs.sglang.io
February 7, 2026 at 12:10 AM
AMD Releases ROCm 6.3 with SGLang, Fortran Compiler, Multi-Node FFT, Vision Libraries, and More
www.techpowerup.com
November 26, 2024 at 11:48 AM
FYI I use sglang on my DeepSeek v4.1 Flash recipe

— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2104097283054686547)
September 27, 2026 at 6:48 AM
While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the SGLang and NVIDIA engineering teams has taken its production performance to the next level.

Find the full technical breakdown in the comments below:
June 23, 2026 at 4:00 PM
Moondream's Photon 2.0: Inference engine for Physical AI

Runs Moondream, Qwen 3.5, and Gemma 4 on NVIDIA H100. Photon outperformed vLLM and SGLang in every throughput test and brought every model online faster, via use of a single 'megakernel' that runs the entire inference on the GPU alone.
August 4, 2026 at 4:26 PM
As a followup to my earlier DeepSeek-v3 performance testing, here's what basically OOTB (2xH100, tp=16) vLLM (0.6.6.post2.dev5+g5ce4627a) vs SGLang (0.4.1.post4) looks like (this is concurrency=64, but scales similarly up to 1024). atm sglang has +125% better throughput and ~10X lower mean TTFT.
January 8, 2025 at 9:01 AM
オープンなLLMとSGLangを活用してJev互換のAPIを提供する仕組みらしい
高速な推論環境をローカルで構築して試したい

yug1224 starred ekzhang/openjev-sglang
https://github.com/ekzhang/openjev-sglang
GitHub - ekzhang/openjev-sglang: Jev-compatible API endpoint based on open models (prefill-only)
github.com
September 19, 2026 at 6:46 PM
Open source. Python, C++, JS and Swift APIs. Integrates with vLLM, SGLang, TensorRT-LLM and more on the listing. www.everydev.ai/tools/xgrammar
XGrammar - LLM Structured Generation Library | EveryDev.ai
XGrammar is an open-source library developed under the mlc-ai organization that brings fast, flexible structured generation to large language model…
www.everydev.ai
September 27, 2026 at 3:30 PM
Liquid AI just made their vision-language model up to 3x faster with zero quality loss. A tiny 280M drafter accelerates the 3B VLM via speculative decoding, shipping day-one for SGLang, llama.cpp, and MLX. https://huggingface.co/blog/liquidai/lfm2-5-vl-dspark
September 25, 2026 at 6:05 AM
SGLang, which originated as an open-source research project at Ion Stoica’s UC Berkeley lab, has raised capital from Accel.
Sources: project SGLang spins out as RadixArk with $400M valuation as inference market explodes | TechCrunch
SGLang, which originated as an open-source research project at Ion Stoica’s UC Berkeley lab, has raised capital from Accel.
techcrunch.com
January 21, 2026 at 11:27 PM
how are you deploying minimax? the model page has some recommendations (sglang/vllm/hf-xformers), claude has other opinions (ollama).
January 13, 2026 at 1:48 AM
modern LLM inference engines like vLLM & SGlang are becoming tough to dive into. to learn how these inference engines work, nano-vllm is a fantastic educational project—complete Page Attention & LLM scheduler in <1k loc.🤯
flaneur2020.github.io/posts/2025-1...
A Walkthrough of nano-vllm | Flaneur2020
Recently, I&rsquo;ve been delving into the architecture of production-grade inference engines. While projects like vLLM and SGLang are crazy sophisticated, …
flaneur2020.github.io
October 12, 2025 at 3:43 PM