mini-sglang
github.com/sgl-project/...
mini-sglang
github.com/sgl-project/...
i’d like to say, “oh, we’ll just use sglang everywhere”, but no. one model only supports vLLM, another only hangs with tokasaurus (what?)
i’d like to say, “oh, we’ll just use sglang everywhere”, but no. one model only supports vLLM, another only hangs with tokasaurus (what?)
📦 sgl-project / sglang
⭐ 8,749 (+130)
🗒 Python
SGLang is a fast serving framework for large language models and vision language models.
📦 sgl-project / sglang
⭐ 8,749 (+130)
🗒 Python
SGLang is a fast serving framework for large language models and vision language models.
that’s…usable
lmsys.org/blog/2025-07...
that’s…usable
lmsys.org/blog/2025-07...
huggingface.co/collections/...
✨ 753B - 1M context
✨ MIT license
✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M
✨ AIME 2026: 99.2 (beats GPT-5.5, Claude Opus 4.8)
✨ vLLM / SGLang / Transformers supported
huggingface.co/collections/...
✨ 753B - 1M context
✨ MIT license
✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M
✨ AIME 2026: 99.2 (beats GPT-5.5, Claude Opus 4.8)
✨ vLLM / SGLang / Transformers supported
🚀 Preview:
projects.localizethedocs.org/sglang-docs-...
🌐 Crowdin:
localizethedocs.crowdin.com/sglang-docs-...
🐙 GitHub:
github.com/localizethed...
#Crowdin #GitHub #Sphinx #SGLang #AI #LLM
🚀 Preview:
projects.localizethedocs.org/sglang-docs-...
🌐 Crowdin:
localizethedocs.crowdin.com/sglang-docs-...
🐙 GitHub:
github.com/localizethed...
#Crowdin #GitHub #Sphinx #SGLang #AI #LLM
github.com/zhaochenyang...
github.com/zhaochenyang...
- Runs new open models on launch day
- Up to 25x faster inference on GB300
- Also handles video and image generation
Explore it here:
osp.fyi/sglang
- Runs new open models on launch day
- Up to 25x faster inference on GB300
- Also handles video and image generation
Explore it here:
osp.fyi/sglang
— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2104097283054686547)
— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2104097283054686547)
Find the full technical breakdown in the comments below:
Find the full technical breakdown in the comments below:
Runs Moondream, Qwen 3.5, and Gemma 4 on NVIDIA H100. Photon outperformed vLLM and SGLang in every throughput test and brought every model online faster, via use of a single 'megakernel' that runs the entire inference on the GPU alone.
Runs Moondream, Qwen 3.5, and Gemma 4 on NVIDIA H100. Photon outperformed vLLM and SGLang in every throughput test and brought every model online faster, via use of a single 'megakernel' that runs the entire inference on the GPU alone.
高速な推論環境をローカルで構築して試したい
yug1224 starred ekzhang/openjev-sglang
https://github.com/ekzhang/openjev-sglang
高速な推論環境をローカルで構築して試したい
yug1224 starred ekzhang/openjev-sglang
https://github.com/ekzhang/openjev-sglang
flaneur2020.github.io/posts/2025-1...
flaneur2020.github.io/posts/2025-1...