#VLLM
nano-vLLM

A lightweight vLLM implementation built from scratch, by one of the DeepSeek researcher.

🚀 Fast offline inference - Comparable inference speeds to vLLM
📖 Readable codebase - Clean implementation in ~ 1,200 lines of Python code

github.com/GeeeekExplor...
GitHub - GeeeekExplorer/nano-vllm: Nano vLLM
Nano vLLM. Contribute to GeeeekExplorer/nano-vllm development by creating an account on GitHub.
github.com
June 24, 2025 at 2:17 AM
knuckle tattoos that say FUCK VLLM
September 19, 2026 at 8:01 AM
vLLM now has a Rust frontend that lets you serve ~5x more HTTP requests in a single process

github.com/vllm-project...
[Frontend][RFC] Rust front-end integration by njhill · Pull Request #40848 · vllm-project/vllm
See corresponding RFC for introducing a rust-based alternative front-end process in vLLM #40846. For now we have staged the poc implementation in https://github.com/Inferact/vllm-frontend-rs. This ...
github.com
May 26, 2026 at 8:07 PM
Vllm est un projet de serveur d'inférence pour vos LLM qui se veut performant et efficace en terme de gestion mémoire ⬇️

github.com/vllm-project...
GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs - vllm-project/vllm
github.com
April 13, 2025 at 5:34 AM
vLLM breakdown blog post

this is an excellent breakdown of how vLLM works (think ollama but for legit production workloads)

great reading if you want a deeper understanding of inference

www.aleksagordic.com/blog/vllm
Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić
From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale.
www.aleksagordic.com
September 2, 2025 at 11:03 AM
Need blazing-fast classifier inference with minimal code?

ModernBERT now runs on vLLM — fast enough to process 200K+ arXiv papers in minutes.

It makes running any of the 100s of ModernBERT models on the @hf.co Hub even quicker.

Guide 👇
danielvanstrien.xyz/posts/2025/v...
Efficient Inference for ModernBERT Classifiers Using vLLM
Modern Inference for modern classifier models. Using vLLM to scale inference for classifiers to clean and curate datasets
danielvanstrien.xyz
April 24, 2025 at 3:11 PM
Inside vLLM: Anatomy of a High-Throughput LLM Inference System by Aleska Gordic

From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale

www.aleksagordic.com/blog/vllm
September 1, 2025 at 10:47 PM
vLLM has been updated to 0.71 with the enhancement to DeepSeek models.

github.com/vllm-project...
February 2, 2025 at 6:31 PM
oomfie you’re so plpilled but i’m talking about vLLM here
September 19, 2026 at 8:08 AM
LLM - Large Linux Manager
vLLM - volume Large Linux Manager
July 23, 2026 at 1:59 AM
If you're running an NVIDIA Spark or any of the clones with a GB10 (ASUS, Lenovo, Dell, HP, MSI, Gigabyte, etc) for local AI. I built a Docker image for it with the latest VLLM and all its component so that you don't have to.

No need to wait weeks for NVIDIA to do it.

github.com/timothystewa...
GitHub - timothystewart6/vllm-gb10: Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a). - timothystewart6/vllm-gb10
github.com
June 19, 2026 at 5:31 PM
Need to customize vLLM? Don't fork it. Use vLLM's plugin system to inject surgical modifications without maintaining a fork.

"Building Clean, Maintainable vLLM Modifications Using the Plugin System"

blog.vllm.ai/2025/11/20/v...
November 21, 2025 at 11:29 PM
Most inference engines make you choose between throughput and memory efficiency.

vLLM skips that tradeoff:
• OpenAI-compatible API
• Wide model support
• Built for speed on real hardware

Just added to OpenAlternative 👇
openalternative.co/vllm
vLLM: Open Source Alternative to Together AI and Fireworks AI
Inference and serving engine for large language models, built for speed and hardware efficiency with an OpenAI-compatible API and support for a wide range of…
openalternative.co
September 24, 2026 at 11:59 AM
We are excited to announce that vLLM inference provider is now available in Llama Stack!

Check out our recent article on vLLM blog that provides an introduction and tutorial to help you get started using it locally or deploying it in a Kubernetes cluster. blog.vllm.ai/2025/01/27/i...
Introducing vLLM Inference Provider in Llama Stack
We are excited to announce that vLLM inference provider is now available in Llama Stack through the collaboration between the Red Hat AI Engineering team and the Llama Stack team from Meta. This artic...
blog.vllm.ai
January 28, 2025 at 1:42 AM
vllm もう Qwen3 サポートしてそう…?
github.com/vllm-project...
仕事がはやい 仕事がはやい人たちだけで世界を回してもろて、おれはもうすべての労働を辞めたい……
[Usage] Qwen3 Usage Guide · Issue #17327 · vllm-project/vllm
vLLM v0.8.4 and higher natively supports all Qwen3 and Qwen3MoE models. Example command: vllm serve Qwen/... --enable-reasoning --reasoning-parser deepseek_r1 All models should work with the comman...
github.com
April 30, 2025 at 3:57 AM
Open source. Integrates with Hugging Face, Kubernetes, Envoy, Prometheus, OpenAI, Claude, Gemini and more. Playground at app.vllm-sr.ai. www.everydev.ai/tools/vllm-...
vLLM Semantic Router - LLM Request Routing Framework | EveryDev.ai
vLLM Semantic Router is an open-source, signal-driven routing framework for heterogeneous LLM inference, developed under the vllm-project…
www.everydev.ai
September 24, 2026 at 1:11 AM
Performance leap: TGI v3 is out. Processes 3x more tokens, 13x faster than vLLM on long prompts. Zero config !
December 10, 2024 at 10:08 AM
someone should make a version of vLLM that doesn't take 10 minutes to start
June 1, 2026 at 12:57 AM
you can go check how much hardware you need to server deepseek r1 or kimi k2 with vllm. it is documented. it is not that much hardware in purely financial terms
November 20, 2025 at 1:47 AM
vLLM CLI

A command-line interface tool for serving Large Language Models using vLLM. Provides both interactive and command-line modes with features for configuration profiles, model management, and server monitoring.

github.com/Chen-zexi/vl...
August 17, 2025 at 1:32 PM
If you want to learn vLLM under the hood, this is the post
www.aleksagordic.com/blog/vllm
Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić
From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale.
www.aleksagordic.com
October 4, 2025 at 4:08 PM
Continuous batching is the secret to why vLLM and transformers are fast.

"Continuous batching" by Remi Ouazan and two others.

huggingface.co/blog/continu...
November 26, 2025 at 12:52 AM
"you can install uv to the conda environment through pip"
vLLM docs who hurt you
November 13, 2025 at 12:50 PM
vLLM running natively on M3 and serving Qwen 🎉 This is huge for quickly spiking ideas without leaving your regular dev environment github.com/vllm-project...
December 15, 2024 at 12:32 PM