A lightweight vLLM implementation built from scratch, by one of the DeepSeek researcher.
🚀 Fast offline inference - Comparable inference speeds to vLLM
📖 Readable codebase - Clean implementation in ~ 1,200 lines of Python code
github.com/GeeeekExplor...
A lightweight vLLM implementation built from scratch, by one of the DeepSeek researcher.
🚀 Fast offline inference - Comparable inference speeds to vLLM
📖 Readable codebase - Clean implementation in ~ 1,200 lines of Python code
github.com/GeeeekExplor...
github.com/vllm-project...
github.com/vllm-project...
github.com/vllm-project...
github.com/vllm-project...
this is an excellent breakdown of how vLLM works (think ollama but for legit production workloads)
great reading if you want a deeper understanding of inference
www.aleksagordic.com/blog/vllm
this is an excellent breakdown of how vLLM works (think ollama but for legit production workloads)
great reading if you want a deeper understanding of inference
www.aleksagordic.com/blog/vllm
ModernBERT now runs on vLLM — fast enough to process 200K+ arXiv papers in minutes.
It makes running any of the 100s of ModernBERT models on the @hf.co Hub even quicker.
Guide 👇
danielvanstrien.xyz/posts/2025/v...
ModernBERT now runs on vLLM — fast enough to process 200K+ arXiv papers in minutes.
It makes running any of the 100s of ModernBERT models on the @hf.co Hub even quicker.
Guide 👇
danielvanstrien.xyz/posts/2025/v...
From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale
www.aleksagordic.com/blog/vllm
From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale
www.aleksagordic.com/blog/vllm
vLLM - volume Large Linux Manager
vLLM - volume Large Linux Manager
No need to wait weeks for NVIDIA to do it.
github.com/timothystewa...
No need to wait weeks for NVIDIA to do it.
github.com/timothystewa...
"Building Clean, Maintainable vLLM Modifications Using the Plugin System"
blog.vllm.ai/2025/11/20/v...
"Building Clean, Maintainable vLLM Modifications Using the Plugin System"
blog.vllm.ai/2025/11/20/v...
Check out our recent article on vLLM blog that provides an introduction and tutorial to help you get started using it locally or deploying it in a Kubernetes cluster. blog.vllm.ai/2025/01/27/i...
Check out our recent article on vLLM blog that provides an introduction and tutorial to help you get started using it locally or deploying it in a Kubernetes cluster. blog.vllm.ai/2025/01/27/i...
vLLM skips that tradeoff:
• OpenAI-compatible API
• Wide model support
• Built for speed on real hardware
Just added to OpenAlternative 👇
openalternative.co/vllm
vLLM skips that tradeoff:
• OpenAI-compatible API
• Wide model support
• Built for speed on real hardware
Just added to OpenAlternative 👇
openalternative.co/vllm
A command-line interface tool for serving Large Language Models using vLLM. Provides both interactive and command-line modes with features for configuration profiles, model management, and server monitoring.
github.com/Chen-zexi/vl...
A command-line interface tool for serving Large Language Models using vLLM. Provides both interactive and command-line modes with features for configuration profiles, model management, and server monitoring.
github.com/Chen-zexi/vl...
www.aleksagordic.com/blog/vllm
www.aleksagordic.com/blog/vllm
"Continuous batching" by Remi Ouazan and two others.
huggingface.co/blog/continu...
"Continuous batching" by Remi Ouazan and two others.
huggingface.co/blog/continu...
vLLM docs who hurt you
vLLM docs who hurt you