#vLLM
Free and open-source (MIT). Supports 65+ LLM providers including Ollama, OpenAI, Anthropic, Groq and local inference via VLLM and Llama.cpp. MCP server support and full docs on the listing. www.everydev.ai/tools/code-...
Code Puppy - CLI AI Code Generation Agent | EveryDev.ai
Code Puppy is an open-source, MIT-licensed AI-powered code generation agent built by Mike Pfaffenberger and installable via `uvx code-puppy`. It was…
www.everydev.ai
October 1, 2026 at 3:33 PM
🦀 Ressources vLLM pour configs bi-GPU Nvidia partagées sur GitHub

Un utilisateur de r/LocalLLaMA a publié sur GitHub un dépôt de ressources vLLM destiné aux configurations à deux GPU Nvidia. Il explique avoir voulu pousser son PC le plus loin possible et avoir constaté que…
#IA #OpenClaw
Ressources vLLM pour configs bi-GPU Nvidia partagées sur GitHub
Un utilisateur de r/LocalLLaMA a publié sur GitHub un dépôt de ressources vLLM destiné aux configurations à deux GPU Nvidia. Il explique avoir voulu pousser son PC le plus loin possible et avoir const
communaute-ia.fr
October 1, 2026 at 1:55 PM
🦀 Un dépôt GitHub de ressources pour faire tourner vLLM sur deux GPU Nvidia

Un membre de r/LocalLLaMA a publié un dépôt GitHub, dual-gpus-vllm, réunissant des ressources pour faire tourner vLLM sur une configuration à deux GPU Nvidia. Il raconte avoir voulu pousser son PC…
#IA #OpenClaw
Un dépôt GitHub de ressources pour faire tourner vLLM sur deux GPU Nvidia
Un membre de r/LocalLLaMA a publié un dépôt GitHub, dual-gpus-vllm, réunissant des ressources pour faire tourner vLLM sur une configuration à deux GPU Nvidia. Il raconte avoir voulu pousser son PC et
communaute-ia.fr
October 1, 2026 at 1:50 PM
9. FRAC Architecture Uses Fractional Dynamics for Long-Sequence SSMs LINK
10. vLLM Adds Support for Pipeline Parallelism with PCP LINK
October 1, 2026 at 1:00 PM
Alibaba's Qwen3.8-2.4T-A95B model is now open weights, ready for the most demanding tasks. Deploy on Amazon SageMaker HyperPod using vLLM for high performance and control. #Alibaba #Qwen38 #AmazonSageMaker #vLLM
Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM | Amazon Web Services
Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.
aws.amazon.com
October 1, 2026 at 12:39 PM
vLLM released vllm-proto 0.4.0.

Useful for teams tracking the latest vLLM protocol release.

https://github.com/vllm-project/vllm/releases/tag/proto-v0.4.0
October 1, 2026 at 12:23 PM
New arXiv work introduces DLFP, a model-free vLLM controller that resizes prefill chunks based on observed decode-latency feedback to cut interference during concurrent inference on a single A100. Strong Qwen3-0.6B BF16 results, though…

#AI #LLMInference #vLLM #GPU
https://arxiv.org/abs/2609.38386
October 1, 2026 at 10:01 AM
One dude called Ash developed an engine that beats vllm and sglang.

Pretty crazy if you think about it.

— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2105578941293465709)
October 1, 2026 at 8:55 AM
AI-discovered flaws are getting weaponized fast. Google Threat Intelligence Group says disclosures doubled in 2026, yet real-world exploitation stayed rare. Attackers are also targeting AI tools like Flowise and vLLM. #BeyondTrust #AI Security #vLLM
The Vulnerabilities AI Finds Are The Ones Attackers Want
Google Threat Intelligence Group found that vulnerability disclosures doubled in 2026 while real-world exploitation remained rare, yet attackers moved quickly to weaponize newly disclosed flaws. The report highlights rapid abuse of an AI-discovered issue in BeyondTrust Privileged Remote Access and Remote Support, along with growing targeting of AI software such as Flowise, Langflow, vLLM, Ollama, and LiteLLM. #CVE-2026-1731 #BeyondTrust #SNOWLIGHT #SPARKRAT #LiteLLM #Langflow #Flowise #vLLM #Ollama
www.hendryadrian.com
October 1, 2026 at 7:00 AM
Fully open source (Apache 2.0). Covers PyTorch, Hugging Face, vLLM, LangChain and more. Link on the listing. www.everydev.ai/tools/llm-i...
LLM Internals - Open Source LLM Learning Resource | EveryDev.ai
LLM Internals is a free, open-source educational resource maintained by Amit Shekhar, founder of Outcome School and IIT alumnus (2010–14). Published…
www.everydev.ai
October 1, 2026 at 12:16 AM
🚨 EUVD-2026-90191
📊 6.9/10
🏢 vllm-project

📝 A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of...

🔗 https://euvd.enisa.europa.eu/vulnerability/EUVD-2026-90191

#cybersecurity #infosec #cve #euvd
September 30, 2026 at 7:02 PM
"Could you not pollute your code with this feature" makes LOTS of sense, especially when the feature is controversial and backed by all the companies obsessed with stealing and selling our data.

Especially when this code pushes a service that gets your data unencrypted, defeating the point of E2EE
September 30, 2026 at 6:46 PM
✍️ New blog post by xbill

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

#aws #sagemaker #gemma #vllm
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.
dev.to
September 30, 2026 at 6:14 PM
Multi-signal routing buys unlimited flexibility — and bills unlimited maintenance. LinkedIn's vLLM grid routes 50+ use cases; humans answer the pager, reading dashboards. The router is not free just because the code is. #openai #ai
https://seldondance.substack.com/p/production-ready-llm-routing
September 30, 2026 at 5:40 PM
✍️ New blog post by xbill

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

#aws #sagemaker #gemma #vllm
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.
dev.to
September 30, 2026 at 5:34 PM
今日のHuggingFaceトレンド

Qwen/Qwen3.8-27B
Qwen3.8-27Bというポストトレーニング済みモデルの重みと設定ファイルを提供するためのリポジトリです。
画像や動画を理解可能な視覚言語モデルであり、コーディングや専門業務、研究、エージェントタスクにおける能力向上を実現しています。
Hugging Face TransformersやvLLMなどに互換性があり、複雑な多段階タスクを信頼性高く遂行できる、デプロイに適したコンパクトな高密度モデルです。
Qwen/Qwen3.8-27B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
September 30, 2026 at 4:27 PM
今日のHuggingFaceトレンド

Edge0/Audio8-ASR-Infinite
応答性に優れたネイティブストリーミング音声認識モデルです。
最適化されたvLLMとRolling KV Cacheにより、時間の経過によるズレなく、24時間365日、無制限の長さの音声を書き起こし続けることができます。
Edge0/Audio8-ASR-Infinite · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
September 30, 2026 at 4:20 PM
Spark X2.5: on-device agentic AI with native 1M token context. No chunking, no cloud dependency, no lost context. Runs on vLLM, Ollama, llama.cpp. Apache 2.0.

Full deployment guide: aiadoptionagency.com/spark-x2-5-o...
Spark X2.5: On-Device Agentic AI with Million Token Context
Discover how Spark X2.5 delivers native 1M token context, agentic AI workflows, and coding assistance on-device. Enterprise deployment guide with implementation strategies.
aiadoptionagency.com
September 30, 2026 at 4:12 PM
Stack on-prem yang beneran 100% lokal hampir selalu bocor di tempat yang sama: bukan di inference-nya. Llama.cpp sama vllm aman. Yang bocor parser pdf sebelum chunking, ocr cloud, judge model yang dihosting orang lain, sama tracing saas. Satu http call dan air-gap-nya hangus. Modelnya gampang, supp…
September 30, 2026 at 4:06 PM
One GPU. 30 People. What Runs Out First? (vLLM)
YouTube video by Cloud Codes
youtu.be
September 30, 2026 at 2:22 PM
Engine inference baru yang ngebut ini pada dasarnya overfit ke satu kombinasi model + hardware. Menang benchmark, tapi 6 bulan lagi kelihatan lagi siapa. Yang general malah untung, trick bagus dari si special ini lama-lama diserap ke llama.cpp sama vllm. Specialisasi jadi r&d gratis buat engine gen…
September 30, 2026 at 1:58 PM
Excited to join forces with
@ashxhart
and
@volatilemarkts
to push local AI forward.

IMO TensorFold could become the "iPhone moment" for local AI. It beats both vLLM and SGLang on what matters the most: speed.

Alongside A…

— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2105280344857399797)
September 30, 2026 at 1:00 PM
✍️ New blog post by xbill

Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

#aws #sagemaker #gemma #vllm
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint: int4 linears, int4 embeddings and lm_head, FP8 and int8, across E2B, E4B, 12B, 26B A4B and 31B.
dev.to
September 30, 2026 at 12:49 PM