Un utilisateur de r/LocalLLaMA a publié sur GitHub un dépôt de ressources vLLM destiné aux configurations à deux GPU Nvidia. Il explique avoir voulu pousser son PC le plus loin possible et avoir constaté que…
#IA #OpenClaw
Un utilisateur de r/LocalLLaMA a publié sur GitHub un dépôt de ressources vLLM destiné aux configurations à deux GPU Nvidia. Il explique avoir voulu pousser son PC le plus loin possible et avoir constaté que…
#IA #OpenClaw
Un membre de r/LocalLLaMA a publié un dépôt GitHub, dual-gpus-vllm, réunissant des ressources pour faire tourner vLLM sur une configuration à deux GPU Nvidia. Il raconte avoir voulu pousser son PC…
#IA #OpenClaw
Un membre de r/LocalLLaMA a publié un dépôt GitHub, dual-gpus-vllm, réunissant des ressources pour faire tourner vLLM sur une configuration à deux GPU Nvidia. Il raconte avoir voulu pousser son PC…
#IA #OpenClaw
Useful for teams tracking the latest vLLM protocol release.
https://github.com/vllm-project/vllm/releases/tag/proto-v0.4.0
Useful for teams tracking the latest vLLM protocol release.
https://github.com/vllm-project/vllm/releases/tag/proto-v0.4.0
#AI #LLMInference #vLLM #GPU
https://arxiv.org/abs/2609.38386
#AI #LLMInference #vLLM #GPU
https://arxiv.org/abs/2609.38386
Pretty crazy if you think about it.
— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2105578941293465709)
Pretty crazy if you think about it.
— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2105578941293465709)
forums.developer.nvidia.com/t/running-qw... 
forums.developer.nvidia.com/t/running-qw... 
📊 6.9/10
🏢 vllm-project
📝 A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of...
🔗 https://euvd.enisa.europa.eu/vulnerability/EUVD-2026-90191
#cybersecurity #infosec #cve #euvd
📊 6.9/10
🏢 vllm-project
📝 A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of...
🔗 https://euvd.enisa.europa.eu/vulnerability/EUVD-2026-90191
#cybersecurity #infosec #cve #euvd
Especially when this code pushes a service that gets your data unencrypted, defeating the point of E2EE
Especially when this code pushes a service that gets your data unencrypted, defeating the point of E2EE
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
#aws #sagemaker #gemma #vllm
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
#aws #sagemaker #gemma #vllm
https://seldondance.substack.com/p/production-ready-llm-routing
https://seldondance.substack.com/p/production-ready-llm-routing
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
#aws #sagemaker #gemma #vllm
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
#aws #sagemaker #gemma #vllm
Qwen/Qwen3.8-27B
Qwen3.8-27Bというポストトレーニング済みモデルの重みと設定ファイルを提供するためのリポジトリです。
画像や動画を理解可能な視覚言語モデルであり、コーディングや専門業務、研究、エージェントタスクにおける能力向上を実現しています。
Hugging Face TransformersやvLLMなどに互換性があり、複雑な多段階タスクを信頼性高く遂行できる、デプロイに適したコンパクトな高密度モデルです。
Qwen/Qwen3.8-27B
Qwen3.8-27Bというポストトレーニング済みモデルの重みと設定ファイルを提供するためのリポジトリです。
画像や動画を理解可能な視覚言語モデルであり、コーディングや専門業務、研究、エージェントタスクにおける能力向上を実現しています。
Hugging Face TransformersやvLLMなどに互換性があり、複雑な多段階タスクを信頼性高く遂行できる、デプロイに適したコンパクトな高密度モデルです。
Edge0/Audio8-ASR-Infinite
応答性に優れたネイティブストリーミング音声認識モデルです。
最適化されたvLLMとRolling KV Cacheにより、時間の経過によるズレなく、24時間365日、無制限の長さの音声を書き起こし続けることができます。
Edge0/Audio8-ASR-Infinite
応答性に優れたネイティブストリーミング音声認識モデルです。
最適化されたvLLMとRolling KV Cacheにより、時間の経過によるズレなく、24時間365日、無制限の長さの音声を書き起こし続けることができます。
Full deployment guide: aiadoptionagency.com/spark-x2-5-o...
Full deployment guide: aiadoptionagency.com/spark-x2-5-o...
@ashxhart
and
@volatilemarkts
to push local AI forward.
IMO TensorFold could become the "iPhone moment" for local AI. It beats both vLLM and SGLang on what matters the most: speed.
Alongside A…
— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2105280344857399797)
@ashxhart
and
@volatilemarkts
to push local AI forward.
IMO TensorFold could become the "iPhone moment" for local AI. It beats both vLLM and SGLang on what matters the most: speed.
Alongside A…
— from @MiaAI_lab (https://x.com/MiaAI_lab/status/2105280344857399797)
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
#aws #sagemaker #gemma #vllm
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
#aws #sagemaker #gemma #vllm