#ModelServing
Big thanks to everyone contributing code, reviews, and ideas — this integration is shaping up to be a game-changer for 𝗞𝘂𝗯𝗲𝗿𝗻𝗲𝘁𝗲𝘀-𝗻𝗮𝘁𝗶𝘃𝗲 𝗟𝗟𝗠 𝘀𝗲𝗿𝘃𝗶𝗻𝗴. Stay tuned for next release!

#KServe #llmd #GenerativeAI #MLOps #Kubernetes #ModelServing #AIInfrastructure
August 11, 2025 at 3:45 PM
This is a big step for the KServe community, and we’re excited about the road ahead in making cloud-native model serving more accessible and production-ready for everyone.

#KServe #CNCF #OpenSource #ModelServing #AI #MLOps #CloudNative @cncf.io @kubernetes.io @kubefloworg.bsky.social
September 9, 2025 at 8:58 PM
Productionizing Machine Learning Trading Models: From Research to Live
beefed.ai
April 18, 2026 at 6:02 PM
🚀 Exploring the latest #AI stack ! From #VerticalAgents to #ModelServing and #Storage, this stack covers the essential tools & frameworks shaping the future of #ArtificialIntelligence. 📊🤖
#AIAgents #MachineLearning #TechStack #AIInfrastructure #DataScience #MLTools
November 15, 2024 at 12:45 PM
Ускорьте обслуживание моделей машинного обучения с помощью FastAPI и кэширования Redis

Вы когда-нибудь ждали слишком долго, чтобы модель вернула прогнозы? Мы все были в этом месте. Машинные модели, особенно большие и сложные, могут быть болезненно медленными при…

#ai #machinelearning #modelserving
Accelerate Machine Learning Model Serving With FastAPI and Redis Caching
www.analyticsvidhya.com
June 12, 2025 at 6:43 PM
Accelerate Machine Learning Model Serving With FastAPI and Redis Caching

Ever waited too long for a model to return predictions? We have all been there. Machine learning models, especially the large, complex ones, can be painfully slow to serve in real time. Use…

#ai #machinelearning #modelserving
Accelerate Machine Learning Model Serving With FastAPI and Redis Caching
Ever waited too long for a model to return predictions? We have all been there. Machine learning models, especially the large, complex ones, can be painfully slow to serve in real time. Users, on the other hand, expect instant feedback. That’s where latency becomes a real problem. Technically speaking, one of the biggest problems is […]
www.analyticsvidhya.com
June 9, 2025 at 11:22 PM
Presentation: Scaling Large Language Model Serving Infrastructure at Meta

Ye (Charlotte) Qi overviews LLM serving infrastructure challenges: fitting & speed (Model Runners, KV cache, and distributed inference), production complexities (latency optimization and continuous …

#llm #meta #modelserving
Presentation: Scaling Large Language Model Serving Infrastructure at Meta
Ye (Charlotte) Qi overviews LLM serving infrastructure challenges: fitting & speed (Model Runners, KV cache, and distributed inference), production complexities (latency optimization and continuous evaluation), and effective scaling strategies (heterogeneous deployment and autoscaling). Learn key concepts for robust LLM deployment. By Ye Qi
www.infoq.com
May 30, 2025 at 2:01 PM
🎄 Happy Holidays! KServe v0.12 release candidate is available! Try it out!

https://github.com/kserve/kserve/releases/tag/v0.12.0-rc0

#KServe #kubernetes #MLOps #DevOps #CloudNative #Kubeflow #ModelServing #AI #MachineLearning @KnativeProject @LFAIDataFdn @CloudNativeFdn
Release v0.12.0-rc0 · kserve/kserve
What's Changed Make storage initializer image configurable by @yuzisun in #3145 chore: Add design doc template links to feature request template by @ckadner in #3155 Increase pytest workers for ko...
github.com
November 20, 2024 at 2:15 AM
Ray Serve Deep Learning Containersが拓くTorchServeワークロードの簡素化と最適化

Ray Serve DLCでTorchServeデプロイを効率化

#RayServe #TorchServe #DeepLearningContainers #MLOps #ModelServing
Ray Serve Deep Learning Containersが拓くTorchServeワークロードの簡素化と最適化
Ray Serve DLCでTorchServeデプロイを効率化
ai.warp-studio.com
September 9, 2026 at 6:11 PM
KServe remains the go-to pattern for model serving on Kubernetes with built-in autoscaling and canary rollouts. Serving inference like you serve web traffic is the core insight of cloud native ML
#KServe #ModelServing #Kubernetes
August 20, 2026 at 3:52 AM
Представление: Масштабирование инфраструктуры для обслуживания больших языковых моделей в Meta.

Е (Шарлотта) Ци рассматривает проблемы инфраструктуры для обслуживания больших языковых моделей (LLM): соответствие требованиям и скорость (Model Runners, KV-кэш и распределенн…

#llm #meta #modelserving
Presentation: Scaling Large Language Model Serving Infrastructure at Meta
www.infoq.com
June 2, 2025 at 10:08 PM
🔔 New chapters on model serving and workflow patterns of Distributed Machine Learning Patterns are now available!

👉 http://bit.ly/2RKv8Zo

#MachineLearning #Kubernetes #DistributedSystems #CloudComputing #DeepLearning #DataScience #DevOps #MLOps #CloudNative #ModelServing
Distributed Machine Learning Patterns
Practical patterns for scaling machine learning from your laptop to a distributed cluster.</b> Distributing machine learning systems allow developers to handle extremely large datasets across multiple clusters, take advantage of automation tools, and benefit from hardware accelerations. This book reveals best practice techniques and insider tips for tackling the challenges of scaling machine learning systems. In Distributed Machine Learning Patterns</i> you will learn how to: Apply distributed systems patterns to build scalable and reliable machine learning projects</li> Build ML pipelines with data ingestion, distributed training, model serving, and more</li> Automate ML tasks with Kubernetes, TensorFlow, Kubeflow, and Argo Workflows</li> Make trade-offs between different patterns and approaches</li> Manage and monitor machine learning workloads at scale</li> </ul> Inside Distributed Machine Learning Patterns</i> you’ll learn to apply established distributed systems patterns to machine learning projects—plus explore cutting-edge new patterns created specifically for machine learning. Firmly rooted in the real world, this book demonstrates how to apply patterns using examples based in TensorFlow, Kubernetes, Kubeflow, and Argo Workflows. Hands-on projects and clear, practical DevOps techniques let you easily launch, manage, and monitor cloud-native distributed machine learning pipelines.
bit.ly
November 20, 2024 at 2:02 AM
🆕 GPU vs CPU Inference: 5 Scenarios, Real Costs & Latency

#MLOps #MLEngineering #ModelServing
https://tildalice.io/gpu-vs-cpu-inference-cost-latency-benchmark/
April 23, 2026 at 3:06 PM
🆕 FastAPI Model Serving: 5 Steps to 50ms Inference

#MLOps #MLEngineering #FastAPI #modelserving
https://tildalice.io/fastapi-model-serving-5-steps-inference/
February 18, 2026 at 3:04 PM