github.com/tijs/mei
github.com/tijs/mei
https://www.kunalganglani.com/blog/intel-arc-b-
https://www.kunalganglani.com/blog/intel-arc-b-
What stayed up: replica count and GPU-seconds per 1k requests. Neither one is an SLO.
What stayed up: replica count and GPU-seconds per 1k requests. Neither one is an SLO.
TTFT (Time to First Token) measures the time elapsed from a request to the first output token, not the overall speed of a large language model.
#hackernews #llm #news
TTFT (Time to First Token) measures the time elapsed from a request to the first output token, not the overall speed of a large language model.
#hackernews #llm #news
#AWS #AmazonEks
#AWS #AmazonEks
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application cha...
#AWS #AmazonEks
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application cha...
#AWS #AmazonEks
That said, pretty happy with it's caching, ha. Though I suppose it matters a lot less when I'm running it locally.
That said, pretty happy with it's caching, ha. Though I suppose it matters a lot less when I'm running it locally.
Haven't been able to hit 50 tps yet... but there's still some juice left to squeeze here.
Having to use a fork of Prism ML's llama.cpp (AMD) isn't helping either, as there are no speculative decoders available.