#ttft
Mei was updated to 0.6.1, for the Ornith 1.5 model weighted avg TPS improved slightly, same for TTFT, and total task time was better. Qwen 3.6 only improved on task time and TTFT. Progress has slowed a bit, definitely harder to get gains now without making another metric worse.

github.com/tijs/mei
GitHub - tijs/mei: Native Swift/MLX OpenAI-compatible inference server for Apple Silicon
Native Swift/MLX OpenAI-compatible inference server for Apple Silicon - tijs/mei
github.com
September 29, 2026 at 5:49 AM
KDDI reduces application response latency by 38% with Google's automated evaluation framework, achieving a 18% improvement in TTFT. 🚀📈 #AI #PerformanceOptimization #KDDI
How KDDI Optimized RAG Performance with Agent Development Kit | Google Cloud Blog
KDDI's RAG app 'Buffmee' case study: Learn how they utilized an automated evaluation framework and Agent Development Kit to reduce total latency by 38% and successfully scale Gen AI performance.
cloud.google.com
September 28, 2026 at 10:40 PM
Intel Arc B‑series can run local LLMs well… if you stop expecting CUDA and start benchmarking reproducibly. I wrote a repeatable harness (TTFT, tok/s, VRAM, power), a Linux/Windows compatibility matrix, and Arc-friendly GGUF quant picks for 8K context.
https://www.kunalganglani.com/blog/intel-arc-b-
September 27, 2026 at 12:42 AM
Nothing went red because more prefill work per request looks like more load. The queue grows, the autoscaler adds replicas, the queue drains, p95 settles back under the line.

What stayed up: replica count and GPU-seconds per 1k requests. Neither one is an SLO.
September 25, 2026 at 2:34 PM
🆕 Devstral 2 2512 (Mistral) is now in the index — 1 provider, from $0.400/MTok input. Pricing, measured speed & capability checks. #LLM #AI
Devstral 2 2512 — pricing & measured performance
1 providers, from $0.400/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 25, 2026 at 5:31 AM
🆕 Mistral Large 3 2512 (Mistral) is now in the index — 2 providers, from $0.500/MTok input. Pricing, measured speed & capability checks. #LLM #AI
Mistral Large 3 2512 — pricing & measured performance
2 providers, from $0.500/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 25, 2026 at 5:31 AM
What Is TTFT? Understanding Time to First Token in LLMs

TTFT (Time to First Token) measures the time elapsed from a request to the first output token, not the overall speed of a large language model.
#hackernews #llm #news
What Is TTFT? Understanding Time to First Token in LLMs
TTFT (Time to First Token) measures the time elapsed from a request to the first output token, not the overall speed of a large language model.
hackernoon.com
September 24, 2026 at 11:54 PM
🆕 Amazon SageMaker HyperPod Inference Gateway cuts LLM inference latency by 82% and TTFT by 97-98% with real-time routing. It employs Envoy, Body-Based Router, and Endpoint Picker for GPU-aware load balancing. Available in all AWS regions with SageMaker HyperPod; cross-cluster and…

#AWS #AmazonEks
Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application changes. By replacing unintelligent round-robin load balancing with real-time inference-signal-driven routing, it reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios. The Gateway is built around 3 core components. The Envoy Endpoint terminates HTTPS traffic and exposes a single private endpoint per cluster. The Body-Based Router reads the model name directly from each incoming request and routes it to the correct GPU pool - enabling one gateway to serve many models from a single endpoint URL with no client-side changes required. The Endpoint Picker continuously scores every model server pod in real time across 6 inference-level signals - KV cache utilization, queue depth, LoRA adapter residency, prefix cache hit rate, predicted latency and running requests - selecting the optimal pod for each individual request. The gateway works with any OpenAI-compatible model server, including vLLM and SGLang, requiring no application code changes. Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Coming soon - cross-cluster and cross-region routing with a centralized fleet gateway, global rate limiting, and cost-tier-aware traffic shaping. To learn more, read the launch blog and explore the documentation
aws.amazon.com
September 24, 2026 at 10:10 PM
Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application changes. By replacing unintelligent round-robin load balancing with real-time inference-signal-driven routing, it reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios. The Gateway is built around 3 core components. The Envoy Endpoint terminates HTTPS traffic and exposes a single private endpoint per cluster. The Body-Based Router reads the model name directly from each incoming request and routes it to the correct GPU pool - enabling one gateway to serve many models from a single endpoint URL with no client-side changes required. The Endpoint Picker continuously scores every model server pod in real time across 6 inference-level signals - KV cache utilization, queue depth, LoRA adapter residency, prefix cache hit rate, predicted latency and running requests - selecting the optimal pod for each individual request. The gateway works with any OpenAI-compatible model server, including vLLM and SGLang, requiring no application code changes. Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Coming soon - cross-cluster and cross-region routing with a centralized fleet gateway, global rate limiting, and cost-tier-aware traffic shaping. To learn more, read the launch blog and explore the documentation
dlvr.it
September 24, 2026 at 10:07 PM
Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference

Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application cha...

#AWS #AmazonEks
Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application changes. By replacing unintelligent round-robin load balancing with real-time inference-signal-driven routing, it reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios. The Gateway is built around 3 core components. The Envoy Endpoint terminates HTTPS traffic and exposes a single private endpoint per cluster. The Body-Based Router reads the model name directly from each incoming request and routes it to the correct GPU pool - enabling one gateway to serve many models from a single endpoint URL with no client-side changes required. The Endpoint Picker continuously scores every model server pod in real time across 6 inference-level signals - KV cache utilization, queue depth, LoRA adapter residency, prefix cache hit rate, predicted latency and running requests - selecting the optimal pod for each individual request. The gateway works with any OpenAI-compatible model server, including vLLM and SGLang, requiring no application code changes. Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Coming soon - cross-cluster and cross-region routing with a centralized fleet gateway, global rate limiting, and cost-tier-aware traffic shaping. To learn more, read the https://aws.amazon.com/blogs/machine-learning/introducing-amazon-sagemaker-hyperpod-inference-gateway/ and explore the https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-model-deployment-inference-gateway.html
aws.amazon.com
September 24, 2026 at 10:05 PM
TTFT (Time to First Token) measures the time elapsed from a request to the first output token, not the overall speed of a large language model. #aiinfrastructure
What Is TTFT? Understanding Time to First Token in LLMs
hackernoon.com
September 24, 2026 at 5:35 AM
cheapestinference.com just increased prices. Maybe it still sounds like a good offer, but beware that performance is fluctuating wildly. While often OK, TTFT goes up to over a minute for extended periods of time.
Cheapest AI Inference — Unlimited LLM API, flat $22/mo
Cheapest AI inference pricing: unlimited Kimi K3, Qwen3.8 Max & more from $22/mo flat — no per-token billing. Works with Claude Code.
cheapestinference.com
September 23, 2026 at 8:06 PM
🆕 Qwen3.8 Omni Flash (Qwen) is now in the index — 1 provider, from $0.150/MTok input. Pricing, measured speed & capability checks. #LLM #AI
Qwen3.8 Omni Flash — pricing & measured performance
1 providers, from $0.150/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 23, 2026 at 5:41 PM
🆕 GPT-6 Sol Pro (OpenAI) is now in the index — 6 providers, from $2.00/MTok input. Pricing, measured speed & capability checks. #LLM #AI
GPT-6 Sol Pro — pricing & measured performance
6 providers, from $2.00/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 23, 2026 at 5:41 PM
🆕 GPT-6 Luna Pro (OpenAI) is now in the index — 6 providers, from $0.100/MTok input. Pricing, measured speed & capability checks. #LLM #AI
GPT-6 Luna Pro — pricing & measured performance
6 providers, from $0.100/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 23, 2026 at 3:29 PM
🆕 GPT-6 Sol (OpenAI) is now in the index — 7 providers, from $2.00/MTok input. Pricing, measured speed & capability checks. #LLM #AI
GPT-6 Sol — pricing & measured performance
7 providers, from $2.00/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 23, 2026 at 3:29 PM
🆕 GPT-6 Luna (OpenAI) is now in the index — 7 providers, from $0.100/MTok input. Pricing, measured speed & capability checks. #LLM #AI
GPT-6 Luna — pricing & measured performance
7 providers, from $0.100/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 23, 2026 at 5:34 AM
🆕 Claude Opus 5.5 (Anthropic) is now in the index — 11 providers, from $4.00/MTok input. Pricing, measured speed & capability checks. #LLM #AI
Claude Opus 5.5 — pricing & measured performance
11 providers, from $4.00/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 23, 2026 at 5:34 AM
🆕 Grok 4.7 (xAI) is now in the index — 4 providers, from $1.60/MTok input. Pricing, measured speed & capability checks. #LLM #AI
Grok 4.7 — pricing & measured performance
4 providers, from $1.60/MTok. Independently measured TTFT/throughput and capability checks.
modelindex.ai
September 22, 2026 at 5:30 AM
as usual this is from my local cloudflare node so latency is super down but i'm sure ttft is noooot
September 21, 2026 at 5:24 PM
The sub-500ms ttft I was seeing at first didn't hold under continual agentic workflows, and tps, while respectable at 31, just don't make up for how much thinking this model does.

That said, pretty happy with it's caching, ha. Though I suppose it matters a lot less when I'm running it locally.
September 21, 2026 at 1:17 AM
Bonsai effective bandwidth not far off Qwen now (97% CPU/85% GPU)
September 20, 2026 at 10:56 PM
Insight: For interactive apps, the initial token (TTFT) latency is key, not total gen time. Streaming is essential, not just a bonus. Example: In-game chat, delays impact real-time experience. #AIforApps #StreamingTech
September 20, 2026 at 10:00 PM
THAT SAID... with kv caching and ttft under 500ms, "real feel" absolutely is around 45-50 tps. It does not feel "slow" by any means.
Increased Bonzai 2 27b's tokens per second by 30% by enabling KV caching.

Haven't been able to hit 50 tps yet... but there's still some juice left to squeeze here.

Having to use a fork of Prism ML's llama.cpp (AMD) isn't helping either, as there are no speculative decoders available.
September 20, 2026 at 7:25 PM