#TokenSpeed
now I see 40 tps vs 400 tps difference

mikeveerman.github.io/tokenspeed

#github #opensource #aitools
tokenspeed — feel LLM tokens-per-second
mikeveerman.github.io
May 29, 2026 at 8:03 AM
LightSeek's TokenSpeed, a speed-of-light LLM inference engine

- TensorRT LLM level performance
- vLLM level usability
- Built by a lean and mission-driven team in two months
- MIT license, open-source

Blog: lightseek.org/blog/lightse...
Repo: github.com/lightseekorg...
May 7, 2026 at 4:32 AM
necroing but i found this visualization for anyone not willing to spin up an actual LLM to get what the numbers mean mikeveerman.github.io/tokenspeed...
May 11, 2026 at 10:40 AM
47 tok/s, 200 tok/s, 800 tok/s — ces chiffres ne veulent rien dire tant qu'on ne les a pas vus défiler. Cet outil permet de visualiser différentes vitesses, en mode code ou prose. La différence de perception entre les deux est réelle.
https://mikeveerman.github.io/tokenspeed/
June 3, 2026 at 10:10 AM
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ
papoo.work
May 24, 2026 at 8:45 PM
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ
papoo.work
May 23, 2026 at 2:54 PM
When open source obsesses over LLM speed, the next AI breakthrough won’t be smarter—it’ll be instant.

https://github.com/lightseekorg/tokenspeed

https://www.projectnothing.ai/from-the-feed?ref=signal
May 7, 2026 at 7:01 PM
580 tokens a second on an open engine.

TokenSpeed is an LLM inference engine built for agent workloads. TensorRT-level speed, vLLM-level usability, KV cache reuse enforced at compile time.

how much of your agent bill is slow inference?

github.com/lightseekorg...
July 31, 2026 at 12:02 PM
Mainio saitti – visualisoi, miltä esim 10 tokenia sekunnissa näyttää ja auttaa muodostamaan kuvan, miten tehokkaan paikallisen mallin kanssa sovellus voi tulla toimeen.
mikeveerman.github.io/tokenspeed/?...
tokenspeed — feel LLM tokens-per-second
How fast is 10 tokens per second really?
mikeveerman.github.io
May 26, 2026 at 10:50 AM
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ
papoo.work
May 28, 2026 at 10:31 AM
LLMの「30 tokens/s」はどれくらい速いのかを体感できる「tokenspeed」の面白さ https://papoo.work/doc/4944bb87c9b6b1e9 #llm #tokens #benchmark #ai #developer-tools

Origin | Interest | Match
Awakari App
awakari.com
May 23, 2026 at 3:02 PM
The speed-of-light optimization for Qwen3.5 on the TokenSpeed inference engine is a significant milestone, achieving a record-breaking 580 tokens per second (tps) for agentic workloads on NVIDIA GPUs. Check out our latest community blog to learn more 👉 bit.ly/4uGUvIS
May 27, 2026 at 4:02 PM
Qwen3.8-27B hits 1M downloads in 2 days, tops Hugging Face trend. 27B dense VLM, 262K context extendable to 1M. Supports image/video, default reasoning mode with adjustable depth. Weights on Transformers, works with vLLM, SGLang, TokenSpeed, Apache 2.0.

#Qwen

aidisruption.ai/p/deepseek-h...
DeepSeek Harness Plugin Recommendations, Round 2
Discover 8 practical DeepSeek Harness plugins for workflow control, file annotation, browser automation, and custom tool creation. Boost your Agent efficiency today.
aidisruption.ai
August 18, 2026 at 7:42 AM
💡 Summary:

Qwen3.8-Maxは、Qwen3.5出自の大型オープンモデルを強化した2.4T規模の因果言語モデルで、長い文脈(最大262,144トークン、拡張で1,010,000まで)と高度な推論・計画能力を備え、コード・専門作業・長期タスクにおいて高性能を発揮します。API経由の利用を推奨しており、推論効率はフレームワークによって異なるため、SGLang/vLLM/TokenSpeedなどの専用サービングエンジンの利用が推奨されています。チャットAPIではthinkingを組み込んだ推論付き出力がデフォルトで有効で、設定例としてtemperatureやtop_p、 (1/2)
August 13, 2026 at 1:44 AM
tokenspeed: a speed-of-light llm inference engine for agentic workloads | lightseek foundation https://lightseek.org/blog/lightseek-tokenspeed.html
September 1, 2026 at 10:29 AM
PyTorch expands with 10 projects, boosting ML tooling.

PyTorch Ecosystem adds 10 projects, enhancing training, inference, and domain-specific tools like Perforated and RLinf. These innovations enable diverse, production-ready ML stacks, improving real-time data efficiency…

Read more on Kimbodo:
How PyTorch’s 10 New Projects Change Production ML — and What Data Science Teams Using Python and R Should Do Next
What Happened The PyTorch Ecosystem Landscape added ten projects that expand training, inference, routing, dataset, visualization and domain-specific tooling: Perforated, AReaL, TorchJD, RLinf, Miles, SMG, FiftyOne, TokenSpeed, VisualTorch, and TorchSurv.…
kimbodo.com
August 31, 2026 at 8:21 PM
💡 Summary:

Qwen3.8-27BはFP8量子化済みの27Bモデルと設定ファイルを提供し、Hugging Face TransformersやvLLM、SGLang、TokenSpeedなどに対応する。 Vision-Language対応の64層モデルで、長い文脈長(デフォルト262,144、最大1,000,000トークン)と強化された推論・思考モード、思考の保持機能を特徴とし、複数のインフェレンスフレームワークとAPIサービス(Qwen Cloud経由の提供予定)と連携する。使い方としては、APIを通じたチャットや画像・動画入力対応のマルチモーダル機能、 (1/2)
August 14, 2026 at 9:44 PM
阿里最近开源了千问旗舰模型Qwen3.8-2.4T-A95B,首次开放Max级别权重,总结如下:

一、技术特性与开源生态
1参数规模创新:
◦总参数2.4万亿(激活参数950亿/Token),原生支持26万Token上下文(可扩展至101万)
◦基于Qwen3.5架构升级,强化编程、科研及长周期Agent任务能力
2开源部署方案:
◦提供SGLang/vLLM/TokenSpeed推理引擎支持,需按GPU配置调整并行策略
◦Unsloth AI通过1-bit量化技术将模型体积从4.9TB压缩至397GB(降91%),410GB内存设备即可本地运行
August 14, 2026 at 10:50 AM
On NVIDIA B200 hardware, TokenSpeed outperformed TensorRT-LLM by 9% for minimum latency and 11% for throughput.
Its multi-head latent attention kernel nearly halves decoding latency on speculative decoding workloads.
Data Points: How Anthropic aligns its models
Hermes, aka the new OpenClaw. DCI, aka the new RAG. NLAs, or how to peek inside Claude. TokenSpeed, aka the new TensorRT.
www.deeplearning.ai
May 21, 2026 at 1:59 AM
6. TokenSpeed Inference Engine

The LightSeek Foundation released TokenSpeed, an inference engine licensed under the MIT License for coding agents that can handle contexts exceeding 50,000 tokens.
May 21, 2026 at 1:59 AM
Really cool tool to visualize what token/s really looks like at different speeds, starting from 0.05 tok/s to 2000 tok/s.
After 800 tok/s you can't really tell the difference it all feels the same!

mikeveerman.github.io/tokenspeed/?...
tokenspeed — feel LLM tokens-per-second
mikeveerman.github.io
June 11, 2026 at 11:44 PM
May 20, 2026 at 9:03 PM
How fast is N tokens per second really? | Discussion
tokenspeed — feel LLM tokens-per-second
mikeveerman.github.io
May 20, 2026 at 6:00 PM
tokenspeed-smg-grpc-proto 0.4.10.post20260621 SMG gRPC proto definitions for vLLM, TRT-LLM, MLX, TokenSpeed, and SGLang

Origin | Interest | Match
Client Challenge
pypi.org
June 21, 2026 at 2:02 AM