- TensorRT LLM level performance
- vLLM level usability
- Built by a lean and mission-driven team in two months
- MIT license, open-source
Blog: lightseek.org/blog/lightse...
Repo: github.com/lightseekorg...
- TensorRT LLM level performance
- vLLM level usability
- Built by a lean and mission-driven team in two months
- MIT license, open-source
Blog: lightseek.org/blog/lightse...
Repo: github.com/lightseekorg...
https://mikeveerman.github.io/tokenspeed/
https://mikeveerman.github.io/tokenspeed/
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
https://github.com/lightseekorg/tokenspeed
https://www.projectnothing.ai/from-the-feed?ref=signal
https://github.com/lightseekorg/tokenspeed
https://www.projectnothing.ai/from-the-feed?ref=signal
TokenSpeed is an LLM inference engine built for agent workloads. TensorRT-level speed, vLLM-level usability, KV cache reuse enforced at compile time.
how much of your agent bill is slow inference?
github.com/lightseekorg...
TokenSpeed is an LLM inference engine built for agent workloads. TensorRT-level speed, vLLM-level usability, KV cache reuse enforced at compile time.
how much of your agent bill is slow inference?
github.com/lightseekorg...
mikeveerman.github.io/tokenspeed/?...
mikeveerman.github.io/tokenspeed/?...
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
https://papoo.work/doc/4944bb87c9b6b1e9
#llm #tokens #benchmark #ai #developer-tools
Origin | Interest | Match
#Qwen
aidisruption.ai/p/deepseek-h...
#Qwen
aidisruption.ai/p/deepseek-h...
Qwen3.8-Maxは、Qwen3.5出自の大型オープンモデルを強化した2.4T規模の因果言語モデルで、長い文脈(最大262,144トークン、拡張で1,010,000まで)と高度な推論・計画能力を備え、コード・専門作業・長期タスクにおいて高性能を発揮します。API経由の利用を推奨しており、推論効率はフレームワークによって異なるため、SGLang/vLLM/TokenSpeedなどの専用サービングエンジンの利用が推奨されています。チャットAPIではthinkingを組み込んだ推論付き出力がデフォルトで有効で、設定例としてtemperatureやtop_p、 (1/2)
Qwen3.8-Maxは、Qwen3.5出自の大型オープンモデルを強化した2.4T規模の因果言語モデルで、長い文脈(最大262,144トークン、拡張で1,010,000まで)と高度な推論・計画能力を備え、コード・専門作業・長期タスクにおいて高性能を発揮します。API経由の利用を推奨しており、推論効率はフレームワークによって異なるため、SGLang/vLLM/TokenSpeedなどの専用サービングエンジンの利用が推奨されています。チャットAPIではthinkingを組み込んだ推論付き出力がデフォルトで有効で、設定例としてtemperatureやtop_p、 (1/2)
PyTorch Ecosystem adds 10 projects, enhancing training, inference, and domain-specific tools like Perforated and RLinf. These innovations enable diverse, production-ready ML stacks, improving real-time data efficiency…
Read more on Kimbodo:
PyTorch Ecosystem adds 10 projects, enhancing training, inference, and domain-specific tools like Perforated and RLinf. These innovations enable diverse, production-ready ML stacks, improving real-time data efficiency…
Read more on Kimbodo:
Qwen3.8-27BはFP8量子化済みの27Bモデルと設定ファイルを提供し、Hugging Face TransformersやvLLM、SGLang、TokenSpeedなどに対応する。 Vision-Language対応の64層モデルで、長い文脈長(デフォルト262,144、最大1,000,000トークン)と強化された推論・思考モード、思考の保持機能を特徴とし、複数のインフェレンスフレームワークとAPIサービス(Qwen Cloud経由の提供予定)と連携する。使い方としては、APIを通じたチャットや画像・動画入力対応のマルチモーダル機能、 (1/2)
Qwen3.8-27BはFP8量子化済みの27Bモデルと設定ファイルを提供し、Hugging Face TransformersやvLLM、SGLang、TokenSpeedなどに対応する。 Vision-Language対応の64層モデルで、長い文脈長(デフォルト262,144、最大1,000,000トークン)と強化された推論・思考モード、思考の保持機能を特徴とし、複数のインフェレンスフレームワークとAPIサービス(Qwen Cloud経由の提供予定)と連携する。使い方としては、APIを通じたチャットや画像・動画入力対応のマルチモーダル機能、 (1/2)
一、技术特性与开源生态
1参数规模创新:
◦总参数2.4万亿(激活参数950亿/Token),原生支持26万Token上下文(可扩展至101万)
◦基于Qwen3.5架构升级,强化编程、科研及长周期Agent任务能力
2开源部署方案:
◦提供SGLang/vLLM/TokenSpeed推理引擎支持,需按GPU配置调整并行策略
◦Unsloth AI通过1-bit量化技术将模型体积从4.9TB压缩至397GB(降91%),410GB内存设备即可本地运行
一、技术特性与开源生态
1参数规模创新:
◦总参数2.4万亿(激活参数950亿/Token),原生支持26万Token上下文(可扩展至101万)
◦基于Qwen3.5架构升级,强化编程、科研及长周期Agent任务能力
2开源部署方案:
◦提供SGLang/vLLM/TokenSpeed推理引擎支持,需按GPU配置调整并行策略
◦Unsloth AI通过1-bit量化技术将模型体积从4.9TB压缩至397GB(降91%),410GB内存设备即可本地运行
Its multi-head latent attention kernel nearly halves decoding latency on speculative decoding workloads.
Its multi-head latent attention kernel nearly halves decoding latency on speculative decoding workloads.
The LightSeek Foundation released TokenSpeed, an inference engine licensed under the MIT License for coding agents that can handle contexts exceeding 50,000 tokens.
The LightSeek Foundation released TokenSpeed, an inference engine licensed under the MIT License for coding agents that can handle contexts exceeding 50,000 tokens.
After 800 tok/s you can't really tell the difference it all feels the same!
mikeveerman.github.io/tokenspeed/?...
After 800 tok/s you can't really tell the difference it all feels the same!
mikeveerman.github.io/tokenspeed/?...
Origin | Interest | Match