#kvcache
Moonshot AI's / Kimi's AgentENV (AENV) is a platform for running agent environments at scale, powering agentic RL training for Kimi K3.

github.com/kvcache-ai/A...
GitHub - kvcache-ai/AgentENV: AgentENV (AENV) is a distributed platform for running agent environments at scale.
AgentENV (AENV) is a distributed platform for running agent environments at scale. - kvcache-ai/AgentENV
github.com
July 27, 2026 at 5:57 PM
CVE-2026-96763 - kvcache-ai mooncake MountSegment Request Processing segment.cpp access control
CVE ID : CVE-2026-96763

Published : Sept. 24, 2026, 1:16 a.m. | 26 minutes ago

Description : A security flaw has been discovered in kvcache-ai mooncake up to 0.3.12/0.3.13.pos...
CVE-2026-96763 - kvcache-ai mooncake MountSegment Request Processing segment.cpp access control
A security flaw has been discovered in kvcache-ai mooncake up to 0.3.12/0.3.13.post1/0.3.14-rc1. This issue affects the function ScopedSegmentAccess::MountSegment of the file segment.cpp of the component MountSegment Request Processing. Performing a manipulation results in improper access controls. The attack is possible to be carried out remotely. The exploit has been released …
cvefeed.io
September 24, 2026 at 2:27 AM
The kvcache is inside the process itself, there is only one obvious continuation.
May 18, 2026 at 7:34 AM
This sounds like someone is dialing the KVcache quantization. You can reproduce this with open models by starting with Q8 and going more aggressive - the model does sound less sure and more forgetful.
September 22, 2026 at 4:01 PM
n8loom

A library for generating trees-of-thought - the kind used in MCTS, variants of majority-voting, etc. - efficiently by splitting the kvcache into fragments at each node, and dynamically concatenating the results together when generating.
February 4, 2025 at 8:12 AM
kvcache-ai/ktransformers
Check out kvcache-ai/ktransformers on GitHub
github.com
February 12, 2025 at 6:24 AM
🚨 EUVD-2026-85753
📊 5.3/10
🏢 kvcache-ai

📝 A weakness has been identified in kvcache-ai mooncake up to 0.3.12/0.3.14-rc1. Impacted is the function MasterService::GetReplicaListByRegex of the com...

🔗 https://euvd.enisa.europa.eu/vulnerability/EUVD-2026-85753

#cybersecurity #infosec #cve #euvd
September 24, 2026 at 1:01 AM
🚨 EUVD-2026-85751
📊 6.9/10
🏢 kvcache-ai

📝 A vulnerability was determined in kvcache-ai mooncake up to 0.3.12/0.3.13.post1. This affects the function UnmountSegment of the component RPC Path Han...

🔗 https://euvd.enisa.europa.eu/vulnerability/EUVD-2026-85751

#cybersecurity #infosec #cve #euvd
September 24, 2026 at 1:01 AM
🚨 EUVD-2026-85752
📊 5.3/10
🏢 kvcache-ai

📝 A security flaw has been discovered in kvcache-ai mooncake up to 0.3.12/0.3.13.post1/0.3.14-rc1. This issue affects the function ScopedSegmentAccess::M...

🔗 https://euvd.enisa.europa.eu/vulnerability/EUVD-2026-85752

#cybersecurity #infosec #cve #euvd
September 24, 2026 at 1:01 AM
CVE-2026-96764 - kvcache-ai mooncake Regular Expression GetReplicaListByRegex allocation of resources
CVE ID : CVE-2026-96764

Published : Sept. 24, 2026, 1:16 a.m. | 26 minutes ago

Description : A weakness has been identified in kvcache-ai mooncake up to 0.3.12/0.3.14-rc...
CVE-2026-96764 - kvcache-ai mooncake Regular Expression GetReplicaListByRegex allocation of resources
A weakness has been identified in kvcache-ai mooncake up to 0.3.12/0.3.14-rc1. Impacted is the function MasterService::GetReplicaListByRegex of the component Regular Expression Handler. Executing a manipulation can lead to allocation of resources. The attack may be performed from remote. The exploit has been made available to the public and could be …
cvefeed.io
September 24, 2026 at 2:21 AM
How Xiaomi decreased their API pricing by over 90%.

"They made Full-Pipeline Inference Optimization, where they have increased effective KVCache capacity by nearly 5x, with server-side cache hit rates averaging 93%–95% across mainstream harness frameworks."

mimo.xiaomi.com/blog/mimo-v2...
Xiaomi MiMo, Explore and Love
An end-to-end engineering practice for the MiMo-V2.5 series inference system, covering KVCache management, tiered caching, SWA-aware prefix cache trees, scheduling, Prefill/Decode pipelines, and multi...
mimo.xiaomi.com
May 30, 2026 at 4:53 PM
, which reduces KV cache size and makes cross-DC PD practical.

Result: Directly translating into lower token cost.

Paper: arxiv.org/abs/2604.150...
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
Prefill-decode (PD) disaggregation has become the standard architecture for large-scale LLM serving, but in practice its deployment boundary is still determined by KVCache transfer. In conventional de...
arxiv.org
April 18, 2026 at 1:10 PM
github.com/shisa-ai/Fas... - the first proper public (that I know of) Dynamic Memory Sparsification (DMS) arxiv.org/abs/2506.05345 - virtually lossly kvcache compression (up to 7.6x from my testing). The "Fast" part is that this implementation actually runs faster than vLLM w/ BF16/FP8 kvcache
GitHub - shisa-ai/FastDMS: Production-speed compact Dynamic Memory Sparsification (DMS) for KV cache compression
Production-speed compact Dynamic Memory Sparsification (DMS) for KV cache compression - shisa-ai/FastDMS
github.com
May 4, 2026 at 10:37 PM
In this new interview, our CEO & co-founder @JunchenJiang explains why KV cache — the internal memory of LLMs, is becoming the 𝗻𝗲𝘅𝘁 𝗕𝗶𝗴 𝗗𝗮𝘁𝗮 layer for AI, and how @tensormesh tackles large-scale inference.

🎥 Watch the full interview: youtu.be/zHW4Zzd7pjI

#LLMInference #KVCache #OpenSource #PyTorch
January 6, 2026 at 5:05 PM
アプリ側の知識足りないのがよくわかった。推論のKVCacheは勉強になった
September 26, 2025 at 2:31 PM
Modular: The Five Eras of KVCache
The Five Eras of KVCache
www.modular.com
February 7, 2026 at 12:56 AM
17時までに家に到着していたいので、MPLS Japan早めに離脱しました。お疲れ様でした。
遠隔KVcacheの話面白かったなー

トラポンは「ちょっと」高い
October 30, 2025 at 6:38 AM
⚡ 3.66 TiB/min throughput on GraySort benchmark in a 25-node cluster
⚡ 40+ GiB/s peak throughput per client node for KVCache lookup
🧬 Disaggregated architecture with strong consistency semantics
February 28, 2025 at 2:37 AM
Lesson probably learned today: maybe don‘t quantize your KVCache. Getting a shockingly sharp performance from a 4bit quant of Ornith-1.0-35B - and noticed that usual Q8_0 was absent… going to test more… maybe it’s just a happy day…
July 7, 2026 at 9:37 AM
so 60GB/s over the pcie then?

if the weights are off-loaded what's resident on the gpu, the kvcache?
April 24, 2026 at 7:23 PM
_technically_ you can have just one H100-class GPU with a whole bunch of RAM and CPU to do hybrid inference using something like ktransformers for a single user, but that still means we're 1 order of magnitude more expensive than what's doable for individuals instead of 2 or 3 orders of magnitude
GitHub - kvcache-ai/ktransformers: A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations - kvcache-ai/ktransformers
github.com
July 29, 2026 at 4:29 PM
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355, B200
AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?
newsletter.semianalysis.com
August 25, 2026 at 10:01 AM
kvcache-ai /ktransformers

kvcache-ai/ktransformers is a flexible framework designed to explore heterogeneous optimizations for LLM inference and fine‑tuning. It provides tools to experiment with various optimization strategies across different ha…
kvcache-ai /ktransformers
kvcache-ai/ktransformers is a flexible framework designed to explore heterogeneous optimizations for LLM inference and fine‑tuning. It provides tools to experiment with various optimization strategies
techyon.pages.dev
July 19, 2026 at 1:20 PM
Third option - models designed to use a segment of kvcache that is trained; so the model itself isn't modified, but it's got the equivalent of privileged context and then normal context. Iirc trained kv cache yielded surprisingly strong results
February 16, 2026 at 2:14 PM