Stanford research found ~62% of tokens sent to AI agents on every call is repeated content, yet most inference systems recompute KV vectors from scratch each time. Prefix caching helps but has a hard ceiling: any change to the cached prefix causes a full miss, breaking RAG multi-document queries, reordered documents, and growing conversation histories. LMCache is an open-source project (10k+ stars) that disaggregates cache management into a separate process, eliminating resource contention with inference. It uses shared GPU memory, zero-copy cross-GPU sharing, and parallel multi-tier loading (GPU, CPU, SSD, remote storage). On H200 GPUs with Qwen3-235B and 50 concurrent users, it delivers 14x faster time-to-first-token and 4x faster decoding. The companion research CacheBlend (EuroSys 2025 Best Paper) solves the prefix-matching limitation by selectively recomputing only the small fraction of tokens with cross-document attention, giving 2-4x faster multi-document processing with no quality loss. LMCache ships with Prometheus/OpenTelemetry, a Kubernetes operator, and fault-tolerant failover.

Table of contents
The pull request now fixes itself before anyone reads itRethinking KV caching for production inferenceQuestions this post answers
Why does prefix caching fail for multi-document RAG queries?
Prefix caching requires the cached portion to be an exact byte-for-byte prefix of the new request, so it fails when a query needs multiple documents cached independently. If document A and document B were each cached alone, combining them causes a cache miss because the second document's cached KV state was computed without awareness of the first, and the same problem occurs when document order changes or conversation history grows. Anyone architecting RAG pipelines can track caching techniques like these as they evolve on daily.dev.
How much faster is LMCache compared to in-process KV caching?
On H200 GPUs running the Qwen3-235B model with 50 concurrent users, LMCache delivers 14x faster time-to-first-token and 4x faster decoding compared to in-process caching, with startup time dropping from over 3 minutes to about 30 seconds. This comes from running cache management as a separate process so it never contends with inference for GPU resources. Teams evaluating inference infrastructure can follow benchmarks like this on daily.dev before choosing a caching approach.
What is CacheBlend and how does it speed up multi-document queries?
CacheBlend is a technique from the LMCache team that won the EuroSys 2025 Best Paper Award, addressing the problem of combining independently cached documents. It identifies the small fraction of tokens with strong cross-document attention connections and selectively recomputes only those, reusing everything else from independent caches, giving 2 to 4x faster processing for multi-document RAG queries without quality loss. Developers building RAG systems can keep up with techniques like CacheBlend through daily.dev.
Share this post