vLLM announces day-0 production support for Kimi K3, Moonshot AI's 2.8-trillion-parameter Mixture-of-Experts model with a 1M-token context window and native vision. Key highlights include: DSpark speculative decoding achieving 370 tok/s (3.14× speedup) on 16 NVIDIA GB300 GPUs; a redesigned hybrid prefix caching system that manages both recurrent KDA state and paged KV blocks; prefill/decode disaggregation with NIXL for large-scale deployments; sequence parallelism with custom reduce-scatter/all-gather kernels 1.7×–4.5× faster than NCCL; and multiple fused CUDA/Triton kernels for KDA decode, attention residuals, and LatentMoE tail fusion. The post covers architecture adaptations, deployment tips, benchmark results (0.976 GSM8K, 0.939 GPQA-Diamond), and a roadmap including Decode Context Parallelism and RL training support. NVIDIA Hopper/Blackwell and AMD MI355X are supported at launch.

•22m read time•From vllm.ai
Post cover image
Table of contents
Quick startTL;DRKimi K3's architecture, and how vLLM serves itBuilt for productionPerformance optimizationsQuality and Performance BenchmarksImportant Deployment TipsKimi K3 vLLM FAQRoadmap and Future WorkQuick linksAcknowledgements

Questions this post answers

How much does DSpark speculative decoding speed up Kimi K3 inference in vLLM?

DSpark speculative decoding gives a 3.14x speedup on single-user requests, raising decode throughput from 118 tok/s to 370 tok/s on 16 NVIDIA GB300 GPUs. It uses a block-diffusion backbone with a low-rank Markov head and confidence head, achieving around 4.73 accepted tokens per step on coding tasks and 2.61 on high-entropy creative writing tasks. Teams tuning inference latency budgets can track speculative decoding gains like these via daily.dev.

How many GPUs are needed to serve Kimi K3 with vLLM?

At least one 8x NVIDIA B300 (or GB300) node is required, and 16x B200 GPUs are also supported. AMD MI355X GPUs with ROCm are supported at launch as well. Most production deployments run multi-node setups with expert and data parallelism connected over RDMA or NVLink. Engineers sizing hardware for massive MoE models can follow deployment guidance like this on daily.dev.

Is prefix caching enabled by default for Kimi K3 in vLLM?

No, prefix caching is not enabled by default for Kimi K3 while its hybrid-cache design continues to evolve, so the --enable-prefix-caching flag must be passed explicitly. It supports caching over both full-attention KV and recurrent Kimi Delta Attention (KDA) state. Anyone configuring inference flags for new model architectures can check details like this on daily.dev.

Share this post