DeepSeek-V3 and similar mixture-of-experts models require high batch sizes to run efficiently due to their architecture. GPUs perform best with large matrix multiplications, so inference servers batch multiple user requests together. Models with many experts and layers need larger batches to avoid pipeline bubbles and keep all experts busy, which increases latency but dramatically improves throughput. This explains why DeepSeek is cost-effective at scale but inefficient for single-user local deployment.

•12m read time•From seangoedecke.com
Post cover image
Table of contents
What is batch inference?Why are some models tuned for high batch sizes?Summary
Share this post