14× faster embeddings: how we rebuilt the ONNX path in Manticore
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
Manticore Search 27.1.5 ships a new ONNX Runtime backend for auto-embeddings that delivers ~14× faster throughput than the previous SentenceTransformers/Candle path on CPU. The old path was stuck at 5–11 docs/sec regardless of concurrency or batch size; the new one reaches 70–233 docs/sec. Key engineering decisions: sharing a single ORT session across concurrent callers (safe on Linux/macOS per ORT's C API docs), processing one document per inference call instead of batching (padding overhead made batching slower with variable-length inputs), and disabling intra-op thread spinning to free CPU for the rest of the system. For maximum bulk ingest throughput, the recommended pattern is large batches (32–128 docs) from a single client thread, since ORT already parallelises internally. GPU support and Windows perf parity are planned for future releases.
Table of contents
TL;DRWhy this mattersWhy ONNX, and not CandleThe concurrency model — the part most readers will find newAdaptive parallelism — the wrong turns we tookNumbersWhat's nextTry itQuestions this post answers
How much faster is the new ONNX Runtime embedding backend in Manticore Search 27.1.5 compared to the old Candle path?
It is roughly 14x faster on average across the full threads x batch workload grid, on the same hardware, model, and weights. The old SentenceTransformers/Candle path stayed at 5-11 docs/sec regardless of configuration, while the new ONNX Runtime backend ranges from 70 to 233 docs/sec, with the advantage holding from 1 to 32 client threads. Teams choosing an embedding backend for search infrastructure can track real-world benchmarks like this one on daily.dev.
Why does batching documents together make ONNX inference slower instead of faster in Manticore's embedding pipeline?
Batching hurts because mixed-length text batches pad every row to the longest row's length, so the model does work proportional to batch_size times max_len times hidden_dim regardless of actual content, wasting cycles on padding tokens. Combined with ORT's intra-op thread pool spinning between dispatches and starving other work of CPU, one-document-per-call with spinning disabled outperformed batched inference in Manticore's tests. Developers debugging unexpected inference slowdowns can compare notes on ONNX batching pitfalls via daily.dev.
How should I configure client threads and batch size to get maximum embedding throughput when bulk inserting into Manticore Search with ONNX auto-embeddings?
Use a single client thread with a large batch size between 32 and 128 rather than many threads with small batches, since the ONNX backend already parallelizes internally. In benchmarks, 1 thread with batch size 64 hit 233 docs/sec, beating 8 threads at batch size 128 which reached only 147 docs/sec, because client-side fan-out just adds coordination overhead on top of ORT's own parallelism. Engineers tuning bulk ingest throughput for vector search can weigh configuration tradeoffs like this on daily.dev.
Share this post