Manticore Search's default keywords dictionary truncates tokens to 42 bytes after normalization, which breaks searching long values like SHA-256 hashes, message IDs, and long email addresses since different IDs sharing the same first 42 bytes become indistinguishable. The dict='keywords_32k' option, available since Manticore Search 27.1.1 (27.1.5+ recommended for migrations), raises the maximum token length to 32768 bytes, skipping oversized tokens with a warning instead of truncating them. The guide covers schema choices (string attribute vs string attribute indexed), exact vs full-text matching semantics, prefix/infix search with min_infix_len, blend_chars for emails, CALL KEYWORDS for inspecting tokenization, migration steps for RT and plain tables, why dict='crc' doesn't help, current limitations (no CALL SUGGEST, no percolate tables, no snippet highlighting, no full-text REGEX), and a warning against indexing secrets like API keys and tokens.

•10m read time•From manticoresearch.com
Post cover image
Table of contents
The problem with the regular dictionaryWhat keywords_32k changesWhen to use keywords_32kSearching by the complete tokenA full-text match is not the same as exact equalityHow to inspect tokenizationPrefix and substring searchEmail addresses and other values with separatorsHow to convert an existing tableWhy dict='crc' does not solve this problemCurrent limitationsDo not index secretsQuick checklistSummaryDocumentation

Questions this post answers

What is the maximum token length for Manticore Search's default keywords dictionary?

The default dict='keywords' dictionary in Manticore Search truncates tokens to 42 bytes after normalization, measured in bytes rather than characters, so UTF-8 characters can consume the limit faster. Tokens exceeding 42 bytes are truncated both when indexing documents and when processing search queries, which can cause two different long IDs sharing the same first 42 bytes to become indistinguishable. Anyone indexing hashes or IDs in Manticore Search can track dictionary limits and workarounds on daily.dev.

How do I search full SHA-256 hashes or long IDs in Manticore Search without truncation?

Use dict='keywords_32k', available starting with Manticore Search 27.1.1 (version 27.1.5+ recommended for converting existing tables), which raises the maximum normalized token length from 42 bytes to 32768 bytes. Tokens exceeding the new limit are skipped with a warning instead of truncated, and morphology is not applied to tokens over 42 bytes since they typically have no useful word stem. Developers building log or ID search on Manticore Search can follow updates like this on daily.dev.

Does dict='crc' in Manticore Search support long token search like keywords_32k does?

No, dict='crc' stores keyword checksums instead of original text but does not increase the allowed token length beyond the regular 42-byte limit. Only dict='keywords_32k' implements the exception to that limit; keywords and keywords_32k also store term text, which allows Manticore to expand prefix and infix wildcard queries against the dictionary, something crc cannot do. Teams choosing between Manticore dictionary types can compare trade-offs like this on daily.dev.

Share this post