Table of Contents

    tokenization - Search, Ranking & Data Pipelines

    SEARCH, RANKING & DATA PIPELINES

    Tokenization

    Learn how tokenization converts raw text into searchable units used by analyzers, inverted indexes, lexical ranking, autocomplete, vectorization, and language models. Understand sentence, word, subword, character, n-gram, keyword, path, and language-aware tokenization, along with positions, offsets, normalization, synonyms, index-time and query-time compatibility, multilingual text, technical identifiers, evaluation, and production monitoring.

    Introduction

    Search engines cannot place raw paragraphs directly into an inverted index. The text must first be converted into smaller, searchable units called tokens.

    Raw text:
    
    "Kafka consumer groups improve scalability."
    
    
    Possible tokens:
    
    kafka
    consumer
    groups
    improve
    scalability

    Tokenization is the process of dividing text into units that the remaining search and ranking pipeline can process.

    The selected tokenization strategy directly affects:

    • Which queries match a document
    • How terms are stored in an inverted index
    • Whether phrase and proximity searches work
    • How autocomplete and partial matching behave
    • How vocabulary size and index storage grow
    • How ranking models calculate term statistics
    • How multilingual and technical content is retrieved

    Core idea: Tokenization defines the searchable vocabulary of a lexical search system. If meaningful text is split incorrectly, later indexing and ranking stages cannot fully repair the lost structure.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Crawling and ingestion Raw content must be discovered, fetched, parsed, and cleaned before text analysis.
    2 Inverted indexes Tokens become terms stored in posting lists.
    3 Unicode text Search content can contain many scripts, symbols, combining characters, and writing systems.
    4 TF-IDF and BM25 Ranking models use statistics calculated from analyzed terms.
    5 Data pipelines Analyzer changes can require reprocessing the complete searchable corpus.
    6 Search evaluation Tokenization choices must be tested using representative queries and relevance judgments.
    7 Vector search Lexical tokenization and model tokenization serve related but different purposes.

    What Is a Token?

    A token is one occurrence of a searchable unit produced from text by an analyzer.

    Input:
    
    "Search systems rank documents."
    
    
    Tokens:
    
    search
    systems
    rank
    documents

    A search token can contain more than its visible text. Depending on the selected search library and index configuration, a token can also carry:

    • Start offset
    • End offset
    • Position
    • Position increment
    • Position length
    • Token type
    • Optional payload or metadata
    Text-analysis Pipeline
    parse content → normalize characters → tokenize text → filter or expand tokens → write terms and postings

    Parsing vs Tokenization vs Analysis

    Stage Responsibility Example
    Parsing Convert HTML, PDF, Word, or another format into structured plain text Extract title, headings, paragraphs, and tables from a PDF
    Character filtering Transform the character stream before tokenization Remove HTML tags while keeping readable text
    Tokenization Divide the character stream into tokens Split a sentence into words or subwords
    Token filtering Normalize, remove, replace, or expand created tokens Lowercase terms, stem words, or introduce synonyms
    Indexing Store final terms and occurrence metadata Add the term and document ID to a posting list

    Sentence Tokenization

    Sentence tokenization, also called sentence segmentation, divides text into sentences.

    Input:
    
    "Kafka stores events. Consumers process them."
    
    
    Sentences:
    
    1. Kafka stores events.
    
    2. Consumers process them.

    Sentence segmentation can support passage extraction, summarization, result snippets, proximity rules, chunking, and natural-language processing.

    A simple split on a full stop is unreliable because punctuation can occur in abbreviations, decimal numbers, URLs, initials, and technical identifiers.

    Word Tokenization

    Word tokenization divides text into word-like units.

    Input:
    
    "Data pipelines improve reliability"
    
    
    Word tokens:
    
    data
    pipelines
    improve
    reliability

    Whitespace tokenization works for simple demonstrations but does not fully handle punctuation, contractions, hyphenated expressions, URLs, email addresses, code identifiers, or languages without whitespace-separated words.

    Subword Tokenization

    Subword tokenization divides a word into smaller reusable pieces.

    Conceptual example:
    
    unbelievable
    
    
    Possible subword units:
    
    un
    believ
    able

    Subword tokenization can help represent:

    • Uncommon words
    • Newly created words
    • Word variations
    • Technical terms
    • Multilingual vocabulary

    Subword tokenization is widely associated with language-model input processing. Search analyzers and language-model tokenizers should not be assumed to produce the same units.

    Character Tokenization

    Character tokenization treats individual characters as units.

    Input:
    
    search
    
    
    Character tokens:
    
    s
    e
    a
    r
    c
    h

    Character-level processing avoids an unknown-word vocabulary problem but creates longer token sequences and loses explicit word boundaries unless the system models them separately.

    N-Gram Tokenization

    An n-gram tokenizer creates overlapping units containing \(n\) characters or terms.

    Character Trigrams

    Input:
    
    search
    
    
    3-character tokens:
    
    sea
    ear
    arc
    rch

    Character n-grams can support partial matching, autocomplete, spelling tolerance, and languages where segmentation is difficult.

    N-grams can greatly increase the number of indexed terms and posting-list entries. Minimum and maximum gram sizes must be selected carefully.

    Keyword Tokenization

    A keyword tokenizer treats the complete field value as one token.

    Input field:
    
    "Search, Ranking & Data Pipelines"
    
    
    One token:
    
    Search, Ranking & Data Pipelines

    Keyword-style processing is useful for:

    • Exact categories
    • Document IDs
    • Status values
    • Product codes
    • Email addresses
    • Controlled tags
    • Sorting, filtering, and faceting

    A field can have both an analyzed representation for full-text search and an exact representation for filtering.

    Path Tokenization

    Hierarchical paths can be tokenized into useful parent and child levels.

    Input:
    
    courses/system-design/search/tokenization
    
    
    Possible hierarchy tokens:
    
    courses
    
    courses/system-design
    
    courses/system-design/search
    
    courses/system-design/search/tokenization

    This can support category navigation, breadcrumbs, and hierarchical filters.

    Technical Tokenization

    Technical search content contains symbols and identifiers for which an ordinary natural-language tokenizer can be unsuitable.

    Examples:
    
    C++
    
    C#
    
    .NET
    
    D365 F&O
    
    group.id
    
    max.poll.records
    
    customerAccountId
    
    snake_case_name
    
    error-26026

    The analyzer must define whether punctuation, periods, plus signs, ampersands, underscores, hyphens, and camel-case boundaries should be preserved, split, or indexed in multiple forms.

    Multi-form Technical Tokens

    Input:
    
    customerAccountId
    
    
    Possible indexed forms:
    
    customerAccountId
    
    customer
    
    account
    
    id

    Preserving the original identifier supports exact matching, while split components improve recall for partial conceptual searches.

    Technical-content rule: Test analyzers using real identifiers, configuration properties, error codes, framework names, and programming syntax from the target corpus.

    Language-aware Tokenization

    Languages differ in word boundaries, morphology, compounds, punctuation, and writing systems.

    Language Situation Tokenization Concern
    Whitespace-separated language Whitespace helps but punctuation and morphology still require analysis
    Language without explicit word spaces Dictionary, morphological, or n-gram segmentation can be required
    Compound-rich language Long compound words can require decomposition
    Highly inflected language Many grammatical forms can represent one underlying concept
    Mixed-language content One document can require script detection and multiple analysis policies
    Right-to-left script Unicode, offsets, highlighting, and display direction require careful testing

    Unicode Normalization

    Visually similar strings can have different Unicode representations. Normalization converts selected equivalent representations into a consistent form before or during tokenization.

    Processing direction:
    
    Raw Unicode text
          |
          v
    Approved normalization
          |
          v
    Tokenizer
          |
          v
    Comparable token forms

    Unicode handling should be tested across the languages and scripts the search system supports. Destructive transliteration should not be applied blindly.

    Token Positions

    Positions record a token's logical location in the analyzed term stream.

    Text:
    
    "Kafka consumer groups scale processing"
    
    
    Positions:
    
    kafka      -> 1
    
    consumer   -> 2
    
    groups     -> 3
    
    scale      -> 4
    
    processing -> 5

    Positions enable phrase and proximity queries by recording the relationship between terms.

    Start and End Offsets

    Offsets identify the token's character range in the source field.

    Source text:
    
    "Kafka streams"
    
    
    Token:
    
    kafka
    
    
    Start offset:
    
    0
    
    
    End offset:
    
    5

    Offsets can support result highlighting and passage extraction by connecting an analyzed token back to the source text.

    Position Increments

    Token filters can place alternatives at the same logical position.

    Original token:
    
    dlq
    
    
    Alternative token at same position:
    
    dead-letter-queue

    Position-aware synonym handling allows alternatives without always making the expanded form appear as a later independent word.

    Character Filters

    Character filters transform text before the tokenizer runs.

    Possible uses include:

    • Removing HTML markup
    • Mapping selected characters
    • Normalizing known formatting patterns
    • Replacing domain-specific separators
    Input:
    
    "<p>Kafka &amp; RabbitMQ</p>"
    
    
    Character-filtered text:
    
    "Kafka & RabbitMQ"

    Offset correction is important when the transformed text has a different length from the original source.

    Token Filters

    Token filters process tokens after tokenization.

    Filter Purpose
    Lowercase filter Normalize case variants
    Stop-word filter Remove selected frequent low-information terms
    Stemmer Reduce related word forms to a common stem
    Lemmatizer Map inflected forms to a dictionary base form
    Synonym filter Add or replace approved equivalent terms
    Length filter Remove tokens outside configured length boundaries
    Unique-token filter Remove repeated tokens within a configured scope
    Edge n-gram filter Create token prefixes for search-as-you-type

    Stemming vs Lemmatization

    Area Stemming Lemmatization
    Method Uses rules to reduce words to stems Uses linguistic information to identify a base form
    Output Can produce a non-dictionary stem Usually aims for a valid base word
    Complexity Generally simpler Generally requires more linguistic processing
    Main risk Over-stemming can combine unrelated terms Incorrect linguistic assumptions can alter intended meaning

    Stop-word Handling

    Stop-word removal can reduce the number of indexed postings, but common words can still matter in phrases, names, legal language, code, and domain-specific terminology.

    Phrase:
    
    "to be or not to be"
    
    
    Aggressive stop-word removal:
    
    Important phrase structure can disappear.

    Stop-word policy should be evaluated with phrase queries and representative domain content rather than inherited blindly from a generic list.

    Synonyms

    Synonym processing connects different expressions used for the same approved concept.

    Possible synonym relationship:
    
    dlq
    
    dead-letter queue
    
    dead letter queue

    Synonyms can be applied during indexing, querying, or both. Index-time expansion increases indexed terms and can require reindexing when synonym rules change. Query-time expansion can be changed without rebuilding the complete index but increases query complexity.

    Index-time vs Query-time Tokenization

    Area Index Time Query Time
    Input Document field text User query text
    Purpose Create searchable index terms Create terms compatible with indexed terms
    Change impact Can require reindexing existing documents Can affect new searches immediately
    Typical expansion Stable normalization and controlled term generation Synonyms, spelling handling, or query-specific expansion
    Main requirement Equivalent document and query concepts must produce compatible searchable terms

    Compatibility rule: Index-time and query-time analyzers do not need to be identical, but their outputs must remain compatible with the intended retrieval behaviour.

    Tokenization and Inverted Indexes

    Document text
          |
          v
    Tokenizer produces terms
          |
          v
    Token filters normalize terms
          |
          v
    Inverted index stores:
    
    term -> document postings

    Tokenization determines the term dictionary. If D365 F&O is divided incorrectly, users searching for D365, F&O, or the complete product name can receive unexpected results.

    Tokenization and Ranking

    TF-IDF and BM25 operate on the terms produced by analysis. Tokenization therefore affects:

    • Term frequency
    • Document frequency
    • Document length
    • Phrase matches
    • Term proximity
    • Field relevance

    If one analyzer produces many small tokens, document length and term statistics can differ significantly from an analyzer producing fewer, larger tokens.

    Vocabulary Size

    Let \(V\) represent the set of unique tokens generated across the corpus. Vocabulary size is:

    \[ VocabularySize = |V| \]

    A tokenizer that creates many n-grams or preserves every surface variation can increase vocabulary size and index storage.

    Token Count

    Let \(T_d\) represent the token sequence of document \(d\). The analyzed document length is:

    \[ DocumentTokenCount(d) = |T_d| \]

    Ranking systems can use analyzed document length for normalization. Analyzer changes can therefore influence relevance even when the visible document text does not change.

    Autocomplete Tokenization

    Search-as-you-type commonly uses a controlled prefix strategy.

    Token:
    
    consumer
    
    
    Possible edge prefixes:
    
    con
    cons
    consu
    consum
    consume
    consumer

    Very short prefixes can match many terms and create large posting lists. Minimum prefix length, field selection, and query limits should be tested.

    Fuzzy and Typo-tolerant Retrieval

    Tokenization can be combined with edit-distance search, phonetic analysis, n-grams, or spelling correction.

    User query:
    
    tokeniztion
    
    
    Intended term:
    
    tokenization

    Typo tolerance should remain conservative for short terms, identifiers, error codes, and security-sensitive searches because broad matching can return unrelated results.

    Search Tokens vs Model Tokens

    Area Search Token Language-model Token
    Primary purpose Lexical indexing and retrieval Model input and output representation
    Typical form Word, normalized term, n-gram, or exact field value Model-specific word, subword, character, or byte-like unit
    Vocabulary Derived from analyzer and indexed corpus Defined by the model tokenizer
    Main concern Recall, precision, index size, and ranking statistics Context length, model compatibility, latency, and usage cost
    Interchangeability The two token types should not be treated as equivalent without verification

    Tokenization in the Data Pipeline

    Source document
          |
          v
    Parse and clean
          |
          v
    Detect language and field type
          |
          v
    Select analyzer
          |
          v
    Tokenize and filter
          |
          v
    Inspect token stream
          |
          v
    Build inverted-index postings
          |
          v
    Validate search behaviour

    Analyzer versions should be recorded with the indexed document or index schema. A change to tokenization can change the searchable representation and require reindexing.

    Analyzer Versioning

    {
      "documentId": "article-1042",
      "sourceVersion": 8,
      "analyzerVersion": "technical-search-v3",
      "language": "en",
      "indexedAt": "indexing-timestamp"
    }

    Versioning supports controlled comparison, rollback, and reindexing after an analyzer change.

    Conceptual Analyzer Configuration

    {
      "analyzer": "technical_english",
      "characterFilters": [
        "html_strip",
        "approved_symbol_mapping"
      ],
      "tokenizer": "technical_standard",
      "tokenFilters": [
        "lowercase",
        "approved_possessive_handling",
        "approved_stemming",
        "technical_synonyms"
      ]
    }

    This is a conceptual configuration. The exact syntax and available components depend on the selected search platform.

    Simple Educational Tokenizer

    import re
    from dataclasses import dataclass
    
    
    @dataclass(frozen=True)
    class Token:
        term: str
        position: int
        start_offset: int
        end_offset: int
    
    
    def tokenize(text: str) -> list[Token\]:
        tokens: list[Token] = []
    
        for position, match in enumerate(
            re.finditer(
                r"[A-Za-z0-9]+(?:[._+-][A-Za-z0-9]+)*",
                text,
            )
        ):
            tokens.append(
                Token(
                    term=match.group(0).lower(),
                    position=position,
                    start_offset=match.start(),
                    end_offset=match.end(),
                )
            )
    
        return tokens
    
    
    text = (
        "Kafka group.id and max.poll.records "
        "control consumer behavior."
    )
    
    for token in tokenize(text):
        print(token)

    This example is intended for learning only. It does not provide complete Unicode, multilingual, URL, email, symbol, code, synonym, sentence, or language-aware analysis.

    Analyzer Inspection

    Before deploying an analyzer, inspect its tokens for representative inputs.

    Input Type Example
    Ordinary sentence Consumer groups process partitions
    Product name Dynamics 365 Finance & Operations
    Programming language C++, C#, .NET
    Configuration property max.poll.records
    Identifier customerAccountId
    Error code ERR-26026
    URL Approved documentation URL
    Mixed language Supported multilingual phrase

    Evaluation

    Tokenizer quality cannot be determined only by inspecting a few token streams. Evaluate retrieval using representative documents, queries, and relevance judgments.

    Useful evaluation measures include:

    • Precision at \(k\)
    • Recall at \(k\)
    • Mean Reciprocal Rank
    • Normalized Discounted Cumulative Gain
    • Zero-result query rate
    • Query reformulation rate
    • Phrase-query success
    • Exact-identifier success

    Tokenization changes should be evaluated separately for natural-language queries, identifiers, names, codes, phrases, and autocomplete.

    Educational-platform Example

    Document title:
    
    "D365 F&O X++ Select Statements"
    
    
    Possible title tokens:
    
    d365
    f&o
    x++
    select
    statements
    
    
    Additional normalized aliases:
    
    dynamics
    365
    finance
    operations
    xpp

    The final analyzer should reflect the site's actual search behaviour. Exact product names should remain searchable while approved aliases improve recall.

    Observability

    Useful tokenization metrics include:

    • Tokens generated per document
    • Vocabulary size
    • Unique terms added per ingestion run
    • Average and maximum token length
    • Documents generating zero searchable tokens
    • Terms removed by stop-word or length filters
    • Synonym expansion count
    • N-gram term growth
    • Unknown or unsupported language count
    • Analyzer execution time
    • Index size after analyzer changes
    • Zero-result and low-recall query patterns

    Alert Conditions

    • Documents unexpectedly generate no tokens
    • Vocabulary size grows sharply after an analyzer change
    • Token counts increase beyond the expected range
    • Analyzer errors increase for one language or content type
    • Exact identifiers stop matching
    • Phrase-query quality decreases
    • Zero-result queries increase after reindexing
    • N-gram generation causes unexpected index growth
    • Index-time and query-time analyzer versions differ unexpectedly
    • Permission, metadata, or exact fields are accidentally analyzed

    Common Tokenization Mistakes

    1

    Splitting Only on Whitespace

    Punctuation, URLs, email addresses, hyphenation, and technical identifiers are handled incorrectly.

    2

    Using One Tokenizer for Every Field

    Titles, body text, product codes, categories, paths, and identifiers have different retrieval requirements.

    3

    Destroying Technical Symbols

    Terms such as C++, C#, .NET, F&O, and group.id become difficult to retrieve correctly.

    4

    Lowercasing Every Exact Identifier

    Case-sensitive codes or identifiers can lose their intended identity.

    5

    Removing Every Stop Word

    Phrase meaning and named expressions can be damaged.

    6

    Applying Aggressive Stemming

    Distinct words can collapse into one term, decreasing precision.

    7

    Generating Excessive N-Grams

    The vocabulary and posting lists grow rapidly, increasing storage and query cost.

    8

    Using Incompatible Index and Query Analysis

    Equivalent query and document text produce different terms and fail to match.

    9

    Ignoring Language Differences

    Whitespace assumptions and generic stemming fail for supported multilingual content.

    10

    Changing the Analyzer without Reindexing

    Existing documents retain terms created by the previous analyzer while new documents use the new policy.

    11

    Confusing Search Tokens with Model Tokens

    Search relevance and language-model context use different tokenization objectives and vocabularies.

    12

    Testing Only Ordinary English Sentences

    The analyzer appears correct but fails on real names, codes, technical properties, symbols, and multilingual content.

    Recommended Test Cases

    Test Expected Evidence
    Case variation Equivalent case variants match according to field policy.
    Unicode variation Approved equivalent Unicode forms generate compatible tokens.
    Phrase query Positions support the expected phrase match.
    Punctuation Meaningful punctuation is retained or normalized deliberately.
    Programming-language name C++, C#, and .NET remain searchable.
    Configuration property group.id and max.poll.records support exact and approved partial search.
    Camel-case identifier The original and approved component forms remain searchable.
    Stop-word phrase Phrase meaning remains correct.
    Synonym Approved variants retrieve the intended documents without broad noise.
    Autocomplete Configured prefixes match while short noisy prefixes remain controlled.
    Multilingual content The correct language-aware analyzer is applied.
    Analyzer upgrade Reindexing produces one consistent searchable representation.

    Tokenization Best Practices

    Recommended Practices

    • Define search requirements before selecting a tokenizer.
    • Use field-specific analyzers.
    • Preserve exact forms for identifiers, categories, tags, and codes.
    • Use analyzed forms for natural-language retrieval.
    • Test domain-specific punctuation and symbols.
    • Use Unicode-aware processing.
    • Use language-aware tokenizers for supported languages.
    • Preserve positions when phrase and proximity search are required.
    • Preserve offsets when highlighting is required.
    • Apply stop words and stemming conservatively.
    • Apply synonyms through controlled, versioned rules.
    • Keep index-time and query-time analysis compatible.
    • Limit n-gram ranges and prefix expansion.
    • Record analyzer versions.
    • Reindex after incompatible analyzer changes.
    • Inspect token streams before deployment.
    • Evaluate retrieval using representative queries.
    • Monitor vocabulary, token count, index growth, and zero-result queries.
    • Separate lexical search tokens from language-model token accounting.
    • Maintain a controlled analyzer rollback strategy.

    Practice Exercise

    Design tokenization for the technical articles on your educational platform.

    Requirements

    1. Create separate mappings for title, body, category, tags, and document ID.
    2. Use exact and analyzed representations where appropriate.
    3. Normalize ordinary case variants.
    4. Preserve C++, C#, .NET, D365 F&O, X++, and SQL identifiers.
    5. Support properties such as group.id and max.poll.records.
    6. Split camel-case identifiers while retaining the original value.
    7. Create an approved technical synonym list.
    8. Preserve positions for phrase search.
    9. Preserve offsets for highlighting.
    10. Add controlled autocomplete prefixes.
    11. Inspect tokens for at least one supported non-English language.
    12. Compare results before and after stemming.
    13. Measure vocabulary and index-size changes.
    14. Evaluate zero-result and phrase queries.
    15. Version the analyzer and test a complete reindex.

    Tokenization-design Template

    Decision Selected Direction Risk Controlled
    Natural-language text Language-aware word analysis Improves lexical recall and ranking statistics
    Technical identifiers Original token plus approved component tokens Preserves exact and partial retrieval
    Exact metadata Keyword-style tokenization Supports reliable filtering and faceting
    Phrase search Store token positions Supports ordered term matching
    Highlighting Store correct offsets Maps matches to source text
    Synonyms Controlled versioned expansion Improves recall without uncontrolled ambiguity
    Autocomplete Bounded edge prefixes Prevents excessive expansion and index growth
    Analyzer change Versioned reindex with rollback Maintains a consistent searchable corpus

    Frequently Asked Questions

    1

    What is tokenization?

    Tokenization divides text into smaller units such as sentences, words, subwords, characters, or n-grams.

    2

    What is a search token?

    A search token is an analyzed term occurrence that can include text, position, offsets, type, and optional metadata.

    3

    Why is tokenization important?

    It determines the searchable terms stored in the inverted index and therefore affects matching, ranking, phrases, recall, and index size.

    4

    Is splitting on spaces enough?

    No. Whitespace splitting does not fully handle punctuation, symbols, URLs, identifiers, or languages without explicit word spaces.

    5

    What is keyword tokenization?

    Keyword tokenization retains the complete field value as one token for exact matching, filtering, sorting, or faceting.

    6

    What are n-grams?

    N-grams are overlapping sequences of characters or terms used for partial matching, autocomplete, and selected language-analysis tasks.

    7

    What are token positions?

    Positions record logical token order and support phrase and proximity queries.

    8

    What are token offsets?

    Offsets record a token's source-text range and support highlighting and passage extraction.

    9

    Should index-time and query-time analyzers be identical?

    Not necessarily, but they must produce compatible terms for the intended search behaviour.

    10

    Does changing a tokenizer require reindexing?

    An incompatible index-time analyzer change normally requires rebuilding existing searchable terms.

    11

    Are search tokens the same as LLM tokens?

    No. Search tokens are designed for indexing and lexical retrieval, while model tokens are defined by a model-specific input and output vocabulary.

    12

    What is a good tokenization strategy?

    Use field-specific, language-aware, Unicode-safe analyzers that preserve technical identifiers, positions, offsets, exact values, and compatibility between indexing and querying.

    Key Takeaway

    Tokenization converts raw text into the discrete units used by inverted indexes and lexical ranking. The best unit depends on the field, language, domain, and search experience. Natural-language fields can use word-aware analysis, exact metadata can use keyword tokenization, autocomplete can use bounded prefixes, and specialized languages can require morphological or n-gram segmentation. Preserve token positions for phrase search and offsets for highlighting. Apply normalization, stop-word removal, stemming, and synonyms conservatively because each transformation changes both matching and ranking statistics. Technical content requires explicit tests for product names, symbols, properties, error codes, and programming identifiers. Keep index-time and query-time analyzers compatible, version every analyzer, and reindex after incompatible changes. Finally, evaluate tokenization with representative queries and monitor vocabulary size, token counts, index growth, zero-result queries, and retrieval quality.