Table of Contents

    full-text search

    STORAGE, FILES, OBJECTS & SEARCH BASICS

    Full-Text Search

    Learn how full-text search analyzes user queries, retrieves matching documents from inverted indexes, ranks results by relevance, supports phrases, Boolean logic, synonyms, fuzzy matching, filters, facets, highlighting, autocomplete, multilingual content, and secure tenant-aware retrieval.

    Introduction

    Applications frequently need to search inside long text fields such as article titles, descriptions, course lessons, documentation, product descriptions, support records, and uploaded documents.

    Exact database comparisons are useful when users know a precise identifier, status, category, or code. Full-text search addresses a different problem: finding documents whose textual content is relevant to a user's search terms.

    Full-text search commonly uses an inverted index to retrieve candidate documents. It then applies query interpretation, structured filters, relevance scoring, authorization, sorting, highlighting, and pagination to produce the final search experience.

    Core idea: An inverted index answers which documents contain the analyzed terms. Full-text search determines how the query should be interpreted, which candidates are allowed, how relevant each candidate is, and how results should be presented.

    In your System Design curriculum, Full-Text Search is Topic 6.7 and completes the Storage, Files, Objects & Search Basics module. It follows inverted indexes.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Inverted indexes Full-text retrieval normally begins with term dictionaries and postings lists.
    2 Metadata Structured metadata supports filtering, security, sorting, and faceting.
    3 Tokenization and normalization Documents and queries must be transformed into compatible searchable terms.
    4 Basic ranking concepts Search systems order candidates according to estimated relevance.
    5 API design Search endpoints require clear query, filters, pagination, and error contracts.
    6 Authorization Search results, snippets, counts, and facets must not disclose restricted content.
    7 Observability Search quality and performance require query, relevance, latency, and freshness metrics.

    What Is Full-Text Search?

    Full-text search locates documents by analyzing and matching terms inside textual content rather than relying only on exact field equality.

    User query
        |
        v
    Query analysis
        |
        v
    Retrieve candidates from inverted index
        |
        v
    Apply structured and security filters
        |
        v
    Calculate relevance scores
        |
        v
    Sort and paginate
        |
        v
    Generate snippets and highlights
        |
        v
    Return search results

    A complete full-text search experience can support:

    • Keyword matching
    • Multi-term queries
    • Phrase matching
    • Boolean operators
    • Language-aware analysis
    • Stemming or lemmatization
    • Synonyms
    • Fuzzy matching
    • Prefix matching
    • Relevance ranking
    • Structured filters
    • Facets
    • Highlighting
    • Autocomplete
    • Security trimming
    Search Flow
    understand query → retrieve candidates → filter securely → rank results → present clearly

    Exact Search vs Full-Text Search

    Area Exact Search Full-Text Search
    Typical question Find order number ORD-1042 Find lessons about scalable database design
    Primary comparison Exact value, prefix, or range Analyzed terms, phrases, and relevance
    Common index B-tree or hash-oriented index Inverted index
    Result ordering Defined column order Relevance score combined with business rules
    Language handling Usually limited to comparison and collation rules Can use tokenization, normalization, stemming, synonyms, and language analyzers

    SQL LIKE Search

    A SQL LIKE predicate can be useful for small datasets, controlled pattern matching, and selected prefix-search requirements.

    SELECT
        article_id,
        article_title
    FROM articles
    WHERE article_title LIKE 'System Design%';

    A leading wildcard commonly requires broader examination because the search does not begin from a known prefix.

    SELECT
        article_id,
        article_title
    FROM articles
    WHERE article_body LIKE '%distributed database%';

    LIKE does not by itself provide a complete relevance-ranking, language-analysis, phrase, synonym, fuzzy, facet, or highlighting system.

    Selection rule: Use ordinary relational predicates for exact structured lookup. Use a full-text index when the application needs analyzed textual retrieval and relevance ranking.

    Query-analysis Pipeline

    Search queries normally pass through an analysis pipeline compatible with the pipeline used for indexed content.

    Raw query:
    
    "Designing Scalable Databases"
    
    
    Tokenization:
    
    designing
    scalable
    databases
    
    
    Normalization:
    
    designing
    scalable
    databases
    
    
    Optional stemming:
    
    design
    scalable
    database
    
    
    Final query terms:
    
    design
    scalable
    database

    Query analysis can also apply synonym expansion, phrase interpretation, spelling logic, field selection, and operator parsing.

    Language-aware Search

    Search behaviour depends on language. Word boundaries, grammatical structure, inflection, character normalization, common words, and compound terms vary between languages.

    A multilingual design can:

    • Store the document language explicitly
    • Use the corresponding analyzer at index time
    • Detect or request the query language
    • Apply compatible query-time analysis
    • Filter or boost results by language
    • Use language-specific stop words and stemming
    {
      "documentId": "tenant-17:article-981",
      "languageCode": "en",
      "analyzerVersion": 3,
      "title": "Full-Text Search",
      "body": "Searchable article content"
    }

    Multilingual rule: Do not silently apply one language's stemming and stop-word rules to content written in another language.

    Single-term Search

    A single-term query uses the term's postings list to retrieve matching documents.

    Search query:
    
    database
    
    
    Postings:
    
    Document 2
    Document 5
    Document 8
    Document 13
    
    
    Candidate result set:
    
    [2, 5, 8, 13]

    The system can then calculate relevance, apply filters, and return the highest-ranked authorized results.

    Multi-term Search

    A multi-term query can require all terms, any term, a minimum number of terms, or a specific phrase.

    Query:
    
    distributed database
    
    
    Possible interpretations:
    
    distributed AND database
    
    distributed OR database
    
    "distributed database"
    
    At least one term,
    with documents matching both terms ranked higher

    The API and user interface should define how unquoted multi-term queries are interpreted.

    Boolean Search

    AND

    database AND indexing
    
    Return documents containing
    both required terms.

    OR

    database OR storage
    
    Return documents containing
    either term.

    NOT

    database NOT relational
    
    Return documents matching database
    after excluding documents
    matching relational.

    Search syntax should be validated and bounded. Unrestricted nested Boolean expressions can create expensive queries.

    Phrase Search

    Phrase search requires terms to occur in the requested order and proximity.

    Query:
    
    "full text search"
    
    
    Document A:
    
    Learn full text search.
    
    
    Document B:
    
    Read the full article about
    text indexing and search.
    
    
    Result:
    
    Document A contains the exact phrase.
    Document B contains the terms,
    but not as the exact phrase.

    Phrase matching normally depends on term-position information in the inverted index.

    Proximity Search

    Proximity search matches terms that appear near each other without requiring exact adjacency.

    Terms:
    
    database
    performance
    
    
    Allowed distance:
    
    5 token positions
    
    
    Document:
    
    Improve database query performance
    with suitable indexes.
    
    
    Result:
    
    The terms occur within
    the allowed distance.

    Proximity can be used as a matching condition or relevance signal.

    Fuzzy Search

    Fuzzy search retrieves terms that differ from the query term by a bounded number of edits.

    User query:
    
    databse
    
    
    Candidate indexed term:
    
    database

    Possible edit operations include:

    • Insert a character
    • Delete a character
    • Replace a character
    • Apply implementation-supported transposition handling

    Fuzzy matching can improve typo tolerance but can also expand a query into many candidate terms.

    Fuzzy-search rule: Bound edit distance, term expansion, input length, and result count. Broad fuzzy matching can increase latency and return irrelevant documents.

    Edit Distance

    Edit distance is the minimum number of supported character operations required to transform one term into another.

    Query:
    
    serch
    
    
    Indexed term:
    
    search
    
    
    One insertion:
    
    Add "a"
    
    
    Edit distance:
    
    1

    The exact distance algorithm and transposition behaviour depend on the implementation.

    Prefix Search

    Prefix search matches indexed terms beginning with a supplied prefix.

    Prefix:
    
    data
    
    
    Possible matching terms:
    
    data
    database
    datacenter
    dataset

    Prefix matching is useful for autocomplete and partially entered terms. Very short prefixes can expand to many terms and should be limited.

    Wildcard Search

    Wildcard search allows a query pattern to represent unknown characters.

    Pattern:
    
    data*
    
    
    Possible matches:
    
    data
    database
    datastore
    dataset

    Leading or unrestricted wildcard patterns can be expensive because they can expand across a large portion of the term dictionary.

    Synonyms

    Synonyms connect terms that users or the business treat as related.

    Query term:
    
    DBMS
    
    
    Possible synonym expansion:
    
    DBMS
    database management system

    Domain-oriented examples include:

    • D365 F&O and Dynamics 365 Finance and Operations
    • DB and database
    • ML and machine learning
    • FTS and full-text search
    • Object store and object storage

    Synonym rules should use an approved vocabulary and should be reviewed when terminology changes.

    Index-time vs Query-time Synonyms

    Approach Behaviour Consideration
    Index-time expansion Additional synonym terms are stored during indexing Changing synonyms can require reindexing
    Query-time expansion The user's query is expanded during search Can increase query complexity and term expansion

    The choice depends on relevance requirements, synonym update frequency, phrase behaviour, index size, and search-engine capabilities.

    Stemming and Lemmatization

    Stemming or lemmatization helps related word forms match.

    Possible related forms:
    
    index
    indexed
    indexing
    indexes

    These techniques can improve recall but can also merge words that should remain distinct. Language and domain testing are essential.

    Relevance Ranking

    Full-text search normally returns the most relevant results first rather than using only creation date or primary-key order.

    Ranking can consider:

    • How often query terms occur in the document
    • How rare each term is across the collection
    • Which field contains the match
    • Document length
    • Phrase and proximity matches
    • Number of query terms matched
    • Freshness
    • Content quality
    • Business importance
    • Language compatibility
    • Popularity or engagement signals

    TF-IDF Intuition

    TF-IDF combines term frequency with inverse document frequency.

    A simplified representation is:

    \[ TFIDF(t,d) = TF(t,d) \times IDF(t) \]

    A simplified inverse document frequency expression is:

    \[ IDF(t) = \log \left( \frac{N}{df(t)} \right) \]

    Where:

    • \(t\) is the query term
    • \(d\) is the document
    • \(N\) is the number of indexed documents
    • \(df(t)\) is the number of documents containing the term

    This gives less-common terms more distinguishing power than terms appearing in nearly every document.

    BM25 Intuition

    BM25 is a relevance-ranking approach that uses term rarity, term-frequency saturation, and document-length normalization.

    A commonly presented form is:

    \[ Score(D,Q) = \sum_{t \in Q} IDF(t) \cdot \frac{ f(t,D)(k_1 + 1) }{ f(t,D) + k_1 \left( 1-b+b\frac{|D|}{avgdl} \right) } \]

    Where:

    • \(f(t,D)\) is the frequency of term \(t\) in document \(D\)
    • \(|D|\) is the document length
    • \(avgdl\) is the average document length
    • \(k_1\) controls term-frequency saturation
    • \(b\) controls document-length normalization

    Search engines can use modified implementations and additional scoring signals. Treat this equation as conceptual guidance rather than a portable score contract.

    Term-frequency Saturation

    Repeating a term many times should not normally make a document proportionally more relevant forever.

    Document A:
    
    "database" appears once
    
    
    Document B:
    
    "database" appears ten times
    
    
    Document C:
    
    "database" appears one hundred times
    
    
    Ranking expectation:
    
    Additional occurrences can help,
    but their contribution eventually provides
    smaller incremental value.

    Saturation reduces the benefit of excessive term repetition.

    Document-length Normalization

    A term match in a short focused document can provide stronger evidence than the same number of occurrences in a very long document.

    Short document:
    
    100 terms
    "index" appears 3 times
    
    
    Long document:
    
    10,000 terms
    "index" appears 3 times
    
    
    Observation:
    
    The term is more concentrated
    in the short document.

    Field Boosting

    Matches in different fields can contribute different ranking weights.

    Possible field importance:
    
    Title:
    High weight
    
    
    Keywords:
    Medium-high weight
    
    
    Description:
    Medium weight
    
    
    Article body:
    Standard weight

    Search-document Example

    {
      "documentId": "tenant-17:article-981",
      "title": "Full-Text Search",
      "description": "Learn query analysis, ranking, and secure retrieval.",
      "body": "Complete article content.",
      "keywords": [
        "search",
        "BM25",
        "ranking"
      ],
      "language": "en",
      "status": "published"
    }

    Excessive boosting can cause title keyword repetition or business metadata to dominate genuine relevance.

    Business Boosting

    Search results can combine textual relevance with approved business signals.

    Possible signals include:

    • Publication state
    • Content freshness
    • Course priority
    • Verified quality
    • Learner completion context
    • Approved editorial promotion

    Business boosts should be transparent, bounded, measurable, and separate from access-control decisions.

    Structured Filters

    Full-text retrieval is commonly combined with exact metadata filters.

    Text query:
    
    database indexing
    
    
    Structured filters:
    
    tenant_id = tenant-17
    status = published
    language = en
    content_type = article
    course_id = course-42

    Filters usually determine eligibility rather than textual relevance.

    Faceted Search

    Facets summarize available result categories and allow users to narrow the result set.

    Search:
    
    system design
    
    
    Content type:
    
    Articles        42
    Videos          18
    Courses          7
    Quizzes          5
    
    
    Difficulty:
    
    Beginner        31
    Intermediate    24
    Advanced        17

    Facet counts must be computed within the caller's authorized result scope. Otherwise, a count can reveal the existence of restricted content.

    Sorting

    Search results can be sorted by relevance or a structured field.

    Common options include:

    • Most relevant
    • Newest first
    • Oldest first
    • Title
    • Popularity
    • Course order

    Sorting only by date can discard relevance ordering. Sorting only by relevance can make time-sensitive content difficult to discover. The user interface should communicate the active ordering.

    Highlighting

    Highlighting shows where query terms matched in a result.

    Query:
    
    inverted index
    
    
    Result snippet:
    
    An [inverted index] maps terms
    to the documents containing them.

    Highlighting can use indexed positions or offsets, reanalyze stored content, or apply provider-specific mechanisms.

    Rendering rule: Search snippets are untrusted text. Encode snippets safely before rendering them in HTML, and allow only controlled highlighting markup.

    Result Snippets

    A snippet presents a short relevant portion of a larger document.

    {
      "documentId": "tenant-17:article-981",
      "title": "Full-Text Search",
      "snippet": "Full-text search retrieves and ranks documents...",
      "score": 8.42,
      "contentType": "article",
      "publishedAt": "publication-time"
    }

    Snippets must not expose text from fields or documents the caller is not permitted to view.

    Autocomplete

    Autocomplete suggests terms, titles, or queries while the user types.

    User enters:
    
    full t
    
    
    Suggestions:
    
    full-text search
    full-text index
    full table scan
    full transaction log

    Autocomplete can use prefixes, edge-oriented tokens, curated suggestions, popular approved queries, or a dedicated suggestion index.

    Avoid executing the complete expensive search pipeline after every keystroke. Apply minimum input length, debouncing, cancellation, and result limits.

    Search-as-you-type

    Keystroke
        |
        v
    Wait for short debounce interval
        |
        v
    Cancel obsolete request
        |
        v
    Send bounded prefix query
        |
        v
    Return limited suggestions

    Suggestions should be tenant-aware and access-controlled when derived from private content.

    Spell Correction

    A search system can suggest a corrected query when an input term is rare or absent.

    User query:
    
    distrubuted systems
    
    
    Suggestion:
    
    Did you mean:
    distributed systems?

    A correction can be shown as a suggestion or applied automatically according to the product's search contract.

    Correction rule: Preserve the user's original query and show when a correction was applied. Automatic correction can be harmful for codes, names, acronyms, and domain-specific terminology.

    Query Expansion

    Query expansion adds related terms to improve recall.

    Original query:
    
    ML
    
    
    Expanded query:
    
    ML
    machine learning

    Expansion sources can include:

    • Approved synonyms
    • Acronym dictionaries
    • Alternate spellings
    • Domain taxonomies
    • Related controlled terms

    Uncontrolled expansion can reduce precision by adding terms with a different meaning.

    Precision and Recall

    Precision measures how many returned results are relevant.

    \[ Precision = \frac{ RelevantRetrievedDocuments }{ RetrievedDocuments } \]

    Recall measures how many of the relevant documents were retrieved.

    \[ Recall = \frac{ RelevantRetrievedDocuments }{ AllRelevantDocuments } \]

    Broad synonym and fuzzy expansion can improve recall while reducing precision. Exact phrases and strict filters can improve precision while reducing recall.

    Precision vs Recall Example

    Search Strategy Likely Effect
    Exact phrase only Higher precision, lower recall
    All terms required More restrictive result set
    Any term allowed Broader result set with possible relevance noise
    Synonym expansion Can retrieve conceptually related terminology
    Fuzzy expansion Can recover typographical variations while adding candidates

    Search-quality Evaluation

    A search system should be evaluated using representative queries and relevance judgments.

    Create a test set containing:

    • Common user queries
    • Rare technical terms
    • Acronyms
    • Misspellings
    • Phrase queries
    • Multilingual queries
    • Queries with no valid result
    • Queries with restricted results
    • Freshly published content
    • Deleted or access-revoked content

    Precision at K

    Precision at \(K\) measures the proportion of relevant documents among the first \(K\) results.

    \[ Precision@K = \frac{ RelevantDocumentsInTopK }{ K } \]

    This is useful when users primarily inspect the first page of results.

    Reciprocal Rank

    Reciprocal rank rewards placing the first relevant result near the top.

    \[ ReciprocalRank = \frac{ 1 }{ RankOfFirstRelevantResult } \]

    Mean Reciprocal Rank averages this value across a query set.

    NDCG Intuition

    Normalized Discounted Cumulative Gain evaluates ranked results when relevance can have several levels.

    It gives more value to highly relevant documents appearing near the top and discounts relevant documents appearing lower in the result list.

    Use the exact formula and relevance scale defined by the selected evaluation framework.

    Pagination

    Search results need stable and efficient pagination.

    Offset Pagination

    Page 1:
    
    Offset 0
    Size 20
    
    
    Page 2:
    
    Offset 20
    Size 20

    Deep offsets can require the search system to identify and skip many earlier ranked results.

    Cursor or Search-after Pagination

    {
      "lastScore": 7.81,
      "lastDocumentId": "tenant-17:article-981",
      "searchSnapshot": "search-context-id"
    }

    The next request continues after the last result using a deterministic ordering and provider-supported search context.

    Pagination rule: Include a deterministic tiebreaker such as a stable document ID. Relevance scores can be identical, and the index can change between requests.

    Search Freshness

    Full-text indexes are commonly derived from an authoritative source. Therefore, a committed source change might not appear in search immediately.

    Article transaction commits
          |
          v
    Indexing event is published
          |
          v
    Search consumer processes event
          |
          v
    Index refresh completes
          |
          v
    Article becomes searchable

    Define freshness objectives for:

    • New documents
    • Document updates
    • Deletions
    • Publication changes
    • Tenant ownership changes
    • Access revocation

    Security Trimming

    Full-text search must not return documents the caller is not authorized to discover.

    Search query
          |
          v
    Apply trusted tenant scope
          |
          v
    Apply publication status
          |
          v
    Apply visibility rules
          |
          v
    Apply course enrollment,
    ownership, role, or group scope
          |
          v
    Rank authorized candidates
          |
          v
    Return results and facets

    Security must cover:

    • Document titles
    • Search snippets
    • Highlight fragments
    • Facet counts
    • Total result count
    • Autocomplete suggestions
    • Cached search responses
    • Search analytics

    Authorization rule: Filtering unauthorized documents after constructing visible results is too late. Apply trusted access scope before returning titles, snippets, counts, facets, or suggestions.

    Tenant-aware Search

    {
      "query": "database indexing",
      "trustedTenantId": "tenant-17",
      "filters": {
        "status": "published",
        "language": "en"
      }
    }

    The authenticated application context should supply the tenant scope. A client-provided tenant identifier alone must not be treated as authorization.

    Search Index as a Derived Store

    A search index is commonly optimized for retrieval rather than treated as the authoritative system of record.

    Authoritative database or content store
            |
            v
    Indexing pipeline
            |
            v
    Search index
            |
            v
    Search API

    The architecture should support:

    • Idempotent indexing
    • Deletion propagation
    • Source-version comparison
    • Reconciliation
    • Schema migration
    • Analyzer migration
    • Complete index rebuild

    Indexing Events

    {
      "eventType": "ArticlePublished",
      "tenantId": "tenant-17",
      "sourceType": "article",
      "sourceId": 981,
      "sourceVersion": 12,
      "occurredAt": "event-time"
    }

    The indexing consumer retrieves the authoritative content, builds the searchable document, and writes it using a stable search-document ID.

    Idempotent Indexing

    Repeated delivery of the same indexing event should not create duplicate search documents.

    Stable search-document ID:
    
    tenant-17:article:981
    
    
    Source version:
    
    12
    
    
    Repeated event:
    
    Update same search document
    or return already-processed outcome.

    An older source version should not overwrite a newer indexed version.

    Reindexing

    Reindexing rebuilds search documents from authoritative content.

    Read authoritative documents
          |
          v
    Create replacement index
          |
          v
    Apply current schema and analyzers
          |
          v
    Validate counts, queries, and security
          |
          v
    Switch search alias or routing
          |
          v
    Retire old index safely

    Reindexing can be required after:

    • Analyzer changes
    • Synonym-strategy changes
    • Search-document schema changes
    • Field-type changes
    • Major ranking changes
    • Index corruption
    • Authoritative-data correction

    Index Versioning

    Logical search name:
    
    course-content
    
    
    Physical index:
    
    course-content-v1
    
    
    Replacement index:
    
    course-content-v2
    
    
    After validation:
    
    Logical search name
    points to v2

    Versioned indexes allow a replacement to be built and validated without destroying the currently serving index.

    Search Architecture

    Web or mobile client
            |
            v
    Search API
            |
            +-- Authentication
            +-- Authorization
            +-- Query validation
            +-- Rate limiting
            +-- Tenant scope
            |
            v
    Search cluster
            |
            +-- Text index
            +-- Metadata filters
            +-- Ranking
            +-- Facets
            |
            v
    Authorized ranked results

    The search API prevents clients from bypassing trusted filters, exposes a stable contract, and centralizes resource limits.

    Search API Request

    POST /api/v1/search/articles HTTP/1.1
    Content-Type: application/json
    Authorization: Bearer access-token
    
    {
      "query": "full text search",
      "language": "en",
      "filters": {
        "courseId": 42,
        "contentType": "article"
      },
      "sort": "relevance",
      "pageSize": 20,
      "cursor": null
    }

    Search API Response

    {
      "query": "full text search",
      "totalRelation": "exact-or-provider-defined",
      "results": [
        {
          "documentId": "tenant-17:article:981",
          "title": "Full-Text Search",
          "snippet": "Learn how full-text search retrieves and ranks...",
          "score": 8.42,
          "contentType": "article",
          "language": "en"
        }
      ],
      "facets": {
        "contentType": [
          {
            "value": "article",
            "count": 12
          },
          {
            "value": "video",
            "count": 5
          }
        ]
      },
      "nextCursor": "opaque-cursor-value"
    }

    The cursor should be treated as an opaque server-generated value. Result counts can be exact or approximate according to the implementation and contract.

    PHP Search-request Validation

    <?php
    
    declare(strict_types=1);
    
    function validateSearchRequest(
        array $request
    ): array {
        $query =
            trim(
                (string)($request['query'] ?? '')
            );
    
        if ($query === '') {
            throw new InvalidArgumentException(
                'The search query is required.'
            );
        }
    
        if (mb_strlen(
            $query,
            'UTF-8'
        ) > 300) {
            throw new InvalidArgumentException(
                'The search query is too long.'
            );
        }
    
        $pageSize =
            filter_var(
                $request['pageSize'] ?? 20,
                FILTER_VALIDATE_INT
            );
    
        if ($pageSize === false ||
            $pageSize < 1 ||
            $pageSize > 100) {
    
            throw new InvalidArgumentException(
                'The page size is invalid.'
            );
        }
    
        $allowedSorts = [
            'relevance',
            'newest',
            'oldest',
            'title'
        ];
    
        $sort =
            strtolower(
                trim(
                    (string)($request['sort'] ?? 'relevance')
                )
            );
    
        if (!in_array(
            $sort,
            $allowedSorts,
            true
        )) {
            throw new InvalidArgumentException(
                'The sort option is invalid.'
            );
        }
    
        return [
            'query' => $query,
            'pageSize' => $pageSize,
            'sort' => $sort,
            'cursor' =>
                isset($request['cursor'])
                    ? (string)$request['cursor']
                    : null
        ];
    }

    Tenant ID, user ID, roles, and permissions must come from authenticated server-side context rather than untrusted request fields.

    Simplified PHP Search Example

    <?php
    
    declare(strict_types=1);
    
    function searchArticles(
        SearchClient $searchClient,
        string $tenantId,
        array $validatedRequest
    ): array {
        return $searchClient->search(
            index: 'course-content',
            query: [
                'text' =>
                    $validatedRequest['query'],
    
                'filters' => [
                    'tenantId' =>
                        $tenantId,
    
                    'status' =>
                        'published'
                ],
    
                'sort' =>
                    $validatedRequest['sort'],
    
                'size' =>
                    $validatedRequest['pageSize'],
    
                'cursor' =>
                    $validatedRequest['cursor']
            ]
        );
    }

    SearchClient is a conceptual abstraction. Use the official supported client and exact query syntax of the selected search platform.

    Query Safety

    User-provided search syntax can create expensive or unintended queries.

    Apply limits to:

    • Query length
    • Number of terms
    • Boolean nesting depth
    • Wildcard expansion
    • Fuzzy edit distance
    • Prefix length
    • Synonym expansion
    • Facet count
    • Page size
    • Highlight fragments
    • Search timeout

    Prefer a structured search request over directly exposing unrestricted provider query syntax to untrusted clients.

    Search Timeouts

    Request deadline
          |
          +-- Authentication time
          +-- Search-API processing
          +-- Search-cluster execution
          +-- Result transformation
          +-- Network response

    Keep the search-engine timeout within the complete API deadline. A timed-out query should not continue consuming unbounded server resources.

    Rate Limiting

    Search endpoints can be abused through automated broad queries, autocomplete storms, deep pagination, large facet requests, and expensive wildcard or fuzzy expressions.

    Rate limits can be scoped by:

    • User
    • Tenant
    • Application
    • Endpoint
    • Query category
    • Anonymous network identity where appropriate

    Sharded Search

    A large search index can be divided into shards.

    Search request
          |
          v
    Coordinator
          |
          +-- Query Shard 1
          +-- Query Shard 2
          +-- Query Shard 3
          |
          v
    Collect shard candidates
          |
          v
    Merge and rank
          |
          v
    Return top results

    Distributed ranking requires care because term and document statistics can differ across shards.

    Search Replicas

    Search shards can have replicas for read capacity and failure recovery.

    Shard A primary
          |
          +-- Replica A1
          |
          +-- Replica A2

    Replication can increase query capacity but also consumes storage and indexing resources. Replica freshness and health should be monitored.

    Search Caching

    Search systems can cache repeated query components or final result pages.

    Cache keys must account for:

    • Normalized query
    • Tenant scope
    • Authorization scope
    • Filters
    • Sort order
    • Language
    • Page or cursor
    • Index version

    Cache-security rule: Never share a search response across users or tenants unless the cache key and authorization design prove that every recipient has the same permitted result set.

    Search Analytics

    Search analytics can reveal how users interact with the search experience.

    Useful signals include:

    • Search-query count
    • Zero-result queries
    • Low-result queries
    • Query reformulations
    • Result clicks
    • Click position
    • Abandoned searches
    • Applied filters
    • Spelling suggestions accepted
    • Search latency

    Search queries can contain personal, confidential, or sensitive information. Apply privacy, retention, access-control, and redaction requirements to search logs and analytics.

    Zero-result Queries

    A zero-result query can indicate:

    • The requested content does not exist
    • The content is not yet indexed
    • The caller is not authorized
    • The analyzer produced unexpected terms
    • A synonym or acronym is missing
    • The query contains a spelling error
    • Filters are too restrictive
    • Content is in another language

    Do not expose whether restricted results exist when the caller is not authorized to discover them.

    Relevance Testing

    Create a controlled relevance test set.

    Query Expected High-ranking Result Reason
    inverted index Inverted Indexes article Exact title and body match
    DBMS transaction reliability ACID article Domain synonym and subject match
    databse indexing Database indexing lesson Approved typo tolerance
    "full text search" Full-Text Search article Exact phrase match
    large video upload Uploads and Multipart Transfer article Concept and content match

    Lexical, Vector, and Hybrid Search

    Search Type Primary Matching Basis Typical Strength
    Lexical full-text Analyzed terms and inverted-index statistics Exact terms, phrases, identifiers, filters, and explainable keyword relevance
    Vector search Distance between vector representations Semantic similarity and paraphrased concepts
    Hybrid search Combination of lexical and vector retrieval Balances exact terminology with semantic similarity

    Full-text search remains important when users search for exact names, codes, quoted phrases, technical terminology, or rare identifiers.

    Hybrid Retrieval Flow

    User query
          |
          +-- Lexical retrieval
          |      |
          |      +-- Exact terms
          |      +-- Phrases
          |      +-- Keyword scoring
          |
          +-- Vector retrieval
                 |
                 +-- Semantic similarity
          |
          v
    Combine candidate sets
          |
          v
    Apply security filters
          |
          v
    Rerank and return results

    Candidate combination and score normalization depend on the selected search architecture.

    Search Observability

    Useful full-text search metrics include:

    • Search-request rate
    • Search latency by percentile
    • Timeout rate
    • Error rate
    • Zero-result rate
    • Result count distribution
    • Click-through rate
    • First-result click rate
    • Query reformulation rate
    • Autocomplete latency
    • Facet latency
    • Indexing lag
    • Stale-document count
    • Deleted documents still searchable
    • Shard and replica health
    • Cache hit rate
    • Security-filter failures

    Alert Conditions

    Alert when:

    • Search latency exceeds its objective
    • Timeouts or errors increase
    • Indexing lag grows
    • Zero-result queries increase unexpectedly
    • One shard receives disproportionate traffic
    • Replicas become unavailable
    • Deleted or restricted content remains searchable
    • Autocomplete or facet requests consume excessive resources
    • Relevance metrics regress after an analyzer or ranking change

    Search Troubleshooting Workflow

    1. Capture the exact query, filters, language, and trusted tenant scope.
    2. Confirm the expected document exists in the authoritative source.
    3. Confirm the document is published and authorized.
    4. Confirm the current document version is indexed.
    5. Analyze the query into terms.
    6. Analyze the expected document field using the same analyzer.
    7. Compare query and indexed terms.
    8. Check synonyms, stop words, stemming, and fuzzy settings.
    9. Check structured and security filters.
    10. Inspect the document's relevance score components.
    11. Check indexing lag and failed indexing events.
    12. Check shard, replica, timeout, and resource health.
    13. Reindex or repair derived search state when required.

    Common Full-text Search Mistakes

    1

    Using SQL LIKE as the Complete Search Engine

    LIKE does not provide the complete language analysis, ranking, phrase, synonym, fuzzy, facet, and highlighting capabilities expected from full-text search.

    2

    Using Different Index and Query Analyzers

    The query can produce terms that do not match the terms stored during indexing.

    3

    Applying One Analyzer to Every Language

    Tokenization, normalization, stemming, and common-word behaviour vary between languages.

    4

    Expanding Every Query Aggressively

    Broad synonyms, prefixes, wildcards, and fuzzy terms can reduce precision and increase query cost.

    5

    Ranking Only by Term Count

    Raw term frequency ignores term rarity, field importance, document length, proximity, and business requirements.

    6

    Boosting Freshness Too Strongly

    A recent but weakly related document can outrank an older highly relevant document.

    7

    Returning Unauthorized Facet Counts

    Counts can disclose the existence and classification of restricted documents.

    8

    Rendering Highlights without Safe Encoding

    Indexed document content is untrusted and can create unsafe browser output when rendered directly.

    9

    Using Deep Offset Pagination

    The engine can perform significant work to identify and discard previous ranked results.

    10

    Ignoring Indexing Lag

    Users can see missing updates, stale titles, old permissions, or deleted content.

    11

    Changing Ranking without Relevance Tests

    A change that improves one query can reduce quality across other important query categories.

    12

    Treating the Search Index as the Only Data Source

    The index is commonly a derived representation and should be rebuildable from authoritative content and metadata.

    Recommended Test Cases

    Test Expected Evidence
    Exact keyword Documents containing the analyzed term are returned
    Case variation Approved uppercase and lowercase forms behave consistently
    Multi-term AND query Only documents satisfying all required terms remain
    Phrase query Terms occur in the required order and proximity
    Synonym query Approved equivalent terminology retrieves expected documents
    Typographical error Bounded fuzzy or correction logic produces the documented behaviour
    Structured filter Only matching language, type, course, or status remains
    Field boosting A strong title match receives the designed relevance treatment
    Highlighting Safe snippets identify the matched terms
    Cursor pagination Results continue without unintended duplicates or omissions
    Tenant isolation Another tenant's results, counts, facets, and suggestions remain hidden
    Access revocation Restricted content stops appearing within the defined objective
    Deleted document The document and its snippets no longer appear
    Index rebuild Document counts, relevance tests, and security filters remain correct

    Full-text Search Best Practices

    Recommended Practices

    • Use an inverted index for scalable analyzed text retrieval.
    • Keep authoritative content separate from the derived search index.
    • Use compatible index-time and query-time analyzers.
    • Select analyzers according to content language.
    • Use approved synonym and acronym dictionaries.
    • Bound fuzzy, wildcard, prefix, and Boolean expansion.
    • Combine textual retrieval with structured metadata filters.
    • Apply trusted tenant and authorization scope before returning any search information.
    • Use stable search-document IDs and authoritative source versions.
    • Make indexing operations idempotent.
    • Reject stale out-of-order indexing events.
    • Use relevance ranking rather than raw term count alone.
    • Test title, body, phrase, synonym, and typo behaviour separately.
    • Use deterministic cursor-based pagination for deep result navigation.
    • Encode snippets and highlights safely.
    • Define freshness objectives for create, update, delete, and access changes.
    • Maintain a tested full-index rebuild procedure.
    • Protect search logs and analytics as potentially sensitive data.
    • Measure precision, recall, ranking quality, latency, and zero-result queries.
    • Evaluate ranking changes against a representative relevance test set.

    Practice Exercise

    Design full-text search for articles, courses, videos, and lessons on your online learning platform.

    Requirements

    1. Create a stable search-document ID for each searchable item.
    2. Index title, description, body, keywords, language, and content type.
    3. Use exact fields for tenant, status, course, and visibility filters.
    4. Use compatible index-time and query-time analyzers.
    5. Support keyword and quoted phrase search.
    6. Support AND and OR logic.
    7. Add approved synonyms for course terminology and acronyms.
    8. Add bounded typo tolerance.
    9. Boost title matches above body-only matches.
    10. Return safely encoded snippets.
    11. Return authorized facet counts for content type and difficulty.
    12. Support relevance and newest-first sorting.
    13. Support cursor-based pagination.
    14. Prevent cross-tenant search disclosure.
    15. Propagate publication, deletion, and access changes.
    16. Measure ranking quality using representative queries.
    17. Track zero-result queries and accepted corrections.
    18. Test a complete index rebuild.

    Search-design Template

    Feature Design Decision Validation
    Searchable fields Title, description, body, and approved keywords Verify expected terms are indexed
    Filter fields Tenant, course, type, language, status, and visibility Verify exact filtering and security trimming
    Analyzer Language-aware tokenization and normalization Compare index and query tokens
    Ranking Text relevance plus bounded field and business boosts Evaluate representative ranked queries
    Typos Bounded correction or fuzzy expansion Test valid words, codes, and mistakes
    Pagination Opaque cursor with deterministic tiebreaker Check duplicate and missing results
    Freshness Idempotent event-driven indexing Measure create, update, and deletion delay
    Recovery Rebuild from authoritative content Validate counts, relevance, and security after rebuild

    Frequently Asked Questions

    1

    What is full-text search?

    Full-text search retrieves documents by analyzing textual content and ranking candidate documents according to relevance.

    2

    How is full-text search different from SQL LIKE?

    Full-text search uses analyzed terms and an inverted index and can provide relevance ranking, phrase matching, language processing, synonyms, and highlighting.

    3

    What is relevance ranking?

    Relevance ranking estimates how well each candidate document satisfies the user query and orders results accordingly.

    4

    What is BM25?

    BM25 is a text-ranking approach that considers term rarity, term-frequency saturation, and document-length normalization.

    5

    What is phrase search?

    Phrase search requires query terms to occur in the requested order and proximity.

    6

    What is fuzzy search?

    Fuzzy search permits bounded character differences between query terms and indexed terms to support selected spelling variations or mistakes.

    7

    What are facets?

    Facets summarize result categories and allow users to narrow a search using structured fields.

    8

    What is search highlighting?

    Highlighting marks matching terms inside a safely rendered result snippet.

    9

    Why can a new article be missing from search?

    The indexing event may not have been processed, the index may not have refreshed, analysis may differ, or filters and authorization may exclude the document.

    10

    Should the search index be the authoritative database?

    It is commonly a derived and rebuildable retrieval structure backed by an authoritative source of content and metadata.

    11

    What is hybrid search?

    Hybrid search combines lexical full-text retrieval with vector-based semantic retrieval.

    12

    What comes after full-text search?

    Full-text search completes the Storage, Files, Objects & Search Basics module.

    Key Takeaway

    Full-text search transforms a user's text into analyzed terms, retrieves candidate documents through an inverted index, applies structured and security filters, calculates relevance, and returns ranked results with snippets, highlights, facets, and pagination. Use language-compatible analyzers, controlled synonyms, bounded fuzzy and wildcard expansion, and relevance models that account for term rarity, frequency saturation, field importance, and document length. Treat the search index as a derived and rebuildable representation, process updates idempotently, reject stale events, and define freshness requirements for publication, deletion, and access changes. Most importantly, apply trusted tenant and authorization scope before returning documents, snippets, counts, facets, suggestions, or cached results.