full-text search
Full-Text Search
Learn how full-text search analyzes user queries, retrieves matching documents from inverted indexes, ranks results by relevance, supports phrases, Boolean logic, synonyms, fuzzy matching, filters, facets, highlighting, autocomplete, multilingual content, and secure tenant-aware retrieval.
Introduction
Applications frequently need to search inside long text fields such as article titles, descriptions, course lessons, documentation, product descriptions, support records, and uploaded documents.
Exact database comparisons are useful when users know a precise identifier, status, category, or code. Full-text search addresses a different problem: finding documents whose textual content is relevant to a user's search terms.
Full-text search commonly uses an inverted index to retrieve candidate documents. It then applies query interpretation, structured filters, relevance scoring, authorization, sorting, highlighting, and pagination to produce the final search experience.
Core idea: An inverted index answers which documents contain the analyzed terms. Full-text search determines how the query should be interpreted, which candidates are allowed, how relevant each candidate is, and how results should be presented.
In your System Design curriculum, Full-Text Search is Topic 6.7 and completes the Storage, Files, Objects & Search Basics module. It follows inverted indexes.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Inverted indexes | Full-text retrieval normally begins with term dictionaries and postings lists. |
| 2 | Metadata | Structured metadata supports filtering, security, sorting, and faceting. |
| 3 | Tokenization and normalization | Documents and queries must be transformed into compatible searchable terms. |
| 4 | Basic ranking concepts | Search systems order candidates according to estimated relevance. |
| 5 | API design | Search endpoints require clear query, filters, pagination, and error contracts. |
| 6 | Authorization | Search results, snippets, counts, and facets must not disclose restricted content. |
| 7 | Observability | Search quality and performance require query, relevance, latency, and freshness metrics. |
What Is Full-Text Search?
Full-text search locates documents by analyzing and matching terms inside textual content rather than relying only on exact field equality.
User query
|
v
Query analysis
|
v
Retrieve candidates from inverted index
|
v
Apply structured and security filters
|
v
Calculate relevance scores
|
v
Sort and paginate
|
v
Generate snippets and highlights
|
v
Return search results
A complete full-text search experience can support:
- Keyword matching
- Multi-term queries
- Phrase matching
- Boolean operators
- Language-aware analysis
- Stemming or lemmatization
- Synonyms
- Fuzzy matching
- Prefix matching
- Relevance ranking
- Structured filters
- Facets
- Highlighting
- Autocomplete
- Security trimming
Exact Search vs Full-Text Search
| Area | Exact Search | Full-Text Search |
|---|---|---|
| Typical question | Find order number ORD-1042 | Find lessons about scalable database design |
| Primary comparison | Exact value, prefix, or range | Analyzed terms, phrases, and relevance |
| Common index | B-tree or hash-oriented index | Inverted index |
| Result ordering | Defined column order | Relevance score combined with business rules |
| Language handling | Usually limited to comparison and collation rules | Can use tokenization, normalization, stemming, synonyms, and language analyzers |
SQL LIKE Search
A SQL LIKE predicate can be useful for small datasets, controlled pattern matching, and selected prefix-search requirements.
SELECT
article_id,
article_title
FROM articles
WHERE article_title LIKE 'System Design%';
A leading wildcard commonly requires broader examination because the search does not begin from a known prefix.
SELECT
article_id,
article_title
FROM articles
WHERE article_body LIKE '%distributed database%';
LIKE does not by itself provide a complete relevance-ranking, language-analysis, phrase, synonym, fuzzy, facet, or highlighting system.
Selection rule: Use ordinary relational predicates for exact structured lookup. Use a full-text index when the application needs analyzed textual retrieval and relevance ranking.
Query-analysis Pipeline
Search queries normally pass through an analysis pipeline compatible with the pipeline used for indexed content.
Raw query:
"Designing Scalable Databases"
Tokenization:
designing
scalable
databases
Normalization:
designing
scalable
databases
Optional stemming:
design
scalable
database
Final query terms:
design
scalable
database
Query analysis can also apply synonym expansion, phrase interpretation, spelling logic, field selection, and operator parsing.
Language-aware Search
Search behaviour depends on language. Word boundaries, grammatical structure, inflection, character normalization, common words, and compound terms vary between languages.
A multilingual design can:
- Store the document language explicitly
- Use the corresponding analyzer at index time
- Detect or request the query language
- Apply compatible query-time analysis
- Filter or boost results by language
- Use language-specific stop words and stemming
{
"documentId": "tenant-17:article-981",
"languageCode": "en",
"analyzerVersion": 3,
"title": "Full-Text Search",
"body": "Searchable article content"
}
Multilingual rule: Do not silently apply one language's stemming and stop-word rules to content written in another language.
Single-term Search
A single-term query uses the term's postings list to retrieve matching documents.
Search query:
database
Postings:
Document 2
Document 5
Document 8
Document 13
Candidate result set:
[2, 5, 8, 13]
The system can then calculate relevance, apply filters, and return the highest-ranked authorized results.
Multi-term Search
A multi-term query can require all terms, any term, a minimum number of terms, or a specific phrase.
Query:
distributed database
Possible interpretations:
distributed AND database
distributed OR database
"distributed database"
At least one term,
with documents matching both terms ranked higher
The API and user interface should define how unquoted multi-term queries are interpreted.
Boolean Search
AND
database AND indexing
Return documents containing
both required terms.
OR
database OR storage
Return documents containing
either term.
NOT
database NOT relational
Return documents matching database
after excluding documents
matching relational.
Search syntax should be validated and bounded. Unrestricted nested Boolean expressions can create expensive queries.
Phrase Search
Phrase search requires terms to occur in the requested order and proximity.
Query:
"full text search"
Document A:
Learn full text search.
Document B:
Read the full article about
text indexing and search.
Result:
Document A contains the exact phrase.
Document B contains the terms,
but not as the exact phrase.
Phrase matching normally depends on term-position information in the inverted index.
Proximity Search
Proximity search matches terms that appear near each other without requiring exact adjacency.
Terms:
database
performance
Allowed distance:
5 token positions
Document:
Improve database query performance
with suitable indexes.
Result:
The terms occur within
the allowed distance.
Proximity can be used as a matching condition or relevance signal.
Fuzzy Search
Fuzzy search retrieves terms that differ from the query term by a bounded number of edits.
User query:
databse
Candidate indexed term:
database
Possible edit operations include:
- Insert a character
- Delete a character
- Replace a character
- Apply implementation-supported transposition handling
Fuzzy matching can improve typo tolerance but can also expand a query into many candidate terms.
Fuzzy-search rule: Bound edit distance, term expansion, input length, and result count. Broad fuzzy matching can increase latency and return irrelevant documents.
Edit Distance
Edit distance is the minimum number of supported character operations required to transform one term into another.
Query:
serch
Indexed term:
search
One insertion:
Add "a"
Edit distance:
1
The exact distance algorithm and transposition behaviour depend on the implementation.
Prefix Search
Prefix search matches indexed terms beginning with a supplied prefix.
Prefix:
data
Possible matching terms:
data
database
datacenter
dataset
Prefix matching is useful for autocomplete and partially entered terms. Very short prefixes can expand to many terms and should be limited.
Wildcard Search
Wildcard search allows a query pattern to represent unknown characters.
Pattern:
data*
Possible matches:
data
database
datastore
dataset
Leading or unrestricted wildcard patterns can be expensive because they can expand across a large portion of the term dictionary.
Synonyms
Synonyms connect terms that users or the business treat as related.
Query term:
DBMS
Possible synonym expansion:
DBMS
database management system
Domain-oriented examples include:
- D365 F&O and Dynamics 365 Finance and Operations
- DB and database
- ML and machine learning
- FTS and full-text search
- Object store and object storage
Synonym rules should use an approved vocabulary and should be reviewed when terminology changes.
Index-time vs Query-time Synonyms
| Approach | Behaviour | Consideration |
|---|---|---|
| Index-time expansion | Additional synonym terms are stored during indexing | Changing synonyms can require reindexing |
| Query-time expansion | The user's query is expanded during search | Can increase query complexity and term expansion |
The choice depends on relevance requirements, synonym update frequency, phrase behaviour, index size, and search-engine capabilities.
Stemming and Lemmatization
Stemming or lemmatization helps related word forms match.
Possible related forms:
index
indexed
indexing
indexes
These techniques can improve recall but can also merge words that should remain distinct. Language and domain testing are essential.
Relevance Ranking
Full-text search normally returns the most relevant results first rather than using only creation date or primary-key order.
Ranking can consider:
- How often query terms occur in the document
- How rare each term is across the collection
- Which field contains the match
- Document length
- Phrase and proximity matches
- Number of query terms matched
- Freshness
- Content quality
- Business importance
- Language compatibility
- Popularity or engagement signals
TF-IDF Intuition
TF-IDF combines term frequency with inverse document frequency.
A simplified representation is:
\[ TFIDF(t,d) = TF(t,d) \times IDF(t) \]
A simplified inverse document frequency expression is:
\[ IDF(t) = \log \left( \frac{N}{df(t)} \right) \]
Where:
- \(t\) is the query term
- \(d\) is the document
- \(N\) is the number of indexed documents
- \(df(t)\) is the number of documents containing the term
This gives less-common terms more distinguishing power than terms appearing in nearly every document.
BM25 Intuition
BM25 is a relevance-ranking approach that uses term rarity, term-frequency saturation, and document-length normalization.
A commonly presented form is:
\[ Score(D,Q) = \sum_{t \in Q} IDF(t) \cdot \frac{ f(t,D)(k_1 + 1) }{ f(t,D) + k_1 \left( 1-b+b\frac{|D|}{avgdl} \right) } \]
Where:
- \(f(t,D)\) is the frequency of term \(t\) in document \(D\)
- \(|D|\) is the document length
- \(avgdl\) is the average document length
- \(k_1\) controls term-frequency saturation
- \(b\) controls document-length normalization
Search engines can use modified implementations and additional scoring signals. Treat this equation as conceptual guidance rather than a portable score contract.
Term-frequency Saturation
Repeating a term many times should not normally make a document proportionally more relevant forever.
Document A:
"database" appears once
Document B:
"database" appears ten times
Document C:
"database" appears one hundred times
Ranking expectation:
Additional occurrences can help,
but their contribution eventually provides
smaller incremental value.
Saturation reduces the benefit of excessive term repetition.
Document-length Normalization
A term match in a short focused document can provide stronger evidence than the same number of occurrences in a very long document.
Short document:
100 terms
"index" appears 3 times
Long document:
10,000 terms
"index" appears 3 times
Observation:
The term is more concentrated
in the short document.
Field Boosting
Matches in different fields can contribute different ranking weights.
Possible field importance:
Title:
High weight
Keywords:
Medium-high weight
Description:
Medium weight
Article body:
Standard weight
Search-document Example
{
"documentId": "tenant-17:article-981",
"title": "Full-Text Search",
"description": "Learn query analysis, ranking, and secure retrieval.",
"body": "Complete article content.",
"keywords": [
"search",
"BM25",
"ranking"
],
"language": "en",
"status": "published"
}
Excessive boosting can cause title keyword repetition or business metadata to dominate genuine relevance.
Business Boosting
Search results can combine textual relevance with approved business signals.
Possible signals include:
- Publication state
- Content freshness
- Course priority
- Verified quality
- Learner completion context
- Approved editorial promotion
Business boosts should be transparent, bounded, measurable, and separate from access-control decisions.
Structured Filters
Full-text retrieval is commonly combined with exact metadata filters.
Text query:
database indexing
Structured filters:
tenant_id = tenant-17
status = published
language = en
content_type = article
course_id = course-42
Filters usually determine eligibility rather than textual relevance.
Faceted Search
Facets summarize available result categories and allow users to narrow the result set.
Search:
system design
Content type:
Articles 42
Videos 18
Courses 7
Quizzes 5
Difficulty:
Beginner 31
Intermediate 24
Advanced 17
Facet counts must be computed within the caller's authorized result scope. Otherwise, a count can reveal the existence of restricted content.
Sorting
Search results can be sorted by relevance or a structured field.
Common options include:
- Most relevant
- Newest first
- Oldest first
- Title
- Popularity
- Course order
Sorting only by date can discard relevance ordering. Sorting only by relevance can make time-sensitive content difficult to discover. The user interface should communicate the active ordering.
Highlighting
Highlighting shows where query terms matched in a result.
Query:
inverted index
Result snippet:
An [inverted index] maps terms
to the documents containing them.
Highlighting can use indexed positions or offsets, reanalyze stored content, or apply provider-specific mechanisms.
Rendering rule: Search snippets are untrusted text. Encode snippets safely before rendering them in HTML, and allow only controlled highlighting markup.
Result Snippets
A snippet presents a short relevant portion of a larger document.
{
"documentId": "tenant-17:article-981",
"title": "Full-Text Search",
"snippet": "Full-text search retrieves and ranks documents...",
"score": 8.42,
"contentType": "article",
"publishedAt": "publication-time"
}
Snippets must not expose text from fields or documents the caller is not permitted to view.
Autocomplete
Autocomplete suggests terms, titles, or queries while the user types.
User enters:
full t
Suggestions:
full-text search
full-text index
full table scan
full transaction log
Autocomplete can use prefixes, edge-oriented tokens, curated suggestions, popular approved queries, or a dedicated suggestion index.
Avoid executing the complete expensive search pipeline after every keystroke. Apply minimum input length, debouncing, cancellation, and result limits.
Search-as-you-type
Keystroke
|
v
Wait for short debounce interval
|
v
Cancel obsolete request
|
v
Send bounded prefix query
|
v
Return limited suggestions
Suggestions should be tenant-aware and access-controlled when derived from private content.
Spell Correction
A search system can suggest a corrected query when an input term is rare or absent.
User query:
distrubuted systems
Suggestion:
Did you mean:
distributed systems?
A correction can be shown as a suggestion or applied automatically according to the product's search contract.
Correction rule: Preserve the user's original query and show when a correction was applied. Automatic correction can be harmful for codes, names, acronyms, and domain-specific terminology.
Query Expansion
Query expansion adds related terms to improve recall.
Original query:
ML
Expanded query:
ML
machine learning
Expansion sources can include:
- Approved synonyms
- Acronym dictionaries
- Alternate spellings
- Domain taxonomies
- Related controlled terms
Uncontrolled expansion can reduce precision by adding terms with a different meaning.
Precision and Recall
Precision measures how many returned results are relevant.
\[ Precision = \frac{ RelevantRetrievedDocuments }{ RetrievedDocuments } \]
Recall measures how many of the relevant documents were retrieved.
\[ Recall = \frac{ RelevantRetrievedDocuments }{ AllRelevantDocuments } \]
Broad synonym and fuzzy expansion can improve recall while reducing precision. Exact phrases and strict filters can improve precision while reducing recall.
Precision vs Recall Example
| Search Strategy | Likely Effect |
|---|---|
| Exact phrase only | Higher precision, lower recall |
| All terms required | More restrictive result set |
| Any term allowed | Broader result set with possible relevance noise |
| Synonym expansion | Can retrieve conceptually related terminology |
| Fuzzy expansion | Can recover typographical variations while adding candidates |
Search-quality Evaluation
A search system should be evaluated using representative queries and relevance judgments.
Create a test set containing:
- Common user queries
- Rare technical terms
- Acronyms
- Misspellings
- Phrase queries
- Multilingual queries
- Queries with no valid result
- Queries with restricted results
- Freshly published content
- Deleted or access-revoked content
Precision at K
Precision at \(K\) measures the proportion of relevant documents among the first \(K\) results.
\[ Precision@K = \frac{ RelevantDocumentsInTopK }{ K } \]
This is useful when users primarily inspect the first page of results.
Reciprocal Rank
Reciprocal rank rewards placing the first relevant result near the top.
\[ ReciprocalRank = \frac{ 1 }{ RankOfFirstRelevantResult } \]
Mean Reciprocal Rank averages this value across a query set.
NDCG Intuition
Normalized Discounted Cumulative Gain evaluates ranked results when relevance can have several levels.
It gives more value to highly relevant documents appearing near the top and discounts relevant documents appearing lower in the result list.
Use the exact formula and relevance scale defined by the selected evaluation framework.
Pagination
Search results need stable and efficient pagination.
Offset Pagination
Page 1:
Offset 0
Size 20
Page 2:
Offset 20
Size 20
Deep offsets can require the search system to identify and skip many earlier ranked results.
Cursor or Search-after Pagination
{
"lastScore": 7.81,
"lastDocumentId": "tenant-17:article-981",
"searchSnapshot": "search-context-id"
}
The next request continues after the last result using a deterministic ordering and provider-supported search context.
Pagination rule: Include a deterministic tiebreaker such as a stable document ID. Relevance scores can be identical, and the index can change between requests.
Search Freshness
Full-text indexes are commonly derived from an authoritative source. Therefore, a committed source change might not appear in search immediately.
Article transaction commits
|
v
Indexing event is published
|
v
Search consumer processes event
|
v
Index refresh completes
|
v
Article becomes searchable
Define freshness objectives for:
- New documents
- Document updates
- Deletions
- Publication changes
- Tenant ownership changes
- Access revocation
Security Trimming
Full-text search must not return documents the caller is not authorized to discover.
Search query
|
v
Apply trusted tenant scope
|
v
Apply publication status
|
v
Apply visibility rules
|
v
Apply course enrollment,
ownership, role, or group scope
|
v
Rank authorized candidates
|
v
Return results and facets
Security must cover:
- Document titles
- Search snippets
- Highlight fragments
- Facet counts
- Total result count
- Autocomplete suggestions
- Cached search responses
- Search analytics
Authorization rule: Filtering unauthorized documents after constructing visible results is too late. Apply trusted access scope before returning titles, snippets, counts, facets, or suggestions.
Tenant-aware Search
{
"query": "database indexing",
"trustedTenantId": "tenant-17",
"filters": {
"status": "published",
"language": "en"
}
}
The authenticated application context should supply the tenant scope. A client-provided tenant identifier alone must not be treated as authorization.
Search Index as a Derived Store
A search index is commonly optimized for retrieval rather than treated as the authoritative system of record.
Authoritative database or content store
|
v
Indexing pipeline
|
v
Search index
|
v
Search API
The architecture should support:
- Idempotent indexing
- Deletion propagation
- Source-version comparison
- Reconciliation
- Schema migration
- Analyzer migration
- Complete index rebuild
Indexing Events
{
"eventType": "ArticlePublished",
"tenantId": "tenant-17",
"sourceType": "article",
"sourceId": 981,
"sourceVersion": 12,
"occurredAt": "event-time"
}
The indexing consumer retrieves the authoritative content, builds the searchable document, and writes it using a stable search-document ID.
Idempotent Indexing
Repeated delivery of the same indexing event should not create duplicate search documents.
Stable search-document ID:
tenant-17:article:981
Source version:
12
Repeated event:
Update same search document
or return already-processed outcome.
An older source version should not overwrite a newer indexed version.
Reindexing
Reindexing rebuilds search documents from authoritative content.
Read authoritative documents
|
v
Create replacement index
|
v
Apply current schema and analyzers
|
v
Validate counts, queries, and security
|
v
Switch search alias or routing
|
v
Retire old index safely
Reindexing can be required after:
- Analyzer changes
- Synonym-strategy changes
- Search-document schema changes
- Field-type changes
- Major ranking changes
- Index corruption
- Authoritative-data correction
Index Versioning
Logical search name:
course-content
Physical index:
course-content-v1
Replacement index:
course-content-v2
After validation:
Logical search name
points to v2
Versioned indexes allow a replacement to be built and validated without destroying the currently serving index.
Search Architecture
Web or mobile client
|
v
Search API
|
+-- Authentication
+-- Authorization
+-- Query validation
+-- Rate limiting
+-- Tenant scope
|
v
Search cluster
|
+-- Text index
+-- Metadata filters
+-- Ranking
+-- Facets
|
v
Authorized ranked results
The search API prevents clients from bypassing trusted filters, exposes a stable contract, and centralizes resource limits.
Search API Request
POST /api/v1/search/articles HTTP/1.1
Content-Type: application/json
Authorization: Bearer access-token
{
"query": "full text search",
"language": "en",
"filters": {
"courseId": 42,
"contentType": "article"
},
"sort": "relevance",
"pageSize": 20,
"cursor": null
}
Search API Response
{
"query": "full text search",
"totalRelation": "exact-or-provider-defined",
"results": [
{
"documentId": "tenant-17:article:981",
"title": "Full-Text Search",
"snippet": "Learn how full-text search retrieves and ranks...",
"score": 8.42,
"contentType": "article",
"language": "en"
}
],
"facets": {
"contentType": [
{
"value": "article",
"count": 12
},
{
"value": "video",
"count": 5
}
]
},
"nextCursor": "opaque-cursor-value"
}
The cursor should be treated as an opaque server-generated value. Result counts can be exact or approximate according to the implementation and contract.
PHP Search-request Validation
<?php
declare(strict_types=1);
function validateSearchRequest(
array $request
): array {
$query =
trim(
(string)($request['query'] ?? '')
);
if ($query === '') {
throw new InvalidArgumentException(
'The search query is required.'
);
}
if (mb_strlen(
$query,
'UTF-8'
) > 300) {
throw new InvalidArgumentException(
'The search query is too long.'
);
}
$pageSize =
filter_var(
$request['pageSize'] ?? 20,
FILTER_VALIDATE_INT
);
if ($pageSize === false ||
$pageSize < 1 ||
$pageSize > 100) {
throw new InvalidArgumentException(
'The page size is invalid.'
);
}
$allowedSorts = [
'relevance',
'newest',
'oldest',
'title'
];
$sort =
strtolower(
trim(
(string)($request['sort'] ?? 'relevance')
)
);
if (!in_array(
$sort,
$allowedSorts,
true
)) {
throw new InvalidArgumentException(
'The sort option is invalid.'
);
}
return [
'query' => $query,
'pageSize' => $pageSize,
'sort' => $sort,
'cursor' =>
isset($request['cursor'])
? (string)$request['cursor']
: null
];
}
Tenant ID, user ID, roles, and permissions must come from authenticated server-side context rather than untrusted request fields.
Simplified PHP Search Example
<?php
declare(strict_types=1);
function searchArticles(
SearchClient $searchClient,
string $tenantId,
array $validatedRequest
): array {
return $searchClient->search(
index: 'course-content',
query: [
'text' =>
$validatedRequest['query'],
'filters' => [
'tenantId' =>
$tenantId,
'status' =>
'published'
],
'sort' =>
$validatedRequest['sort'],
'size' =>
$validatedRequest['pageSize'],
'cursor' =>
$validatedRequest['cursor']
]
);
}
SearchClient is a conceptual abstraction. Use the official
supported client and exact query syntax of the selected search platform.
Query Safety
User-provided search syntax can create expensive or unintended queries.
Apply limits to:
- Query length
- Number of terms
- Boolean nesting depth
- Wildcard expansion
- Fuzzy edit distance
- Prefix length
- Synonym expansion
- Facet count
- Page size
- Highlight fragments
- Search timeout
Prefer a structured search request over directly exposing unrestricted provider query syntax to untrusted clients.
Search Timeouts
Request deadline
|
+-- Authentication time
+-- Search-API processing
+-- Search-cluster execution
+-- Result transformation
+-- Network response
Keep the search-engine timeout within the complete API deadline. A timed-out query should not continue consuming unbounded server resources.
Rate Limiting
Search endpoints can be abused through automated broad queries, autocomplete storms, deep pagination, large facet requests, and expensive wildcard or fuzzy expressions.
Rate limits can be scoped by:
- User
- Tenant
- Application
- Endpoint
- Query category
- Anonymous network identity where appropriate
Sharded Search
A large search index can be divided into shards.
Search request
|
v
Coordinator
|
+-- Query Shard 1
+-- Query Shard 2
+-- Query Shard 3
|
v
Collect shard candidates
|
v
Merge and rank
|
v
Return top results
Distributed ranking requires care because term and document statistics can differ across shards.
Search Replicas
Search shards can have replicas for read capacity and failure recovery.
Shard A primary
|
+-- Replica A1
|
+-- Replica A2
Replication can increase query capacity but also consumes storage and indexing resources. Replica freshness and health should be monitored.
Search Caching
Search systems can cache repeated query components or final result pages.
Cache keys must account for:
- Normalized query
- Tenant scope
- Authorization scope
- Filters
- Sort order
- Language
- Page or cursor
- Index version
Cache-security rule: Never share a search response across users or tenants unless the cache key and authorization design prove that every recipient has the same permitted result set.
Search Analytics
Search analytics can reveal how users interact with the search experience.
Useful signals include:
- Search-query count
- Zero-result queries
- Low-result queries
- Query reformulations
- Result clicks
- Click position
- Abandoned searches
- Applied filters
- Spelling suggestions accepted
- Search latency
Search queries can contain personal, confidential, or sensitive information. Apply privacy, retention, access-control, and redaction requirements to search logs and analytics.
Zero-result Queries
A zero-result query can indicate:
- The requested content does not exist
- The content is not yet indexed
- The caller is not authorized
- The analyzer produced unexpected terms
- A synonym or acronym is missing
- The query contains a spelling error
- Filters are too restrictive
- Content is in another language
Do not expose whether restricted results exist when the caller is not authorized to discover them.
Relevance Testing
Create a controlled relevance test set.
| Query | Expected High-ranking Result | Reason |
|---|---|---|
| inverted index | Inverted Indexes article | Exact title and body match |
| DBMS transaction reliability | ACID article | Domain synonym and subject match |
| databse indexing | Database indexing lesson | Approved typo tolerance |
| "full text search" | Full-Text Search article | Exact phrase match |
| large video upload | Uploads and Multipart Transfer article | Concept and content match |
Lexical, Vector, and Hybrid Search
| Search Type | Primary Matching Basis | Typical Strength |
|---|---|---|
| Lexical full-text | Analyzed terms and inverted-index statistics | Exact terms, phrases, identifiers, filters, and explainable keyword relevance |
| Vector search | Distance between vector representations | Semantic similarity and paraphrased concepts |
| Hybrid search | Combination of lexical and vector retrieval | Balances exact terminology with semantic similarity |
Full-text search remains important when users search for exact names, codes, quoted phrases, technical terminology, or rare identifiers.
Hybrid Retrieval Flow
User query
|
+-- Lexical retrieval
| |
| +-- Exact terms
| +-- Phrases
| +-- Keyword scoring
|
+-- Vector retrieval
|
+-- Semantic similarity
|
v
Combine candidate sets
|
v
Apply security filters
|
v
Rerank and return results
Candidate combination and score normalization depend on the selected search architecture.
Search Observability
Useful full-text search metrics include:
- Search-request rate
- Search latency by percentile
- Timeout rate
- Error rate
- Zero-result rate
- Result count distribution
- Click-through rate
- First-result click rate
- Query reformulation rate
- Autocomplete latency
- Facet latency
- Indexing lag
- Stale-document count
- Deleted documents still searchable
- Shard and replica health
- Cache hit rate
- Security-filter failures
Alert Conditions
Alert when:
- Search latency exceeds its objective
- Timeouts or errors increase
- Indexing lag grows
- Zero-result queries increase unexpectedly
- One shard receives disproportionate traffic
- Replicas become unavailable
- Deleted or restricted content remains searchable
- Autocomplete or facet requests consume excessive resources
- Relevance metrics regress after an analyzer or ranking change
Search Troubleshooting Workflow
- Capture the exact query, filters, language, and trusted tenant scope.
- Confirm the expected document exists in the authoritative source.
- Confirm the document is published and authorized.
- Confirm the current document version is indexed.
- Analyze the query into terms.
- Analyze the expected document field using the same analyzer.
- Compare query and indexed terms.
- Check synonyms, stop words, stemming, and fuzzy settings.
- Check structured and security filters.
- Inspect the document's relevance score components.
- Check indexing lag and failed indexing events.
- Check shard, replica, timeout, and resource health.
- Reindex or repair derived search state when required.
Common Full-text Search Mistakes
Using SQL LIKE as the Complete Search Engine
LIKE does not provide the complete language analysis, ranking, phrase, synonym, fuzzy, facet, and highlighting capabilities expected from full-text search.
Using Different Index and Query Analyzers
The query can produce terms that do not match the terms stored during indexing.
Applying One Analyzer to Every Language
Tokenization, normalization, stemming, and common-word behaviour vary between languages.
Expanding Every Query Aggressively
Broad synonyms, prefixes, wildcards, and fuzzy terms can reduce precision and increase query cost.
Ranking Only by Term Count
Raw term frequency ignores term rarity, field importance, document length, proximity, and business requirements.
Boosting Freshness Too Strongly
A recent but weakly related document can outrank an older highly relevant document.
Returning Unauthorized Facet Counts
Counts can disclose the existence and classification of restricted documents.
Rendering Highlights without Safe Encoding
Indexed document content is untrusted and can create unsafe browser output when rendered directly.
Using Deep Offset Pagination
The engine can perform significant work to identify and discard previous ranked results.
Ignoring Indexing Lag
Users can see missing updates, stale titles, old permissions, or deleted content.
Changing Ranking without Relevance Tests
A change that improves one query can reduce quality across other important query categories.
Treating the Search Index as the Only Data Source
The index is commonly a derived representation and should be rebuildable from authoritative content and metadata.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Exact keyword | Documents containing the analyzed term are returned |
| Case variation | Approved uppercase and lowercase forms behave consistently |
| Multi-term AND query | Only documents satisfying all required terms remain |
| Phrase query | Terms occur in the required order and proximity |
| Synonym query | Approved equivalent terminology retrieves expected documents |
| Typographical error | Bounded fuzzy or correction logic produces the documented behaviour |
| Structured filter | Only matching language, type, course, or status remains |
| Field boosting | A strong title match receives the designed relevance treatment |
| Highlighting | Safe snippets identify the matched terms |
| Cursor pagination | Results continue without unintended duplicates or omissions |
| Tenant isolation | Another tenant's results, counts, facets, and suggestions remain hidden |
| Access revocation | Restricted content stops appearing within the defined objective |
| Deleted document | The document and its snippets no longer appear |
| Index rebuild | Document counts, relevance tests, and security filters remain correct |
Full-text Search Best Practices
Recommended Practices
- Use an inverted index for scalable analyzed text retrieval.
- Keep authoritative content separate from the derived search index.
- Use compatible index-time and query-time analyzers.
- Select analyzers according to content language.
- Use approved synonym and acronym dictionaries.
- Bound fuzzy, wildcard, prefix, and Boolean expansion.
- Combine textual retrieval with structured metadata filters.
- Apply trusted tenant and authorization scope before returning any search information.
- Use stable search-document IDs and authoritative source versions.
- Make indexing operations idempotent.
- Reject stale out-of-order indexing events.
- Use relevance ranking rather than raw term count alone.
- Test title, body, phrase, synonym, and typo behaviour separately.
- Use deterministic cursor-based pagination for deep result navigation.
- Encode snippets and highlights safely.
- Define freshness objectives for create, update, delete, and access changes.
- Maintain a tested full-index rebuild procedure.
- Protect search logs and analytics as potentially sensitive data.
- Measure precision, recall, ranking quality, latency, and zero-result queries.
- Evaluate ranking changes against a representative relevance test set.
Practice Exercise
Design full-text search for articles, courses, videos, and lessons on your online learning platform.
Requirements
- Create a stable search-document ID for each searchable item.
- Index title, description, body, keywords, language, and content type.
- Use exact fields for tenant, status, course, and visibility filters.
- Use compatible index-time and query-time analyzers.
- Support keyword and quoted phrase search.
- Support AND and OR logic.
- Add approved synonyms for course terminology and acronyms.
- Add bounded typo tolerance.
- Boost title matches above body-only matches.
- Return safely encoded snippets.
- Return authorized facet counts for content type and difficulty.
- Support relevance and newest-first sorting.
- Support cursor-based pagination.
- Prevent cross-tenant search disclosure.
- Propagate publication, deletion, and access changes.
- Measure ranking quality using representative queries.
- Track zero-result queries and accepted corrections.
- Test a complete index rebuild.
Search-design Template
| Feature | Design Decision | Validation |
|---|---|---|
| Searchable fields | Title, description, body, and approved keywords | Verify expected terms are indexed |
| Filter fields | Tenant, course, type, language, status, and visibility | Verify exact filtering and security trimming |
| Analyzer | Language-aware tokenization and normalization | Compare index and query tokens |
| Ranking | Text relevance plus bounded field and business boosts | Evaluate representative ranked queries |
| Typos | Bounded correction or fuzzy expansion | Test valid words, codes, and mistakes |
| Pagination | Opaque cursor with deterministic tiebreaker | Check duplicate and missing results |
| Freshness | Idempotent event-driven indexing | Measure create, update, and deletion delay |
| Recovery | Rebuild from authoritative content | Validate counts, relevance, and security after rebuild |
Frequently Asked Questions
What is full-text search?
Full-text search retrieves documents by analyzing textual content and ranking candidate documents according to relevance.
How is full-text search different from SQL LIKE?
Full-text search uses analyzed terms and an inverted index and can provide relevance ranking, phrase matching, language processing, synonyms, and highlighting.
What is relevance ranking?
Relevance ranking estimates how well each candidate document satisfies the user query and orders results accordingly.
What is BM25?
BM25 is a text-ranking approach that considers term rarity, term-frequency saturation, and document-length normalization.
What is phrase search?
Phrase search requires query terms to occur in the requested order and proximity.
What is fuzzy search?
Fuzzy search permits bounded character differences between query terms and indexed terms to support selected spelling variations or mistakes.
What are facets?
Facets summarize result categories and allow users to narrow a search using structured fields.
What is search highlighting?
Highlighting marks matching terms inside a safely rendered result snippet.
Why can a new article be missing from search?
The indexing event may not have been processed, the index may not have refreshed, analysis may differ, or filters and authorization may exclude the document.
Should the search index be the authoritative database?
It is commonly a derived and rebuildable retrieval structure backed by an authoritative source of content and metadata.
What is hybrid search?
Hybrid search combines lexical full-text retrieval with vector-based semantic retrieval.
What comes after full-text search?
Full-text search completes the Storage, Files, Objects & Search Basics module.
Key Takeaway
Full-text search transforms a user's text into analyzed terms, retrieves candidate documents through an inverted index, applies structured and security filters, calculates relevance, and returns ranked results with snippets, highlights, facets, and pagination. Use language-compatible analyzers, controlled synonyms, bounded fuzzy and wildcard expansion, and relevance models that account for term rarity, frequency saturation, field importance, and document length. Treat the search index as a derived and rebuildable representation, process updates idempotently, reject stale events, and define freshness requirements for publication, deletion, and access changes. Most importantly, apply trusted tenant and authorization scope before returning documents, snippets, counts, facets, suggestions, or cached results.