Crawling and ingestion
Crawling and Ingestion
Learn how search engines and knowledge systems discover content, fetch pages and documents, extract useful information, normalize metadata, remove duplicates, create chunks, generate embeddings, and index processed content. Understand crawl frontiers, URL normalization, robots rules, sitemaps, rate limits, dynamic content, incremental crawling, parsing, schema validation, retries, dead-letter queues, backpressure, freshness, and observability.
Introduction
A search engine cannot answer questions about content that it has not discovered, processed, and indexed.
Content sources
|
v
Discovery and crawling
|
v
Content extraction
|
v
Normalization and validation
|
v
Deduplication and chunking
|
v
Embedding and indexing
|
v
Search and retrieval
Crawling and ingestion are related but different responsibilities.
- Crawling discovers and fetches content from connected sources.
- Ingestion converts fetched content into validated, searchable records.
Core idea: Crawling answers, “What content exists and may be fetched?” Ingestion answers, “How should that content be transformed, governed, and stored so downstream search can use it?”
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | HTTP, URLs, and DNS | A crawler retrieves resources using web protocols and addresses. |
| 2 | HTML and document formats | Content can arrive as HTML, PDF, Word, PowerPoint, Markdown, CSV, or Excel. |
| 3 | Queues and asynchronous processing | URL fetching, parsing, chunking, and indexing are commonly separated into stages. |
| 4 | Retries, DLQs, and backpressure | Remote sources and downstream indexes can fail or become overloaded. |
| 5 | Search indexes | The final processed content must be written into a retrieval structure. |
| 6 | Embeddings and vector search | Semantic retrieval can require vector representations of content chunks. |
| 7 | Security and data governance | Access controls, retention, provenance, and sensitive data must survive ingestion. |
Crawling vs Scraping vs Ingestion
| Concept | Primary Responsibility | Example Output |
|---|---|---|
| Crawling | Discover and fetch pages or documents | Fetched HTML, PDF, or source metadata |
| Scraping | Extract selected fields from fetched content | Title, author, price, headings, or article body |
| Parsing | Convert source bytes into structured content | Text blocks, tables, images, and metadata |
| Ingestion | Validate, transform, govern, and load content | Search documents, chunks, vectors, and lineage records |
| Indexing | Create retrieval structures from processed content | Keyword index, vector index, or hybrid-search document |
End-to-End Architecture
Seeds and connectors
|
v
URL discovery
|
v
Crawl frontier
|
v
Fetcher
|
v
Raw-content store
|
v
Parser and extractor
|
v
Normalizer and validator
|
v
Deduplicator
|
v
Chunker
|
+-- Keyword indexing
|
+-- Embedding generation
|
v
Vector indexing
|
v
Search and RAG
Content Discovery
Discovery begins with one or more trusted seeds.
Common discovery sources include:
- Configured starting URLs
- XML sitemaps
- Links extracted from fetched pages
- RSS or Atom feeds
- CMS and repository connectors
- Object-storage notifications
- Database change feeds
- User-uploaded files
Seed URL
|
v
Fetch page
|
v
Extract allowed links
|
v
Normalize discovered URLs
|
v
Add unseen URLs
to crawl frontier
Crawl Frontier
The crawl frontier is the managed collection of URLs that are waiting to be fetched.
A useful frontier records:
- Normalized URL
- Source and domain
- Discovery time
- Priority
- Next eligible fetch time
- Attempt count
- Previous fetch status
- Content-change hints
- Applicable access policy
URL Normalization
The same logical resource can appear through several syntactically different URLs.
Possible forms:
https://example.test/course/42
https://example.test/course/42/
https://example.test/course/42?tracking=campaign
https://EXAMPLE.test/course/42
A normalization policy can consider:
- Protocol and hostname normalization
- Default ports
- Fragment removal
- Trailing-slash policy
- Query-parameter filtering
- Canonical URL declarations
- Redirect targets
Query parameters must not be removed blindly because some parameters identify genuinely different resources.
Robots Rules and Sitemaps
A crawler should evaluate the permissions and discovery guidance applicable to the source.
robots.txtcommunicates crawler-access directives.sitemap.xmlcan provide a list of discoverable URLs.
Authorization to access content must come from the applicable legal, contractual, organizational, and technical permissions. A robots file is not a substitute for authentication or data-usage authorization.
Politeness and Rate Limiting
An uncontrolled crawler can overload the source system.
Before fetch:
- Check domain concurrency
- Check crawl delay
- Check recent failures
- Check server guidance
- Check retry budget
- Check authorization
Useful controls include:
- Per-domain request limits
- Per-host concurrency limits
- Connection and response timeouts
- Capped exponential backoff with jitter
- Respect for valid server retry guidance
- Circuit breakers for failing sources
- Bounded crawl queues
Fetching
The fetcher retrieves a resource and records enough evidence to process and audit it.
Useful fetch metadata includes:
- Requested URL
- Final URL after redirects
- HTTP status
- Content type
- Content length
- Response timestamp
- ETag or modification metadata when available
- Content hash
- Source identity
Static vs Dynamic Content
| Content Type | Fetch Direction |
|---|---|
| Server-rendered HTML | Use an ordinary HTTP fetcher where sufficient |
| JavaScript-rendered content | Use an approved browser-rendering path when required |
| API-backed content | Prefer an authorized API or connector where available |
| Files and documents | Download and route according to MIME type and parser support |
| Authenticated content | Use an approved identity and preserve source permissions |
Connector rule: Prefer an authorized source connector or API when it provides more reliable content, metadata, permissions, and change tracking than HTML scraping.
Parsing and Extraction
Parsing converts downloaded content into a structured representation.
Extracted elements can include:
- Page title and headings
- Main body text
- Tables and lists
- Links and canonical metadata
- Document properties
- Image captions and approved image-derived text
- Code blocks
- Language and publication metadata
The internal IAI4E_KB_Ingestion_Agent_Ingestion_Tutorial_1_1.pdf describes a multi-format ingestion service that processes PDF, Word, PowerPoint, HTML, Markdown, CSV, and Excel content through parsing, chunking, embedding, and indexing stages. 【1-73f1a5】
Content Cleaning
Extracted text can contain repeated navigation, cookie banners, headers, footers, menus, legal notices, advertisements, and hidden elements.
Raw page
|
v
Remove repeated boilerplate
|
v
Preserve meaningful structure
|
v
Normalize text and metadata
|
v
Produce canonical document
Cleaning must not remove meaningful warnings, qualifications, citations, or section boundaries.
Deduplication
Duplicate content wastes storage and can cause repeated or biased search results.
| Duplicate Type | Example |
|---|---|
| URL duplicate | Different URL spellings resolve to the same page |
| Exact-content duplicate | Two files have the same content hash |
| Near duplicate | Pages differ only by navigation or a small timestamp |
| Version duplicate | An older copy remains beside a newer authoritative version |
| Chunk duplicate | Repeated boilerplate appears in many indexed chunks |
Chunking
Long documents are commonly divided into smaller retrieval units called chunks.
Document
|
+-- Chunk 1: Introduction
+-- Chunk 2: Architecture
+-- Chunk 3: Configuration
+-- Chunk 4: Failure handling
+-- Chunk 5: Troubleshooting
A chunk should preserve:
- Document ID
- Chunk ID
- Heading path
- Source URL or file reference
- Content position
- Permissions
- Language
- Version and ingestion timestamp
Chunk-size Trade-off
| Chunk Direction | Advantage | Risk |
|---|---|---|
| Small chunks | More focused matching | Important context can be separated |
| Large chunks | More surrounding context | Matches can become less precise and more expensive |
| Overlapping chunks | Preserves context across boundaries | Creates duplicate indexed text |
| Structure-aware chunks | Preserves sections, headings, or table boundaries | Requires format-aware parsing |
Embeddings and Indexing
Processed chunks can be indexed for keyword retrieval, vector retrieval, or hybrid retrieval.
Clean chunk
|
+-- Text and metadata
| |
| v
| Keyword index
|
+-- Embedding model
|
v
Vector
|
v
Vector index
The internal IAI4E_KB_Ingestion_Agent_Ingestion_Tutorial_1_1.pdf describes extraction, chunking, embedding, and indexing into Azure AI Search for downstream semantic vector search. 【1-73f1a5】
Full vs Incremental Ingestion
| Mode | Purpose |
|---|---|
| Full ingestion | Process the complete approved source |
| Incremental ingestion | Process only new or changed content |
| Parse only | Preview extracted content without writing to the index |
| Reindex | Rebuild the index after schema, parser, chunker, or embedding changes |
| Deletion synchronization | Remove or deactivate indexed content that no longer exists or is no longer authorized |
Incremental Recrawling
Incremental crawling avoids downloading and processing every document during every run.
Previously observed metadata
|
v
Check source for change
|
+-- Unchanged:
| skip expensive processing
|
+-- Changed:
fetch, parse,
validate, and reindex
Change signals can include:
- Source-provided modification metadata
- ETag or conditional-fetch metadata
- Content hash
- Repository change token
- Database Change Data Capture
- Object-storage event
- Explicit version number
Deleted and Revoked Content
An ingestion system must account for content that was removed, renamed, expired, or made inaccessible.
Previously indexed document
|
v
Source document deleted
or permission revoked
|
v
Deletion synchronization
|
v
Remove, deactivate,
or restrict indexed content
Failing to synchronize deletions can expose stale or unauthorized content in search results.
Security Trimming
Search results must not reveal content that the requesting user is not permitted to access.
Preserve:
- Source identity
- Tenant identity
- Document permissions
- Group or role restrictions
- Classification labels
- Retention and deletion policy
- Permission-change timestamps
Security rule: Do not make private content public merely because the crawler could technically fetch it. Source authorization and retrieval authorization must be preserved end to end.
Backpressure
Fetchers can retrieve content faster than parsers, embedding services, or search indexes can process it.
Crawl rate
|
v
Raw-document queue grows
|
v
Parser backlog grows
|
v
Embedding service throttles
|
v
Index freshness decreases
Useful controls include:
- Bounded crawl and ingestion queues
- Per-stage concurrency limits
- Batching
- Source rate limiting
- Embedding-service quotas
- Index-write backoff
- Priority for high-value or recently changed content
- Autoscaling within downstream capacity
Retries and DLQs
Ingestion stage fails
|
v
Classify failure
|
+-- Temporary:
| retry with backoff
|
+-- Permanent:
| route to DLQ
|
+-- Unknown:
bounded retry
then DLQ
Examples include:
- Retry a temporary source timeout.
- Retry temporary embedding throttling.
- Dead-letter a corrupt document.
- Dead-letter an unsupported schema after validation.
- Quarantine content that violates the ingestion policy.
Conceptual Pipeline Configuration
crawlingAndIngestion:
discovery:
seeds:
- approved-source-root
sitemapDiscovery: enabled
linkDiscovery: enabled
crawling:
robotsPolicy: respected
perHostConcurrency: approved-limit
requestTimeout: approved-boundary
retries:
bounded: true
backoff: exponential-with-jitter
extraction:
supportedTypes:
- html
- pdf
- docx
- pptx
- markdown
- csv
- xlsx
normalization:
urlCanonicalization: enabled
contentHashing: enabled
exactDeduplication: enabled
chunking:
strategy: structure-aware
maximumSize: approved-limit
overlap: workload-specific
indexing:
keywordIndex: enabled
vectorIndex: enabled
preservePermissions: true
failureHandling:
deadLetterDestination: ingestion-failures
observability:
crawlRate: enabled
fetchFailures: enabled
ingestionLag: enabled
indexingFailures: enabled
freshness: enabled
Python Pipeline Skeleton
from dataclasses import dataclass
from typing import Iterable
@dataclass(frozen=True)
class CrawledDocument:
source_url: str
content_type: str
content: bytes
content_hash: str
@dataclass(frozen=True)
class SearchChunk:
document_id: str
chunk_id: str
text: str
metadata: dict
def ingest_document(
document: CrawledDocument
) -> Iterable[SearchChunk\]:
parsed_document = parse_by_content_type(
document.content_type,
document.content,
)
normalized_document = normalize_document(
parsed_document
)
validate_document(
normalized_document
)
for chunk in create_structure_aware_chunks(
normalized_document
):
yield SearchChunk(
document_id=normalized_document.document_id,
chunk_id=chunk.chunk_id,
text=chunk.text,
metadata={
"source_url": document.source_url,
"heading_path": chunk.heading_path,
"content_hash": document.content_hash,
},
)
This educational example omits access control, robots evaluation, network safety, retries, DLQs, malware scanning, parser isolation, embedding, persistence, deletion synchronization, and index-write guarantees.
Ingestion Tracking Table
CREATE TABLE ingestion_documents
(
document_id VARCHAR(150) PRIMARY KEY,
source_uri TEXT NOT NULL,
source_type VARCHAR(50) NOT NULL,
content_hash VARCHAR(150) NOT NULL,
ingestion_status VARCHAR(30) NOT NULL,
schema_version INTEGER NOT NULL,
discovered_at TIMESTAMP NOT NULL,
fetched_at TIMESTAMP NULL,
indexed_at TIMESTAMP NULL,
deleted_at TIMESTAMP NULL,
failure_code VARCHAR(100) NULL
);
Freshness
Search freshness is the delay between a source change and the corresponding searchable index update.
A simplified measure is:
\[ FreshnessLag = IndexedAt - SourceChangedAt \]
When a reliable source-change timestamp is unavailable, freshness must be estimated using available discovery, fetch, and index timestamps.
Learning-platform Example
Sources:
Course articles
Tutorial pages
PDF notes
Code examples
Quiz explanations
Pipeline:
Discover approved content
|
v
Fetch changed documents
|
v
Extract title, headings, and body
|
v
Remove navigation and repeated footer
|
v
Create section-aware chunks
|
v
Generate keyword and vector records
|
v
Update searchable index
Observability
Useful metrics include:
- URLs discovered
- URLs pending in the frontier
- Fetch success and failure rate
- Responses by status and content type
- Bytes downloaded
- Robots or policy exclusions
- Parse failures by file type
- Documents and chunks created
- Exact and near duplicates
- Embedding requests and failures
- Index writes and failures
- Ingestion backlog and oldest-item age
- Source-to-index freshness
- Deleted-document synchronization lag
- DLQ depth and oldest failure age
Alert Conditions
- Frontier depth continues growing
- One source returns repeated failures
- Parse failures increase for one format
- Embedding or indexing becomes throttled
- Freshness lag exceeds the search objective
- Deletion synchronization stops progressing
- Duplicate-content rate increases unexpectedly
- One source generates unlimited URL variations
- DLQ depth or failure age increases
- Indexed permissions differ from source permissions
Common Mistakes
Crawling without URL Normalization
The frontier repeatedly fetches the same logical resource through different URLs.
Ignoring Source Policies
The crawler fetches content outside the approved access and usage scope.
Using Browser Rendering for Every Page
Processing becomes unnecessarily expensive when ordinary HTTP fetching is sufficient.
Indexing Raw HTML
Navigation, scripts, banners, and repeated boilerplate reduce retrieval quality.
Using One Chunk for a Large Document
Search matches become broad and relevant passages are difficult to isolate.
Discarding Document Structure
Chunks lose heading, section, table, and source context.
Ignoring Duplicate Content
Storage costs grow and repeated content can dominate retrieval results.
Performing Full Ingestion Every Time
Unchanged content repeatedly consumes fetching, parsing, embedding, and indexing capacity.
Ignoring Deleted Content
Search continues returning stale or unauthorized documents.
Dropping Source Permissions
Indexed content can become visible to users who could not access the source.
Retrying Corrupt Files Indefinitely
Permanent parsing failures consume ingestion capacity without progress.
Monitoring Index Size but Not Freshness
The index appears healthy while recent source changes remain unavailable.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Duplicate URL forms | They resolve to one canonical crawl identity. |
| Disallowed path | The crawler excludes it according to the approved policy. |
| Redirect chain | The final canonical source is recorded safely. |
| JavaScript page | The approved rendering path extracts required content. |
| Corrupt PDF | Bounded retries route the document to the DLQ. |
| Unchanged document | Incremental ingestion avoids unnecessary reprocessing. |
| Changed document | New chunks replace or version the previous searchable representation. |
| Deleted source | Indexed content is removed or deactivated. |
| Permission change | Search access reflects the updated source authorization. |
| Embedding throttling | Backoff and bounded concurrency protect the service. |
| Repeated boilerplate | Cleaning and deduplication prevent repeated low-value chunks. |
| Pipeline replay | Idempotent writes prevent duplicate indexed documents. |
Best Practices
Recommended Practices
- Define approved sources and content-usage boundaries.
- Prefer authorized connectors and APIs where available.
- Respect applicable crawl and access policies.
- Normalize and deduplicate URLs before adding them to the frontier.
- Apply per-host rate and concurrency limits.
- Store raw content and provenance when governance requires it.
- Parse according to verified content type.
- Remove boilerplate without removing meaningful context.
- Use exact-content and near-duplicate detection.
- Create structure-aware chunks.
- Preserve document, section, version, and source metadata.
- Preserve source permissions through indexing and retrieval.
- Use incremental ingestion for new and changed content.
- Synchronize deletions and permission changes.
- Make index writes idempotent.
- Use bounded retries with backoff and jitter.
- Route permanent failures to a controlled DLQ.
- Apply backpressure between pipeline stages.
- Monitor freshness as well as ingestion throughput.
- Test parsing, deduplication, deletion, permissions, and recovery.
Frequently Asked Questions
What is web crawling?
Web crawling is the automated discovery and fetching of permitted web resources.
What is data ingestion?
Data ingestion validates, transforms, governs, and loads source content into a downstream system.
What is a crawl frontier?
A crawl frontier is the managed collection of discovered URLs waiting for an eligible fetch.
Why is URL normalization important?
It prevents syntactically different representations of one resource from being crawled and indexed repeatedly.
What is incremental ingestion?
Incremental ingestion processes only new, changed, deleted, or permission-modified content.
Why is content chunked?
Chunking creates smaller retrieval units that can match focused queries while retaining useful context.
What is ingestion freshness?
Freshness describes how quickly a source change becomes available in the searchable index.
Why preserve source metadata?
Source metadata supports provenance, citations, permissions, deletion, versioning, and troubleshooting.
Should every page use browser rendering?
No. Use browser rendering only when required content is unavailable through the approved ordinary fetch or connector path.
How are failed documents handled?
Temporary failures use bounded retries, while permanent parsing or validation failures follow a controlled DLQ workflow.
Why synchronize deletions?
Without deletion synchronization, search can continue returning stale or unauthorized content.
What is a good ingestion architecture?
A good architecture separates discovery, fetching, raw storage, parsing, normalization, deduplication, chunking, embedding, indexing, and failure handling into observable and bounded stages.
Key Takeaway
Crawling discovers and retrieves approved content, while ingestion converts that content into governed and searchable records. Start from trusted seeds, connectors, feeds, or sitemaps. Normalize and deduplicate URLs, respect source policies, and protect websites with per-host rate and concurrency enance and fetch metadata, parse content according to its verified format, remove low-value boilerplate, and preserve headings, tables, permissions, versions, and source references. Use structure-aware chunks for retrieval and create keyword, vector, or hybrid indexes according to search requirements. Prefer incremental ingestion over repeatedly processing unchanged content, and synchronize deletions and permission changes so stale data does not remain searchable. Finally, make every stage idempotent, use bounded retries and DLQs, apply backpressure, and monitor crawl coverage, failures, backlog age, indexing progress, and freshness.