metadata
Metadata
Learn how metadata describes files, objects, blocks, and business assets; supports discovery, integrity, authorization, lifecycle management, auditing, and search; and must be modeled, validated, indexed, protected, synchronized, and governed separately from the underlying content.
Introduction
Applications rarely manage content as anonymous bytes. A stored video needs a title, content type, size, owner, checksum, processing state, access classification, and retention policy. A document can require a course, chapter, lesson, language, version, and publication status.
This descriptive and operational information is called metadata.
Metadata helps a system understand what stored content represents, where it belongs, how it should be processed, who may access it, how long it should remain, and whether it passed integrity verification.
Core idea: Data is the stored content. Metadata describes, identifies, classifies, governs, and connects that content to the application domain.
In your System Design curriculum, Metadata is Topic 6.4 under Storage, Files, Objects & Search Basics. It follows checksums and precedes uploads and multipart transfer, inverted indexes, and full-text search.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Block, file, and object storage | Each storage model maintains different forms of descriptive and operational metadata. |
| 2 | Entities and relationships | Application metadata commonly connects stored content to business entities. |
| 3 | Keys and constraints | Metadata records require stable identities, uniqueness, validation, and referential integrity. |
| 4 | Checksums | Checksum algorithm, value, and verification status are important integrity metadata. |
| 5 | Security fundamentals | Metadata can contain sensitive information and influence access-control decisions. |
| 6 | Search fundamentals | Metadata provides structured fields for filtering, faceting, ranking, and discovery. |
What Is Metadata?
Metadata is data that describes another data item, resource, object, file, event, or dataset.
Stored content:
lesson-07.mp4
Metadata:
- Asset ID
- Object key
- Original filename
- Content type
- Content length
- Course ID
- Lesson ID
- Language
- Checksum
- Upload time
- Processing status
- Visibility
- Retention class
The video bytes contain the media content. The metadata describes how the application should identify, organize, validate, process, protect, and present that content.
Data vs Metadata
| Area | Data | Metadata |
|---|---|---|
| Video asset | Encoded audio and video bytes | Title, duration, content type, checksum, and lesson ID |
| PDF document | Document bytes and embedded content | Filename, author, page count, language, and publication state |
| Database table | Rows stored in the table | Column names, data types, constraints, indexes, and ownership |
| Image | Pixel and encoding data | Dimensions, format, orientation, checksum, and alt text |
| Backup | Copied application or database content | Backup time, source, checksum, retention, and recovery status |
Categories of Metadata
Metadata can be grouped according to its purpose.
| Category | Purpose | Examples |
|---|---|---|
| Descriptive | Explains what the content is | Title, description, keywords, language, and subject |
| Structural | Explains how components relate | Course, chapter, lesson, page order, and media segments |
| Administrative | Supports management and governance | Owner, visibility, retention class, and storage tier |
| Technical | Describes technical properties | Content type, size, dimensions, codec, and checksum |
| Operational | Tracks processing and lifecycle state | Pending, verified, scanning, published, archived, and failed |
| Provenance | Records origin and transformation history | Uploader, source system, creation time, and processing version |
| Security | Supports protection and access decisions | Classification, tenant ID, access scope, and legal hold |
Descriptive Metadata
Descriptive metadata helps users and applications understand and discover content.
{
"title": "Introduction to System Design",
"description": "Foundational concepts for designing scalable systems.",
"language": "en",
"contentType": "video/mp4",
"keywords": [
"system design",
"scalability",
"architecture"
]
}
Descriptive fields can later feed catalog search, filters, recommendations, accessibility features, and user-interface labels.
Structural Metadata
Structural metadata describes relationships among content components.
Course
|
+-- Chapter 1
| |
| +-- Lesson 1
| +-- Lesson 2
|
+-- Chapter 2
|
+-- Lesson 3
+-- Lesson 4
Structural metadata can define:
- Parent-child relationships
- Display order
- Document page sequence
- Video segments
- Related attachments
- Alternative renditions
- Course-module composition
Technical Metadata
Technical metadata describes properties required to store, validate, deliver, or process content correctly.
| Content | Technical Metadata |
|---|---|
| Image | Format, dimensions, color profile, orientation, and size |
| Video | Codec, duration, resolution, frame rate, and bitrate |
| Audio | Codec, sample rate, channel count, duration, and bitrate |
| Document | Format, page count, character encoding, and content length |
| Object | Key, version, content type, length, checksum, and storage class |
Administrative Metadata
Administrative metadata supports ownership, governance, lifecycle, and operational management.
{
"tenantId": "tenant-17",
"ownerId": "course-team-42",
"visibility": "private",
"retentionClass": "course-content",
"storageTier": "standard",
"publicationStatus": "published",
"legalHold": false
}
Administrative metadata can influence important system behaviour, so its accepted values and update permissions should be controlled carefully.
Provenance Metadata
Provenance metadata records where content came from and how it changed.
Original upload
|
v
Malware scan
|
v
Metadata extraction
|
v
Video transcoding
|
v
Thumbnail generation
|
v
Publication
Useful provenance fields can include:
- Source system
- Ingestion time
- Uploader or service identity
- Original asset ID
- Transformation name
- Transformation version
- Parent object version
- Processing time
- Processing outcome
File-system Metadata
A file system maintains metadata that allows the operating system to locate and manage files.
File metadata can include:
- Filename
- Path
- File type
- Size
- Owner
- Permissions
- Creation and modification timestamps
- Links
- Block-location references
stat course-notes.pdf
The exact fields and timestamp semantics depend on the operating system and file system.
Object Metadata
An object-storage service maintains metadata describing each stored object. Metadata can be created by the storage system or supplied by the application.
{
"key": "tenants/17/courses/42/lessons/7/video.mp4",
"contentType": "video/mp4",
"contentLength": 52428800,
"checksumAlgorithm": "SHA-256",
"checksumValue": "expected-checksum-value",
"storageClass": "standard",
"versionId": "object-version-id",
"courseId": "42",
"lessonId": "7"
}
Object storage commonly distinguishes system-managed properties from custom metadata supplied by the application.
System-defined vs User-defined Metadata
| Metadata Type | Managed By | Examples |
|---|---|---|
| System-defined | Storage service | Content length, last-modified time, version, encryption status, and storage class |
| User-defined | Application or uploader | Course ID, lesson ID, source system, and processing correlation ID |
| Tags or labels | Application or operations | Environment, retention class, cost category, and workflow state |
Provider rule: Metadata limits, mutability, searchability, naming rules, and API behaviour differ between storage services. Verify the selected platform's contract before designing around custom metadata.
Application Metadata
Business metadata often belongs in an application database rather than only in the storage service.
The application database can represent:
- Tenant ownership
- Course, chapter, and lesson relationships
- Publication state
- Authorization rules
- Processing workflow
- Audit history
- Search visibility
- Retention requirements
- Object key and version references
Course-asset Metadata Table
CREATE TABLE course_assets
(
asset_id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
course_id BIGINT NOT NULL,
chapter_id BIGINT NULL,
lesson_id BIGINT NULL,
object_key VARCHAR(500) NOT NULL,
object_version VARCHAR(200) NULL,
original_file_name VARCHAR(255) NOT NULL,
display_name VARCHAR(255) NOT NULL,
content_type VARCHAR(100) NOT NULL,
content_length BIGINT NOT NULL,
checksum_algorithm VARCHAR(30) NOT NULL,
checksum_value VARCHAR(200) NOT NULL,
asset_type VARCHAR(30) NOT NULL,
asset_status VARCHAR(30) NOT NULL,
visibility VARCHAR(30) NOT NULL,
language_code VARCHAR(20) NULL,
created_by BIGINT NOT NULL,
created_at TIMESTAMP NOT NULL,
verified_at TIMESTAMP NULL,
published_at TIMESTAMP NULL,
CONSTRAINT uq_course_assets_object
UNIQUE
(
tenant_id,
object_key
),
CONSTRAINT fk_course_assets_course
FOREIGN KEY (course_id)
REFERENCES courses (course_id),
CONSTRAINT fk_course_assets_chapter
FOREIGN KEY (chapter_id)
REFERENCES chapters (chapter_id),
CONSTRAINT fk_course_assets_lesson
FOREIGN KEY (lesson_id)
REFERENCES lessons (lesson_id),
CONSTRAINT ck_course_asset_length
CHECK (content_length >= 0),
CONSTRAINT ck_course_asset_status
CHECK
(
asset_status IN
(
'pending',
'uploaded',
'verifying',
'processing',
'published',
'failed',
'quarantined',
'archived',
'deleted'
)
),
CONSTRAINT ck_course_asset_visibility
CHECK
(
visibility IN
(
'private',
'course',
'tenant',
'public'
)
)
);
Storage Metadata vs Database Metadata
| Area | Storage Metadata | Application Database Metadata |
|---|---|---|
| Primary focus | Properties of the stored object | Business meaning and relationships |
| Object identity | Bucket, key, and object version | Asset ID and domain relationships |
| Technical details | Size, content type, checksum, and storage class | Can retain validated copies of required technical fields |
| Business workflow | Limited or provider-specific | Pending, processing, published, archived, or failed |
| Relational queries | Provider-specific listing or metadata queries | SQL joins, constraints, indexes, and transactions |
| Authorization | Storage policies and object permissions | Tenant, ownership, course enrollment, and application roles |
Many systems use both. The storage service manages object properties, while the application database provides business context and queryable relationships.
Metadata Source of Truth
When the same field exists in several systems, define which source is authoritative.
Content length:
Authoritative source:
Verified object-store property
Course ID:
Authoritative source:
Application database
Video duration:
Authoritative source:
Approved media-analysis output
Search document:
Derived projection from
authoritative metadata
Ownership rule: Every duplicated metadata field should have one authoritative source, a synchronization mechanism, an acceptable staleness policy, and a reconciliation process.
Stable Identifiers
Metadata records should use stable identifiers that do not depend on mutable display values.
Object key:
courses/system-design-introduction/video.mp4
Problem:
Changing the course title can require
renaming or copying the stored object.
Object key:
tenants/17/courses/42/lessons/7/assets/981/video.mp4
Display title:
Stored separately as metadata
Original Filename vs Storage Key
The original filename and storage key serve different purposes.
| Value | Purpose |
|---|---|
| Original filename | User-facing reference retained after validation and normalization |
| Display name | Application-controlled title shown to users |
| Storage key | Stable generated identifier used to access stored content |
| Asset ID | Stable business identifier used by the application |
Do not use an untrusted original filename directly as an object key or local storage path.
Content Type
Content type describes the media type of a stored resource.
Examples:
application/pdf
image/png
image/jpeg
video/mp4
audio/mpeg
text/plain
text/csv
application/json
Content type can influence browser handling, media processing, security checks, indexing, caching, and download behaviour.
Validation rule: Do not trust only the client-provided filename extension or content type. Validate the content using the approved upload and security process.
Content Length
Content length records the number of stored bytes.
It can support:
- Upload-size limits
- Transfer progress
- Quota enforcement
- Integrity verification
- Storage-cost reporting
- Download headers
- Multipart-upload validation
Content length should be verified against the stored object rather than accepted only from the initial client request.
Checksum Metadata
{
"checksumAlgorithm": "SHA-256",
"checksumValue": "expected-checksum-value",
"checksumScope": "complete-object",
"verificationStatus": "verified",
"verifiedAt": "stored-verification-time"
}
Store the algorithm, value, calculation scope, object version, and verification status together. A checksum without context is ambiguous.
Metadata Lifecycle
Pending
|
v
Uploaded
|
v
Verifying
|
+-- Checksum failure --> Quarantined
|
v
Processing
|
+-- Processing failure --> Failed
|
v
Published
|
v
Archived
|
v
Deleted
Explicit lifecycle states make incomplete, failed, and published content distinguishable.
Lifecycle-state Example
| State | Meaning | Typical Access |
|---|---|---|
| Pending | The upload workflow has been created | Uploader and workflow services |
| Uploaded | The object exists but is not fully verified | Verification service |
| Processing | Metadata extraction, scanning, or transformation is running | Processing services |
| Published | The asset passed required checks and can be served | Authorized learners or public users |
| Quarantined | The asset requires restricted investigation | Approved administrators only |
| Archived | The asset remains retained but is no longer active | Authorized recovery or historical workflows |
Metadata during Upload
Client requests upload
|
v
Application validates business metadata
|
v
Create pending metadata record
|
v
Generate stable object key
|
v
Upload content
|
v
Storage creates technical metadata
|
v
Application verifies size and checksum
|
v
Extract additional metadata
|
v
Mark asset ready or quarantined
Metadata should not be marked published before the required upload, integrity, security, and processing checks complete.
Conceptual Object-upload Request
PUT /course-assets/tenants/17/courses/42/assets/981/video.mp4 HTTP/1.1
Host: storage.example.com
Content-Type: video/mp4
Content-Length: 52428800
X-Checksum-Algorithm: SHA-256
X-Checksum-Value: expected-checksum-value
X-Meta-Asset-Id: 981
X-Meta-Course-Id: 42
[BINARY CONTENT]
Header names, metadata limits, allowed characters, encoding, and update semantics are provider-specific.
PHP Metadata-validation Example
<?php
declare(strict_types=1);
function validateAssetMetadata(
array $metadata
): array {
$allowedTypes = [
'application/pdf',
'image/jpeg',
'image/png',
'video/mp4',
'text/plain'
];
$displayName =
trim(
(string)($metadata['displayName'] ?? '')
);
if ($displayName === '') {
throw new InvalidArgumentException(
'The display name is required.'
);
}
$contentType =
strtolower(
trim(
(string)($metadata['contentType'] ?? '')
)
);
if (!in_array(
$contentType,
$allowedTypes,
true
)) {
throw new InvalidArgumentException(
'The content type is not allowed.'
);
}
$contentLength =
filter_var(
$metadata['contentLength'] ?? null,
FILTER_VALIDATE_INT
);
if ($contentLength === false ||
$contentLength <= 0) {
throw new InvalidArgumentException(
'The content length is invalid.'
);
}
$visibility =
strtolower(
trim(
(string)($metadata['visibility'] ?? 'private')
)
);
$allowedVisibility = [
'private',
'course',
'tenant',
'public'
];
if (!in_array(
$visibility,
$allowedVisibility,
true
)) {
throw new InvalidArgumentException(
'The visibility value is invalid.'
);
}
return [
'displayName' => $displayName,
'contentType' => $contentType,
'contentLength' => $contentLength,
'visibility' => $visibility
];
}
Authorization, ownership, tenant scope, actual file inspection, checksum verification, quotas, and malware controls must be enforced separately.
Metadata Extraction
Some metadata can be extracted automatically from content after upload.
Stored asset
|
v
Metadata extractor
|
+-- Detect content type
+-- Read dimensions
+-- Read duration
+-- Count pages
+-- Detect language
+-- Extract document title
|
v
Validated extracted metadata
|
v
Update application record
Extracted metadata should be treated as untrusted processing output until it passes validation.
Metadata Versioning
Metadata can change independently from the underlying content.
Object version 1:
Same video bytes
Metadata revision 1:
Title = Introduction
Metadata revision 2:
Title = Introduction to System Design
Metadata revision 3:
Title = System Design Foundations
Decide whether metadata edits should:
- Update the current metadata record
- Create an audit entry
- Create a new metadata revision
- Create a new object version
- Trigger search reindexing
- Invalidate caches
Metadata History
CREATE TABLE asset_metadata_history
(
metadata_history_id BIGINT PRIMARY KEY,
asset_id BIGINT NOT NULL,
revision_number INT NOT NULL,
metadata_snapshot TEXT NOT NULL,
changed_by BIGINT NOT NULL,
changed_at TIMESTAMP NOT NULL,
change_reason VARCHAR(500) NULL,
CONSTRAINT fk_asset_metadata_history_asset
FOREIGN KEY (asset_id)
REFERENCES course_assets (asset_id),
CONSTRAINT uq_asset_metadata_revision
UNIQUE
(
asset_id,
revision_number
)
);
The snapshot format and retained fields should follow audit, privacy, retention, and operational requirements.
Metadata and Authorization
Metadata can participate in authorization but should not be trusted without validation and protected update rules.
Download request
|
v
Load authoritative asset metadata
|
v
Verify tenant scope
|
v
Verify asset status
|
v
Verify visibility
|
v
Verify caller relationship or permission
|
v
Grant narrowly scoped access
A user should not be able to change metadata such as tenant ownership, publication state, security classification, or legal hold merely by submitting those fields in a client request.
Sensitive Metadata
Metadata can reveal sensitive information even when the underlying file remains encrypted or inaccessible.
Potentially sensitive metadata includes:
- Original filenames
- User or tenant identifiers
- Document titles
- Storage paths and object keys
- Location data
- Device information
- Processing history
- Security classification
- Access history
Privacy rule: Protect metadata according to the sensitivity of the information it reveals, not merely according to the size of the metadata record.
Metadata Sanitization
Uploaded files can contain embedded metadata that should not always be preserved or published.
Examples include:
- Image-camera information
- Location coordinates
- Document author names
- Editing-software details
- Internal file paths
- Revision history
- Hidden document properties
Define which embedded metadata is extracted, retained, removed, displayed, indexed, or shared.
Metadata and Search
Metadata provides structured fields for search filtering and discovery.
Search request:
System Design videos in English
Possible metadata filters:
asset_type = video
language_code = en
course_status = published
visibility allows current learner
title or keywords match System Design
Metadata can support:
- Filters
- Facets
- Sorting
- Grouping
- Ranking signals
- Security trimming
- Lifecycle queries
- Operational dashboards
Metadata-query Example
SELECT
asset_id,
display_name,
content_type,
content_length,
language_code,
published_at
FROM course_assets
WHERE tenant_id = :tenant_id
AND course_id = :course_id
AND asset_status = 'published'
AND visibility IN
(
'course',
'tenant',
'public'
)
ORDER BY
published_at DESC,
asset_id DESC;
Candidate Index
CREATE INDEX ix_assets_tenant_course_status_date
ON course_assets
(
tenant_id,
course_id,
asset_status,
published_at DESC,
asset_id DESC
);
Index design should be validated against representative queries and execution plans.
Tags
Tags are commonly represented as key-value labels used for categorization, workflow, cost allocation, lifecycle, or policy evaluation.
{
"environment": "production",
"content-category": "course-video",
"retention-class": "published-learning-content",
"processing-status": "verified"
}
Tags should use a controlled vocabulary. Unrestricted tag names and values can create spelling variants, inconsistent automation, and unmanageable governance.
Controlled Vocabularies
Language:
English
english
EN
en-US
Eng
Status:
Published
publish
live
active
visible
language_code:
en
bn
hi
asset_status:
pending
processing
published
archived
Define whether codes represent language generally, regional variants, or a separate business classification.
Metadata Validation
Validate metadata for:
- Required fields
- Data types
- Length limits
- Allowed values
- Character encoding
- Tenant scope
- Identifier existence
- Date and time format
- Cross-field consistency
- Authorization to set protected fields
Cross-field Example
CONSTRAINT ck_asset_publication
CHECK
(
published_at IS NULL
OR asset_status = 'published'
);
More complex lifecycle rules can require transactional application logic or database-specific mechanisms.
Metadata Synchronization
Metadata can exist in storage, a relational database, a search index, a cache, and analytics systems.
Authoritative metadata database
|
+-- Object-storage metadata
|
+-- Search index
|
+-- Cache
|
+-- Analytics projection
Each derived representation should define:
- Source of truth
- Update trigger
- Acceptable delay
- Retry strategy
- Idempotency
- Reconciliation process
- Rebuild procedure
Event-driven Metadata Projection
Metadata transaction commits
|
v
Outbox event becomes available
|
v
Publisher sends event
|
v
Search-index consumer updates document
|
v
Cache or analytics projection updates
Consumers should handle duplicate events safely and support replay or rebuild from authoritative metadata.
Metadata Drift
Metadata drift occurs when two representations no longer agree.
Database metadata:
asset_status = published
Object missing:
No object exists for stored key
Search metadata:
asset_status = processing
Result:
Three representations disagree.
Drift can result from partial failures, delayed events, manual changes, failed processing, incorrect retries, or deleted objects.
Metadata Reconciliation
A reconciliation process compares authoritative metadata with actual stored objects and derived representations.
Read authoritative asset records
|
v
Check referenced storage objects
|
v
Compare size, checksum, and version
|
v
Check search or cache projection
|
+-- Match:
| mark healthy
|
+-- Mismatch:
repair, reindex,
quarantine, or investigate
Find Metadata without Objects
For each active metadata record:
1. Resolve its object key.
2. Request object properties.
3. Verify existence.
4. Compare object version, length, and checksum.
5. Record or repair mismatches.
Find Objects without Metadata
For each stored object in managed prefix:
1. Extract stable asset identifier.
2. Find authoritative metadata record.
3. Confirm tenant and lifecycle state.
4. Quarantine or clean abandoned content
according to policy.
Metadata Deletion
Deleting metadata before stored content can create an object that is no longer associated with an active application record. Deleting content first can leave a record pointing to missing bytes.
Safer deletion workflow:
1. Authorize deletion.
2. Mark metadata as pending deletion.
3. Apply retention and legal-hold checks.
4. Delete or queue content deletion.
5. Verify storage outcome.
6. Remove derived search and cache entries.
7. Finalize metadata state.
8. Retain required audit evidence.
Retention Metadata
{
"retentionClass": "course-content",
"retainUntil": "approved-retention-date",
"legalHold": false,
"deletionStatus": "not-requested"
}
Retention values can have legal, privacy, audit, and operational impact. Only authorized workflows should modify them.
Time Metadata
Time fields should have clearly defined semantics.
| Field | Meaning |
|---|---|
| created_at | Business metadata record creation time |
| uploaded_at | Object-upload completion time |
| verified_at | Integrity-verification completion time |
| published_at | Time content became published |
| last_modified_at | Time the tracked metadata or object changed according to the defined source |
| archived_at | Time content entered archived state |
Use a consistent time standard and retain the original timezone only when it has business meaning.
Schema Evolution
Metadata schemas change as applications add new fields, classifications, workflows, and processing outputs.
Metadata schema version 1:
title
content_type
object_key
Metadata schema version 2:
title
content_type
object_key
language_code
checksum
visibility
Metadata schema version 3:
Adds processing and retention fields
Plan for:
- Optional fields during migration
- Default values
- Backfilling older records
- Version-aware consumers
- Search-index updates
- Compatibility during rolling deployment
- Removal of obsolete fields
Flexible vs Structured Metadata
| Approach | Benefit | Consideration |
|---|---|---|
| Structured columns | Strong typing, validation, indexes, and queryability | Schema changes require migrations |
| Key-value metadata | Supports varying optional attributes | Can weaken consistency and increase query complexity |
| JSON metadata | Provides a flexible nested document | Frequently queried fields can require validation and specialized indexing |
| Separate related tables | Supports repeatable and relational metadata values | Requires joins and additional schema objects |
Use structured columns for critical and frequently queried business fields. Use flexible metadata deliberately for genuinely variable attributes.
Multi-valued Metadata
Metadata such as keywords, contributors, captions, and categories can contain several values.
keywords:
system design,database,storage,cloud
CREATE TABLE asset_keywords
(
asset_id BIGINT NOT NULL,
keyword VARCHAR(100) NOT NULL,
PRIMARY KEY
(
asset_id,
keyword
),
CONSTRAINT fk_asset_keywords_asset
FOREIGN KEY (asset_id)
REFERENCES course_assets (asset_id)
ON DELETE CASCADE
);
Derived Metadata
Derived metadata is calculated or extracted from existing content or authoritative metadata.
Examples include:
- Video duration
- Image dimensions
- Document page count
- Detected language
- Generated thumbnail key
- Searchable extracted text
- Computed storage cost category
- Latest processing result
Derived metadata should record how and when it was generated when reproducibility or troubleshooting matters.
Generated Metadata
Automated systems can produce classifications, summaries, labels, or extracted entities.
{
"generator": "metadata-processing-service",
"generatorVersion": "approved-version",
"generatedAt": "processing-time",
"classification": "technical-course-content",
"confidence": 0.94,
"reviewStatus": "pending"
}
Generated metadata should remain distinguishable from human-approved business metadata. Confidence, model or processor version, and review state can be important for governance.
Metadata Performance
A storage platform can contain many small metadata records even when the stored files are large. Metadata operations can become a distinct performance bottleneck.
Important metadata operations include:
- Lookup by object key
- Directory listing
- Object listing by prefix
- Permission checks
- Version lookup
- Tag retrieval
- Lifecycle scans
- Search-index projection
- Reconciliation
Design metadata storage for its own query rate, update rate, consistency, index size, and failure behaviour.
Metadata Service
Clients
|
v
Metadata API
|
+-- Metadata database
|
+-- Authorization service
|
+-- Search projection
|
+-- Object storage
|
+-- Processing events
A dedicated metadata service can centralize validation, ownership, lifecycle, object references, security decisions, and change events.
Whether a separate service is justified depends on scale, team boundaries, latency, consistency, and operational complexity.
Metadata Observability
Useful metadata metrics include:
- Metadata record count
- Records by lifecycle state
- Metadata query latency
- Metadata update failure count
- Objects missing business metadata
- Metadata records referencing missing objects
- Checksum and length mismatches
- Search-index synchronization delay
- Metadata extraction failures
- Quarantined asset count
- Schema-version distribution
- Reconciliation backlog
- Unauthorized metadata-update attempts
- Retention-policy exceptions
Metadata Alert Conditions
Alert when:
- Published metadata points to missing content
- Metadata extraction repeatedly fails
- Search-projection delay exceeds its objective
- Checksum or content-length drift is detected
- Quarantined assets increase unexpectedly
- Required metadata fields are missing
- Retention or legal-hold fields change unexpectedly
- Reconciliation cannot repair a mismatch
- Metadata queries exceed latency thresholds
Metadata Troubleshooting Workflow
- Identify the affected asset and stable identifier.
- Determine the authoritative metadata source.
- Retrieve the application metadata record.
- Retrieve the current storage-object properties.
- Compare object key, version, length, and checksum.
- Check lifecycle and publication state.
- Check tenant ownership and authorization fields.
- Check extraction and processing history.
- Check search and cache projections.
- Review recent metadata-change events.
- Repair derived copies from the authoritative source.
- Quarantine unresolved content or metadata conflicts.
Common Metadata Mistakes
Storing Content without Business Metadata
The application cannot reliably identify ownership, purpose, lifecycle, visibility, or relationships.
Using the Original Filename as the Storage Key
Filenames can be untrusted, duplicated, mutable, excessively long, or incompatible with storage-key rules.
Trusting Client-supplied Metadata
Clients can submit false content types, sizes, ownership, visibility, or workflow states.
Keeping No Source of Truth
Storage, database, cache, and search metadata can diverge when no authoritative source is defined.
Mixing Technical and Business Meaning
A storage object's last-modified time is not automatically the business publication time or course-update time.
Using Uncontrolled Tags
Inconsistent spelling, capitalization, and terminology make querying and automation unreliable.
Placing Every Attribute in Flexible JSON
Critical fields can lose strong typing, constraints, indexes, and clear ownership.
Duplicating Metadata without Synchronization
Search, cache, analytics, storage, and relational copies can show contradictory values.
Publishing before Verification Completes
Incomplete, corrupted, unsupported, or unsafe content can become visible to users.
Ignoring Embedded Sensitive Metadata
Uploaded images and documents can disclose location, author, software, or internal-path information.
Deleting Metadata and Content Independently
Partial failures can create orphaned objects or broken application references.
Assuming Metadata Is Too Small to Need Governance
Metadata can expose sensitive business context and control consequential access, retention, and processing decisions.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Required metadata | Missing required fields are rejected |
| Invalid controlled value | Unsupported status, visibility, or asset type is rejected |
| Untrusted filename | The application generates a safe, stable storage key |
| False client content type | Server-side validation identifies the actual approved type |
| Content-length mismatch | The asset remains unpublished and enters failed or quarantined state |
| Checksum mismatch | The metadata records the integrity failure |
| Missing object | Reconciliation detects the broken metadata reference |
| Orphaned object | Reconciliation identifies content without an active metadata record |
| Unauthorized metadata update | Protected ownership, visibility, and retention fields are not changed |
| Search synchronization | Published metadata appears in the search projection within its objective |
| Metadata revision | The change is audited or versioned according to policy |
| Deletion workflow | Metadata, content, cache, and search state converge correctly |
Metadata Best Practices
Recommended Practices
- Separate stored content from the metadata that describes it.
- Use stable asset identifiers and generated storage keys.
- Retain original filenames only as validated display metadata.
- Distinguish system, technical, business, operational, and security metadata.
- Define an authoritative source for every important field.
- Use structured columns for critical and frequently queried metadata.
- Use flexible metadata only for genuinely variable attributes.
- Validate all client-provided metadata.
- Verify content type, length, checksum, and object version after upload.
- Use controlled vocabularies for status, language, type, and classification.
- Model multi-valued metadata through appropriate related structures.
- Record provenance for important automated transformations.
- Distinguish generated metadata from approved business metadata.
- Protect metadata according to its sensitivity.
- Remove embedded sensitive metadata when required.
- Publish content only after required checks complete.
- Synchronize search, cache, and analytics projections reliably.
- Implement reconciliation between metadata and stored content.
- Audit consequential metadata changes.
- Monitor incomplete, inconsistent, unauthorized, and orphaned metadata.
Practice Exercise
Design metadata management for course assets on your online learning platform.
Requirements
- Create metadata for videos, PDFs, images, presentations, and source-code files.
- Generate a stable asset ID and object key.
- Store the validated original filename separately.
- Associate each asset with a tenant, course, chapter, or lesson.
- Store content type, content length, checksum, and object version.
- Record language and accessibility metadata.
- Define pending, processing, published, failed, quarantined, and archived states.
- Record uploader and processing provenance.
- Extract technical metadata after upload.
- Prevent clients from assigning protected metadata fields.
- Project published metadata into the search index.
- Track metadata revisions.
- Detect database records referencing missing objects.
- Detect stored objects without metadata records.
- Test the complete deletion and retention workflow.
Metadata-design Template
| Metadata Field | Source of Truth | Validation | Purpose |
|---|---|---|---|
| Asset ID | Application database | Primary key | Stable business identity |
| Object key | Application-generated metadata | Tenant-scoped uniqueness | Storage identity |
| Content type | Approved upload validation | Allowed media-type list | Delivery and processing |
| Content length | Verified stored object | Non-negative and within quota | Integrity, quota, and transfer handling |
| Checksum | Verified checksum calculation | Algorithm and value validation | Content integrity |
| Publication status | Application workflow | Controlled lifecycle transition | Content visibility |
| Search document | Derived from authoritative metadata | Projection and reconciliation | Discovery and filtering |
Frequently Asked Questions
What is metadata?
Metadata is information that describes, identifies, organizes, governs, or connects another data item or resource.
What is descriptive metadata?
Descriptive metadata explains what content represents, using fields such as title, description, language, subject, and keywords.
What is structural metadata?
Structural metadata describes how content components relate, such as course, chapter, lesson, page, and display-order relationships.
What is technical metadata?
Technical metadata records properties such as format, content type, size, checksum, dimensions, duration, codec, and object version.
What is provenance metadata?
Provenance metadata records the origin, processing history, and transformations applied to content.
What is object metadata?
Object metadata describes a stored object through system-managed or application-defined properties such as key, size, content type, checksum, storage class, and custom values.
Should all metadata be stored in object storage?
No. Business relationships, workflow, authorization, audit, and searchable structured fields commonly belong in an application database or dedicated metadata system.
Why should object keys be stable?
Stable keys avoid storage changes when mutable display values such as filenames, titles, or course names change.
Can metadata be sensitive?
Yes. Filenames, ownership, locations, titles, object keys, classifications, and processing history can reveal sensitive information.
How does metadata support search?
Metadata provides structured fields for filtering, faceting, sorting, grouping, authorization, and ranking.
What is metadata drift?
Metadata drift occurs when storage, database, search, cache, or analytics representations no longer agree.
What comes after metadata?
The next topic is uploads and multipart transfer, followed by inverted indexes and full-text search.
Key Takeaway
Metadata gives stored content its identity, meaning, relationships, technical context, lifecycle, security classification, and searchability. Distinguish system-defined, user-defined, descriptive, structural, technical, administrative, operational, and provenance metadata. Use stable identifiers and generated object keys, validate all untrusted fields, define a source of truth for duplicated values, protect sensitive metadata, and publish content only after required verification completes. Keep business metadata in a constrained and queryable system where appropriate, synchronize derived search and cache representations reliably, and reconcile metadata records with actual stored objects.