Table of Contents

    metadata

    STORAGE, FILES, OBJECTS & SEARCH BASICS

    Metadata

    Learn how metadata describes files, objects, blocks, and business assets; supports discovery, integrity, authorization, lifecycle management, auditing, and search; and must be modeled, validated, indexed, protected, synchronized, and governed separately from the underlying content.

    Introduction

    Applications rarely manage content as anonymous bytes. A stored video needs a title, content type, size, owner, checksum, processing state, access classification, and retention policy. A document can require a course, chapter, lesson, language, version, and publication status.

    This descriptive and operational information is called metadata.

    Metadata helps a system understand what stored content represents, where it belongs, how it should be processed, who may access it, how long it should remain, and whether it passed integrity verification.

    Core idea: Data is the stored content. Metadata describes, identifies, classifies, governs, and connects that content to the application domain.

    In your System Design curriculum, Metadata is Topic 6.4 under Storage, Files, Objects & Search Basics. It follows checksums and precedes uploads and multipart transfer, inverted indexes, and full-text search.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Block, file, and object storage Each storage model maintains different forms of descriptive and operational metadata.
    2 Entities and relationships Application metadata commonly connects stored content to business entities.
    3 Keys and constraints Metadata records require stable identities, uniqueness, validation, and referential integrity.
    4 Checksums Checksum algorithm, value, and verification status are important integrity metadata.
    5 Security fundamentals Metadata can contain sensitive information and influence access-control decisions.
    6 Search fundamentals Metadata provides structured fields for filtering, faceting, ranking, and discovery.

    What Is Metadata?

    Metadata is data that describes another data item, resource, object, file, event, or dataset.

    Stored content:
    
    lesson-07.mp4
    
    
    Metadata:
    
    - Asset ID
    - Object key
    - Original filename
    - Content type
    - Content length
    - Course ID
    - Lesson ID
    - Language
    - Checksum
    - Upload time
    - Processing status
    - Visibility
    - Retention class

    The video bytes contain the media content. The metadata describes how the application should identify, organize, validate, process, protect, and present that content.

    Metadata Flow
    content created → metadata captured → metadata validated → content processed → metadata updated → content discovered and governed

    Data vs Metadata

    Area Data Metadata
    Video asset Encoded audio and video bytes Title, duration, content type, checksum, and lesson ID
    PDF document Document bytes and embedded content Filename, author, page count, language, and publication state
    Database table Rows stored in the table Column names, data types, constraints, indexes, and ownership
    Image Pixel and encoding data Dimensions, format, orientation, checksum, and alt text
    Backup Copied application or database content Backup time, source, checksum, retention, and recovery status

    Categories of Metadata

    Metadata can be grouped according to its purpose.

    Category Purpose Examples
    Descriptive Explains what the content is Title, description, keywords, language, and subject
    Structural Explains how components relate Course, chapter, lesson, page order, and media segments
    Administrative Supports management and governance Owner, visibility, retention class, and storage tier
    Technical Describes technical properties Content type, size, dimensions, codec, and checksum
    Operational Tracks processing and lifecycle state Pending, verified, scanning, published, archived, and failed
    Provenance Records origin and transformation history Uploader, source system, creation time, and processing version
    Security Supports protection and access decisions Classification, tenant ID, access scope, and legal hold

    Descriptive Metadata

    Descriptive metadata helps users and applications understand and discover content.

    {
      "title": "Introduction to System Design",
      "description": "Foundational concepts for designing scalable systems.",
      "language": "en",
      "contentType": "video/mp4",
      "keywords": [
        "system design",
        "scalability",
        "architecture"
      ]
    }

    Descriptive fields can later feed catalog search, filters, recommendations, accessibility features, and user-interface labels.

    Structural Metadata

    Structural metadata describes relationships among content components.

    Course
      |
      +-- Chapter 1
      |     |
      |     +-- Lesson 1
      |     +-- Lesson 2
      |
      +-- Chapter 2
            |
            +-- Lesson 3
            +-- Lesson 4

    Structural metadata can define:

    • Parent-child relationships
    • Display order
    • Document page sequence
    • Video segments
    • Related attachments
    • Alternative renditions
    • Course-module composition

    Technical Metadata

    Technical metadata describes properties required to store, validate, deliver, or process content correctly.

    Content Technical Metadata
    Image Format, dimensions, color profile, orientation, and size
    Video Codec, duration, resolution, frame rate, and bitrate
    Audio Codec, sample rate, channel count, duration, and bitrate
    Document Format, page count, character encoding, and content length
    Object Key, version, content type, length, checksum, and storage class

    Administrative Metadata

    Administrative metadata supports ownership, governance, lifecycle, and operational management.

    {
      "tenantId": "tenant-17",
      "ownerId": "course-team-42",
      "visibility": "private",
      "retentionClass": "course-content",
      "storageTier": "standard",
      "publicationStatus": "published",
      "legalHold": false
    }

    Administrative metadata can influence important system behaviour, so its accepted values and update permissions should be controlled carefully.

    Provenance Metadata

    Provenance metadata records where content came from and how it changed.

    Original upload
          |
          v
    Malware scan
          |
          v
    Metadata extraction
          |
          v
    Video transcoding
          |
          v
    Thumbnail generation
          |
          v
    Publication

    Useful provenance fields can include:

    • Source system
    • Ingestion time
    • Uploader or service identity
    • Original asset ID
    • Transformation name
    • Transformation version
    • Parent object version
    • Processing time
    • Processing outcome

    File-system Metadata

    A file system maintains metadata that allows the operating system to locate and manage files.

    File metadata can include:

    • Filename
    • Path
    • File type
    • Size
    • Owner
    • Permissions
    • Creation and modification timestamps
    • Links
    • Block-location references
    stat course-notes.pdf

    The exact fields and timestamp semantics depend on the operating system and file system.

    Object Metadata

    An object-storage service maintains metadata describing each stored object. Metadata can be created by the storage system or supplied by the application.

    {
      "key": "tenants/17/courses/42/lessons/7/video.mp4",
      "contentType": "video/mp4",
      "contentLength": 52428800,
      "checksumAlgorithm": "SHA-256",
      "checksumValue": "expected-checksum-value",
      "storageClass": "standard",
      "versionId": "object-version-id",
      "courseId": "42",
      "lessonId": "7"
    }

    Object storage commonly distinguishes system-managed properties from custom metadata supplied by the application.

    System-defined vs User-defined Metadata

    Metadata Type Managed By Examples
    System-defined Storage service Content length, last-modified time, version, encryption status, and storage class
    User-defined Application or uploader Course ID, lesson ID, source system, and processing correlation ID
    Tags or labels Application or operations Environment, retention class, cost category, and workflow state

    Provider rule: Metadata limits, mutability, searchability, naming rules, and API behaviour differ between storage services. Verify the selected platform's contract before designing around custom metadata.

    Application Metadata

    Business metadata often belongs in an application database rather than only in the storage service.

    The application database can represent:

    • Tenant ownership
    • Course, chapter, and lesson relationships
    • Publication state
    • Authorization rules
    • Processing workflow
    • Audit history
    • Search visibility
    • Retention requirements
    • Object key and version references

    Course-asset Metadata Table

    CREATE TABLE course_assets
    (
        asset_id BIGINT PRIMARY KEY,
        tenant_id BIGINT NOT NULL,
        course_id BIGINT NOT NULL,
        chapter_id BIGINT NULL,
        lesson_id BIGINT NULL,
    
        object_key VARCHAR(500) NOT NULL,
        object_version VARCHAR(200) NULL,
    
        original_file_name VARCHAR(255) NOT NULL,
        display_name VARCHAR(255) NOT NULL,
        content_type VARCHAR(100) NOT NULL,
        content_length BIGINT NOT NULL,
    
        checksum_algorithm VARCHAR(30) NOT NULL,
        checksum_value VARCHAR(200) NOT NULL,
    
        asset_type VARCHAR(30) NOT NULL,
        asset_status VARCHAR(30) NOT NULL,
        visibility VARCHAR(30) NOT NULL,
        language_code VARCHAR(20) NULL,
    
        created_by BIGINT NOT NULL,
        created_at TIMESTAMP NOT NULL,
        verified_at TIMESTAMP NULL,
        published_at TIMESTAMP NULL,
    
        CONSTRAINT uq_course_assets_object
            UNIQUE
            (
                tenant_id,
                object_key
            ),
    
        CONSTRAINT fk_course_assets_course
            FOREIGN KEY (course_id)
            REFERENCES courses (course_id),
    
        CONSTRAINT fk_course_assets_chapter
            FOREIGN KEY (chapter_id)
            REFERENCES chapters (chapter_id),
    
        CONSTRAINT fk_course_assets_lesson
            FOREIGN KEY (lesson_id)
            REFERENCES lessons (lesson_id),
    
        CONSTRAINT ck_course_asset_length
            CHECK (content_length >= 0),
    
        CONSTRAINT ck_course_asset_status
            CHECK
            (
                asset_status IN
                (
                    'pending',
                    'uploaded',
                    'verifying',
                    'processing',
                    'published',
                    'failed',
                    'quarantined',
                    'archived',
                    'deleted'
                )
            ),
    
        CONSTRAINT ck_course_asset_visibility
            CHECK
            (
                visibility IN
                (
                    'private',
                    'course',
                    'tenant',
                    'public'
                )
            )
    );

    Storage Metadata vs Database Metadata

    Area Storage Metadata Application Database Metadata
    Primary focus Properties of the stored object Business meaning and relationships
    Object identity Bucket, key, and object version Asset ID and domain relationships
    Technical details Size, content type, checksum, and storage class Can retain validated copies of required technical fields
    Business workflow Limited or provider-specific Pending, processing, published, archived, or failed
    Relational queries Provider-specific listing or metadata queries SQL joins, constraints, indexes, and transactions
    Authorization Storage policies and object permissions Tenant, ownership, course enrollment, and application roles

    Many systems use both. The storage service manages object properties, while the application database provides business context and queryable relationships.

    Metadata Source of Truth

    When the same field exists in several systems, define which source is authoritative.

    Content length:
    
    Authoritative source:
    Verified object-store property
    
    
    Course ID:
    
    Authoritative source:
    Application database
    
    
    Video duration:
    
    Authoritative source:
    Approved media-analysis output
    
    
    Search document:
    
    Derived projection from
    authoritative metadata

    Ownership rule: Every duplicated metadata field should have one authoritative source, a synchronization mechanism, an acceptable staleness policy, and a reconciliation process.

    Stable Identifiers

    Metadata records should use stable identifiers that do not depend on mutable display values.

    Mutable identity
    Object key:
    
    courses/system-design-introduction/video.mp4
    
    
    Problem:
    
    Changing the course title can require
    renaming or copying the stored object.
    Stable identity
    Object key:
    
    tenants/17/courses/42/lessons/7/assets/981/video.mp4
    
    
    Display title:
    
    Stored separately as metadata

    Original Filename vs Storage Key

    The original filename and storage key serve different purposes.

    Value Purpose
    Original filename User-facing reference retained after validation and normalization
    Display name Application-controlled title shown to users
    Storage key Stable generated identifier used to access stored content
    Asset ID Stable business identifier used by the application

    Do not use an untrusted original filename directly as an object key or local storage path.

    Content Type

    Content type describes the media type of a stored resource.

    Examples:
    
    application/pdf
    image/png
    image/jpeg
    video/mp4
    audio/mpeg
    text/plain
    text/csv
    application/json

    Content type can influence browser handling, media processing, security checks, indexing, caching, and download behaviour.

    Validation rule: Do not trust only the client-provided filename extension or content type. Validate the content using the approved upload and security process.

    Content Length

    Content length records the number of stored bytes.

    It can support:

    • Upload-size limits
    • Transfer progress
    • Quota enforcement
    • Integrity verification
    • Storage-cost reporting
    • Download headers
    • Multipart-upload validation

    Content length should be verified against the stored object rather than accepted only from the initial client request.

    Checksum Metadata

    {
      "checksumAlgorithm": "SHA-256",
      "checksumValue": "expected-checksum-value",
      "checksumScope": "complete-object",
      "verificationStatus": "verified",
      "verifiedAt": "stored-verification-time"
    }

    Store the algorithm, value, calculation scope, object version, and verification status together. A checksum without context is ambiguous.

    Metadata Lifecycle

    Pending
       |
       v
    Uploaded
       |
       v
    Verifying
       |
       +-- Checksum failure --> Quarantined
       |
       v
    Processing
       |
       +-- Processing failure --> Failed
       |
       v
    Published
       |
       v
    Archived
       |
       v
    Deleted

    Explicit lifecycle states make incomplete, failed, and published content distinguishable.

    Lifecycle-state Example

    State Meaning Typical Access
    Pending The upload workflow has been created Uploader and workflow services
    Uploaded The object exists but is not fully verified Verification service
    Processing Metadata extraction, scanning, or transformation is running Processing services
    Published The asset passed required checks and can be served Authorized learners or public users
    Quarantined The asset requires restricted investigation Approved administrators only
    Archived The asset remains retained but is no longer active Authorized recovery or historical workflows

    Metadata during Upload

    Client requests upload
          |
          v
    Application validates business metadata
          |
          v
    Create pending metadata record
          |
          v
    Generate stable object key
          |
          v
    Upload content
          |
          v
    Storage creates technical metadata
          |
          v
    Application verifies size and checksum
          |
          v
    Extract additional metadata
          |
          v
    Mark asset ready or quarantined

    Metadata should not be marked published before the required upload, integrity, security, and processing checks complete.

    Conceptual Object-upload Request

    PUT /course-assets/tenants/17/courses/42/assets/981/video.mp4 HTTP/1.1
    Host: storage.example.com
    Content-Type: video/mp4
    Content-Length: 52428800
    X-Checksum-Algorithm: SHA-256
    X-Checksum-Value: expected-checksum-value
    X-Meta-Asset-Id: 981
    X-Meta-Course-Id: 42
    
    [BINARY CONTENT]

    Header names, metadata limits, allowed characters, encoding, and update semantics are provider-specific.

    PHP Metadata-validation Example

    <?php
    
    declare(strict_types=1);
    
    function validateAssetMetadata(
        array $metadata
    ): array {
        $allowedTypes = [
            'application/pdf',
            'image/jpeg',
            'image/png',
            'video/mp4',
            'text/plain'
        ];
    
        $displayName =
            trim(
                (string)($metadata['displayName'] ?? '')
            );
    
        if ($displayName === '') {
            throw new InvalidArgumentException(
                'The display name is required.'
            );
        }
    
        $contentType =
            strtolower(
                trim(
                    (string)($metadata['contentType'] ?? '')
                )
            );
    
        if (!in_array(
            $contentType,
            $allowedTypes,
            true
        )) {
            throw new InvalidArgumentException(
                'The content type is not allowed.'
            );
        }
    
        $contentLength =
            filter_var(
                $metadata['contentLength'] ?? null,
                FILTER_VALIDATE_INT
            );
    
        if ($contentLength === false ||
            $contentLength <= 0) {
    
            throw new InvalidArgumentException(
                'The content length is invalid.'
            );
        }
    
        $visibility =
            strtolower(
                trim(
                    (string)($metadata['visibility'] ?? 'private')
                )
            );
    
        $allowedVisibility = [
            'private',
            'course',
            'tenant',
            'public'
        ];
    
        if (!in_array(
            $visibility,
            $allowedVisibility,
            true
        )) {
            throw new InvalidArgumentException(
                'The visibility value is invalid.'
            );
        }
    
        return [
            'displayName' => $displayName,
            'contentType' => $contentType,
            'contentLength' => $contentLength,
            'visibility' => $visibility
        ];
    }

    Authorization, ownership, tenant scope, actual file inspection, checksum verification, quotas, and malware controls must be enforced separately.

    Metadata Extraction

    Some metadata can be extracted automatically from content after upload.

    Stored asset
          |
          v
    Metadata extractor
          |
          +-- Detect content type
          +-- Read dimensions
          +-- Read duration
          +-- Count pages
          +-- Detect language
          +-- Extract document title
          |
          v
    Validated extracted metadata
          |
          v
    Update application record

    Extracted metadata should be treated as untrusted processing output until it passes validation.

    Metadata Versioning

    Metadata can change independently from the underlying content.

    Object version 1:
    
    Same video bytes
    
    
    Metadata revision 1:
    
    Title = Introduction
    
    
    Metadata revision 2:
    
    Title = Introduction to System Design
    
    
    Metadata revision 3:
    
    Title = System Design Foundations

    Decide whether metadata edits should:

    • Update the current metadata record
    • Create an audit entry
    • Create a new metadata revision
    • Create a new object version
    • Trigger search reindexing
    • Invalidate caches

    Metadata History

    CREATE TABLE asset_metadata_history
    (
        metadata_history_id BIGINT PRIMARY KEY,
        asset_id BIGINT NOT NULL,
        revision_number INT NOT NULL,
        metadata_snapshot TEXT NOT NULL,
        changed_by BIGINT NOT NULL,
        changed_at TIMESTAMP NOT NULL,
        change_reason VARCHAR(500) NULL,
    
        CONSTRAINT fk_asset_metadata_history_asset
            FOREIGN KEY (asset_id)
            REFERENCES course_assets (asset_id),
    
        CONSTRAINT uq_asset_metadata_revision
            UNIQUE
            (
                asset_id,
                revision_number
            )
    );

    The snapshot format and retained fields should follow audit, privacy, retention, and operational requirements.

    Metadata and Authorization

    Metadata can participate in authorization but should not be trusted without validation and protected update rules.

    Download request
          |
          v
    Load authoritative asset metadata
          |
          v
    Verify tenant scope
          |
          v
    Verify asset status
          |
          v
    Verify visibility
          |
          v
    Verify caller relationship or permission
          |
          v
    Grant narrowly scoped access

    A user should not be able to change metadata such as tenant ownership, publication state, security classification, or legal hold merely by submitting those fields in a client request.

    Sensitive Metadata

    Metadata can reveal sensitive information even when the underlying file remains encrypted or inaccessible.

    Potentially sensitive metadata includes:

    • Original filenames
    • User or tenant identifiers
    • Document titles
    • Storage paths and object keys
    • Location data
    • Device information
    • Processing history
    • Security classification
    • Access history

    Privacy rule: Protect metadata according to the sensitivity of the information it reveals, not merely according to the size of the metadata record.

    Metadata Sanitization

    Uploaded files can contain embedded metadata that should not always be preserved or published.

    Examples include:

    • Image-camera information
    • Location coordinates
    • Document author names
    • Editing-software details
    • Internal file paths
    • Revision history
    • Hidden document properties

    Define which embedded metadata is extracted, retained, removed, displayed, indexed, or shared.

    Metadata and Search

    Metadata provides structured fields for search filtering and discovery.

    Search request:
    
    System Design videos in English
    
    
    Possible metadata filters:
    
    asset_type = video
    language_code = en
    course_status = published
    visibility allows current learner
    title or keywords match System Design

    Metadata can support:

    • Filters
    • Facets
    • Sorting
    • Grouping
    • Ranking signals
    • Security trimming
    • Lifecycle queries
    • Operational dashboards

    Metadata-query Example

    SELECT
        asset_id,
        display_name,
        content_type,
        content_length,
        language_code,
        published_at
    FROM course_assets
    WHERE tenant_id = :tenant_id
      AND course_id = :course_id
      AND asset_status = 'published'
      AND visibility IN
      (
          'course',
          'tenant',
          'public'
      )
    ORDER BY
        published_at DESC,
        asset_id DESC;

    Candidate Index

    CREATE INDEX ix_assets_tenant_course_status_date
    ON course_assets
    (
        tenant_id,
        course_id,
        asset_status,
        published_at DESC,
        asset_id DESC
    );

    Index design should be validated against representative queries and execution plans.

    Tags

    Tags are commonly represented as key-value labels used for categorization, workflow, cost allocation, lifecycle, or policy evaluation.

    {
      "environment": "production",
      "content-category": "course-video",
      "retention-class": "published-learning-content",
      "processing-status": "verified"
    }

    Tags should use a controlled vocabulary. Unrestricted tag names and values can create spelling variants, inconsistent automation, and unmanageable governance.

    Controlled Vocabularies

    Uncontrolled values
    Language:
    
    English
    english
    EN
    en-US
    Eng
    
    
    Status:
    
    Published
    publish
    live
    active
    visible
    Controlled values
    language_code:
    
    en
    bn
    hi
    
    
    asset_status:
    
    pending
    processing
    published
    archived

    Define whether codes represent language generally, regional variants, or a separate business classification.

    Metadata Validation

    Validate metadata for:

    • Required fields
    • Data types
    • Length limits
    • Allowed values
    • Character encoding
    • Tenant scope
    • Identifier existence
    • Date and time format
    • Cross-field consistency
    • Authorization to set protected fields

    Cross-field Example

    CONSTRAINT ck_asset_publication
    CHECK
    (
        published_at IS NULL
        OR asset_status = 'published'
    );

    More complex lifecycle rules can require transactional application logic or database-specific mechanisms.

    Metadata Synchronization

    Metadata can exist in storage, a relational database, a search index, a cache, and analytics systems.

    Authoritative metadata database
          |
          +-- Object-storage metadata
          |
          +-- Search index
          |
          +-- Cache
          |
          +-- Analytics projection

    Each derived representation should define:

    • Source of truth
    • Update trigger
    • Acceptable delay
    • Retry strategy
    • Idempotency
    • Reconciliation process
    • Rebuild procedure

    Event-driven Metadata Projection

    Metadata transaction commits
            |
            v
    Outbox event becomes available
            |
            v
    Publisher sends event
            |
            v
    Search-index consumer updates document
            |
            v
    Cache or analytics projection updates

    Consumers should handle duplicate events safely and support replay or rebuild from authoritative metadata.

    Metadata Drift

    Metadata drift occurs when two representations no longer agree.

    Database metadata:
    
    asset_status = published
    
    
    Object missing:
    
    No object exists for stored key
    
    
    Search metadata:
    
    asset_status = processing
    
    
    Result:
    
    Three representations disagree.

    Drift can result from partial failures, delayed events, manual changes, failed processing, incorrect retries, or deleted objects.

    Metadata Reconciliation

    A reconciliation process compares authoritative metadata with actual stored objects and derived representations.

    Read authoritative asset records
            |
            v
    Check referenced storage objects
            |
            v
    Compare size, checksum, and version
            |
            v
    Check search or cache projection
            |
            +-- Match:
            |      mark healthy
            |
            +-- Mismatch:
                   repair, reindex,
                   quarantine, or investigate

    Find Metadata without Objects

    For each active metadata record:
    
    1. Resolve its object key.
    2. Request object properties.
    3. Verify existence.
    4. Compare object version, length, and checksum.
    5. Record or repair mismatches.

    Find Objects without Metadata

    For each stored object in managed prefix:
    
    1. Extract stable asset identifier.
    2. Find authoritative metadata record.
    3. Confirm tenant and lifecycle state.
    4. Quarantine or clean abandoned content
       according to policy.

    Metadata Deletion

    Deleting metadata before stored content can create an object that is no longer associated with an active application record. Deleting content first can leave a record pointing to missing bytes.

    Safer deletion workflow:
    
    1. Authorize deletion.
    2. Mark metadata as pending deletion.
    3. Apply retention and legal-hold checks.
    4. Delete or queue content deletion.
    5. Verify storage outcome.
    6. Remove derived search and cache entries.
    7. Finalize metadata state.
    8. Retain required audit evidence.

    Retention Metadata

    {
      "retentionClass": "course-content",
      "retainUntil": "approved-retention-date",
      "legalHold": false,
      "deletionStatus": "not-requested"
    }

    Retention values can have legal, privacy, audit, and operational impact. Only authorized workflows should modify them.

    Time Metadata

    Time fields should have clearly defined semantics.

    Field Meaning
    created_at Business metadata record creation time
    uploaded_at Object-upload completion time
    verified_at Integrity-verification completion time
    published_at Time content became published
    last_modified_at Time the tracked metadata or object changed according to the defined source
    archived_at Time content entered archived state

    Use a consistent time standard and retain the original timezone only when it has business meaning.

    Schema Evolution

    Metadata schemas change as applications add new fields, classifications, workflows, and processing outputs.

    Metadata schema version 1:
    
    title
    content_type
    object_key
    
    
    Metadata schema version 2:
    
    title
    content_type
    object_key
    language_code
    checksum
    visibility
    
    
    Metadata schema version 3:
    
    Adds processing and retention fields

    Plan for:

    • Optional fields during migration
    • Default values
    • Backfilling older records
    • Version-aware consumers
    • Search-index updates
    • Compatibility during rolling deployment
    • Removal of obsolete fields

    Flexible vs Structured Metadata

    Approach Benefit Consideration
    Structured columns Strong typing, validation, indexes, and queryability Schema changes require migrations
    Key-value metadata Supports varying optional attributes Can weaken consistency and increase query complexity
    JSON metadata Provides a flexible nested document Frequently queried fields can require validation and specialized indexing
    Separate related tables Supports repeatable and relational metadata values Requires joins and additional schema objects

    Use structured columns for critical and frequently queried business fields. Use flexible metadata deliberately for genuinely variable attributes.

    Multi-valued Metadata

    Metadata such as keywords, contributors, captions, and categories can contain several values.

    Comma-separated values
    keywords:
    
    system design,database,storage,cloud
    Related metadata rows
    CREATE TABLE asset_keywords
    (
        asset_id BIGINT NOT NULL,
        keyword VARCHAR(100) NOT NULL,
    
        PRIMARY KEY
        (
            asset_id,
            keyword
        ),
    
        CONSTRAINT fk_asset_keywords_asset
            FOREIGN KEY (asset_id)
            REFERENCES course_assets (asset_id)
            ON DELETE CASCADE
    );

    Derived Metadata

    Derived metadata is calculated or extracted from existing content or authoritative metadata.

    Examples include:

    • Video duration
    • Image dimensions
    • Document page count
    • Detected language
    • Generated thumbnail key
    • Searchable extracted text
    • Computed storage cost category
    • Latest processing result

    Derived metadata should record how and when it was generated when reproducibility or troubleshooting matters.

    Generated Metadata

    Automated systems can produce classifications, summaries, labels, or extracted entities.

    {
      "generator": "metadata-processing-service",
      "generatorVersion": "approved-version",
      "generatedAt": "processing-time",
      "classification": "technical-course-content",
      "confidence": 0.94,
      "reviewStatus": "pending"
    }

    Generated metadata should remain distinguishable from human-approved business metadata. Confidence, model or processor version, and review state can be important for governance.

    Metadata Performance

    A storage platform can contain many small metadata records even when the stored files are large. Metadata operations can become a distinct performance bottleneck.

    Important metadata operations include:

    • Lookup by object key
    • Directory listing
    • Object listing by prefix
    • Permission checks
    • Version lookup
    • Tag retrieval
    • Lifecycle scans
    • Search-index projection
    • Reconciliation

    Design metadata storage for its own query rate, update rate, consistency, index size, and failure behaviour.

    Metadata Service

    Clients
       |
       v
    Metadata API
       |
       +-- Metadata database
       |
       +-- Authorization service
       |
       +-- Search projection
       |
       +-- Object storage
       |
       +-- Processing events

    A dedicated metadata service can centralize validation, ownership, lifecycle, object references, security decisions, and change events.

    Whether a separate service is justified depends on scale, team boundaries, latency, consistency, and operational complexity.

    Metadata Observability

    Useful metadata metrics include:

    • Metadata record count
    • Records by lifecycle state
    • Metadata query latency
    • Metadata update failure count
    • Objects missing business metadata
    • Metadata records referencing missing objects
    • Checksum and length mismatches
    • Search-index synchronization delay
    • Metadata extraction failures
    • Quarantined asset count
    • Schema-version distribution
    • Reconciliation backlog
    • Unauthorized metadata-update attempts
    • Retention-policy exceptions

    Metadata Alert Conditions

    Alert when:

    • Published metadata points to missing content
    • Metadata extraction repeatedly fails
    • Search-projection delay exceeds its objective
    • Checksum or content-length drift is detected
    • Quarantined assets increase unexpectedly
    • Required metadata fields are missing
    • Retention or legal-hold fields change unexpectedly
    • Reconciliation cannot repair a mismatch
    • Metadata queries exceed latency thresholds

    Metadata Troubleshooting Workflow

    1. Identify the affected asset and stable identifier.
    2. Determine the authoritative metadata source.
    3. Retrieve the application metadata record.
    4. Retrieve the current storage-object properties.
    5. Compare object key, version, length, and checksum.
    6. Check lifecycle and publication state.
    7. Check tenant ownership and authorization fields.
    8. Check extraction and processing history.
    9. Check search and cache projections.
    10. Review recent metadata-change events.
    11. Repair derived copies from the authoritative source.
    12. Quarantine unresolved content or metadata conflicts.

    Common Metadata Mistakes

    1

    Storing Content without Business Metadata

    The application cannot reliably identify ownership, purpose, lifecycle, visibility, or relationships.

    2

    Using the Original Filename as the Storage Key

    Filenames can be untrusted, duplicated, mutable, excessively long, or incompatible with storage-key rules.

    3

    Trusting Client-supplied Metadata

    Clients can submit false content types, sizes, ownership, visibility, or workflow states.

    4

    Keeping No Source of Truth

    Storage, database, cache, and search metadata can diverge when no authoritative source is defined.

    5

    Mixing Technical and Business Meaning

    A storage object's last-modified time is not automatically the business publication time or course-update time.

    6

    Using Uncontrolled Tags

    Inconsistent spelling, capitalization, and terminology make querying and automation unreliable.

    7

    Placing Every Attribute in Flexible JSON

    Critical fields can lose strong typing, constraints, indexes, and clear ownership.

    8

    Duplicating Metadata without Synchronization

    Search, cache, analytics, storage, and relational copies can show contradictory values.

    9

    Publishing before Verification Completes

    Incomplete, corrupted, unsupported, or unsafe content can become visible to users.

    10

    Ignoring Embedded Sensitive Metadata

    Uploaded images and documents can disclose location, author, software, or internal-path information.

    11

    Deleting Metadata and Content Independently

    Partial failures can create orphaned objects or broken application references.

    12

    Assuming Metadata Is Too Small to Need Governance

    Metadata can expose sensitive business context and control consequential access, retention, and processing decisions.

    Recommended Test Cases

    Test Expected Evidence
    Required metadata Missing required fields are rejected
    Invalid controlled value Unsupported status, visibility, or asset type is rejected
    Untrusted filename The application generates a safe, stable storage key
    False client content type Server-side validation identifies the actual approved type
    Content-length mismatch The asset remains unpublished and enters failed or quarantined state
    Checksum mismatch The metadata records the integrity failure
    Missing object Reconciliation detects the broken metadata reference
    Orphaned object Reconciliation identifies content without an active metadata record
    Unauthorized metadata update Protected ownership, visibility, and retention fields are not changed
    Search synchronization Published metadata appears in the search projection within its objective
    Metadata revision The change is audited or versioned according to policy
    Deletion workflow Metadata, content, cache, and search state converge correctly

    Metadata Best Practices

    Recommended Practices

    • Separate stored content from the metadata that describes it.
    • Use stable asset identifiers and generated storage keys.
    • Retain original filenames only as validated display metadata.
    • Distinguish system, technical, business, operational, and security metadata.
    • Define an authoritative source for every important field.
    • Use structured columns for critical and frequently queried metadata.
    • Use flexible metadata only for genuinely variable attributes.
    • Validate all client-provided metadata.
    • Verify content type, length, checksum, and object version after upload.
    • Use controlled vocabularies for status, language, type, and classification.
    • Model multi-valued metadata through appropriate related structures.
    • Record provenance for important automated transformations.
    • Distinguish generated metadata from approved business metadata.
    • Protect metadata according to its sensitivity.
    • Remove embedded sensitive metadata when required.
    • Publish content only after required checks complete.
    • Synchronize search, cache, and analytics projections reliably.
    • Implement reconciliation between metadata and stored content.
    • Audit consequential metadata changes.
    • Monitor incomplete, inconsistent, unauthorized, and orphaned metadata.

    Practice Exercise

    Design metadata management for course assets on your online learning platform.

    Requirements

    1. Create metadata for videos, PDFs, images, presentations, and source-code files.
    2. Generate a stable asset ID and object key.
    3. Store the validated original filename separately.
    4. Associate each asset with a tenant, course, chapter, or lesson.
    5. Store content type, content length, checksum, and object version.
    6. Record language and accessibility metadata.
    7. Define pending, processing, published, failed, quarantined, and archived states.
    8. Record uploader and processing provenance.
    9. Extract technical metadata after upload.
    10. Prevent clients from assigning protected metadata fields.
    11. Project published metadata into the search index.
    12. Track metadata revisions.
    13. Detect database records referencing missing objects.
    14. Detect stored objects without metadata records.
    15. Test the complete deletion and retention workflow.

    Metadata-design Template

    Metadata Field Source of Truth Validation Purpose
    Asset ID Application database Primary key Stable business identity
    Object key Application-generated metadata Tenant-scoped uniqueness Storage identity
    Content type Approved upload validation Allowed media-type list Delivery and processing
    Content length Verified stored object Non-negative and within quota Integrity, quota, and transfer handling
    Checksum Verified checksum calculation Algorithm and value validation Content integrity
    Publication status Application workflow Controlled lifecycle transition Content visibility
    Search document Derived from authoritative metadata Projection and reconciliation Discovery and filtering

    Frequently Asked Questions

    1

    What is metadata?

    Metadata is information that describes, identifies, organizes, governs, or connects another data item or resource.

    2

    What is descriptive metadata?

    Descriptive metadata explains what content represents, using fields such as title, description, language, subject, and keywords.

    3

    What is structural metadata?

    Structural metadata describes how content components relate, such as course, chapter, lesson, page, and display-order relationships.

    4

    What is technical metadata?

    Technical metadata records properties such as format, content type, size, checksum, dimensions, duration, codec, and object version.

    5

    What is provenance metadata?

    Provenance metadata records the origin, processing history, and transformations applied to content.

    6

    What is object metadata?

    Object metadata describes a stored object through system-managed or application-defined properties such as key, size, content type, checksum, storage class, and custom values.

    7

    Should all metadata be stored in object storage?

    No. Business relationships, workflow, authorization, audit, and searchable structured fields commonly belong in an application database or dedicated metadata system.

    8

    Why should object keys be stable?

    Stable keys avoid storage changes when mutable display values such as filenames, titles, or course names change.

    9

    Can metadata be sensitive?

    Yes. Filenames, ownership, locations, titles, object keys, classifications, and processing history can reveal sensitive information.

    10

    How does metadata support search?

    Metadata provides structured fields for filtering, faceting, sorting, grouping, authorization, and ranking.

    11

    What is metadata drift?

    Metadata drift occurs when storage, database, search, cache, or analytics representations no longer agree.

    12

    What comes after metadata?

    The next topic is uploads and multipart transfer, followed by inverted indexes and full-text search.

    Key Takeaway

    Metadata gives stored content its identity, meaning, relationships, technical context, lifecycle, security classification, and searchability. Distinguish system-defined, user-defined, descriptive, structural, technical, administrative, operational, and provenance metadata. Use stable identifiers and generated object keys, validate all untrusted fields, define a source of truth for duplicated values, protect sensitive metadata, and publish content only after required verification completes. Keep business metadata in a constrained and queryable system where appropriate, synchronize derived search and cache representations reliably, and reconcile metadata records with actual stored objects.