Table of Contents

    checksums

    STORAGE, FILES, OBJECTS & SEARCH BASICS

    Checksums

    Learn how checksums detect accidental data corruption during storage and transmission, how CRC and cryptographic hashes differ, and how to verify files, objects, multipart uploads, backups, and replicated data safely.

    Introduction

    Data can change unexpectedly while it is being created, transferred, stored, copied, restored, or downloaded. A network interruption, damaged storage device, software defect, incomplete upload, or memory error can produce content that differs from the original.

    A checksum is a fixed-size value calculated from data using a defined algorithm. The checksum can be stored or transmitted separately and later compared with a newly calculated value.

    When the values match, the content passed that checksum verification. When they differ, the content, algorithm, encoding, or calculation scope does not match the expected input.

    Core idea: A checksum does not prevent corruption. It provides evidence that content has changed, allowing the system to reject, retry, repair, quarantine, or restore the affected data.

    In your System Design curriculum, Checksums is Topic 6.3 under Storage, Files, Objects & Search Basics. It follows durability and precedes metadata, uploads and multipart transfer, inverted indexes, and full-text search.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Block, file, and object storage Checksums can be calculated for blocks, files, objects, and transferred parts.
    2 Durability Integrity verification is required to detect durable but corrupted data.
    3 Binary data Checksums are calculated from the exact input bytes.
    4 File uploads Client and server checksums can verify transferred content.
    5 Basic security concepts Accidental corruption and intentional modification require different integrity mechanisms.
    6 Metadata The algorithm name and checksum value must be stored with sufficient context.

    What Is a Checksum?

    A checksum is the output produced when a checksum or hash algorithm processes a defined sequence of bytes.

    If \(D\) represents the input data and \(H\) represents the selected algorithm, the checksum can be expressed as:

    \[ C = H(D) \]

    Where:

    • \(D\) is the exact input data
    • \(H\) is the checksum or hash algorithm
    • \(C\) is the resulting checksum value
    Original data
          |
          v
    Checksum algorithm
          |
          v
    Fixed-size checksum value

    Even when a checksum is much smaller than the original data, it can provide useful evidence about whether the content changed.

    Checksum-verification Flow

    Before storage or transfer:
    
    Original bytes
          |
          v
    Calculate expected checksum
          |
          v
    Store or transmit data
    and checksum
    
    
    During verification:
    
    Received or stored bytes
          |
          v
    Calculate actual checksum
          |
          v
    Compare expected and actual values
          |
          +-- Match:
          |      integrity check passes
          |
          +-- Mismatch:
                 reject, retry, repair,
                 quarantine, or investigate

    Verification rule: Both calculations must use the same algorithm, exact byte sequence, checksum encoding, and calculation scope.

    What Checksums Can Detect

    Checksums can help detect:

    • Incomplete file transfers
    • Changed bytes
    • Truncated files
    • Corrupted storage blocks
    • Incorrectly assembled multipart uploads
    • Damaged backup content
    • Unexpected file replacement
    • Differences between replicas
    • Transmission errors

    What Checksums Do Not Provide Automatically

    A checksum alone does not automatically provide:

    • Encryption
    • Confidentiality
    • Authentication
    • Authorization
    • Proof of who created the content
    • Recovery of corrupted bytes
    • Protection from an attacker who can replace both data and checksum
    • Guaranteed uniqueness for every possible input

    Security rule: When an attacker can modify both the data and its ordinary checksum, the attacker can calculate a new checksum. Security-sensitive integrity requires an authenticated mechanism such as a keyed message authentication code or a verified digital signature.

    Checksum, CRC, Hash, HMAC, and Signature

    Mechanism Main Purpose Secret Required? Typical Use
    Simple checksum Basic accidental-error detection No Simple integrity checks
    CRC Efficient detection of common transmission and storage errors No Networks, storage, archives, and protocols
    Cryptographic hash Strong content fingerprinting and integrity comparison No Files, releases, objects, backups, and content addressing
    HMAC Integrity and authenticity using a shared secret Yes Authenticated API messages and protected metadata
    Digital signature Integrity and signer authenticity using asymmetric cryptography Private signing key Software releases, documents, and signed artifacts

    Cyclic Redundancy Check

    Cyclic Redundancy Check, commonly abbreviated as CRC, is a family of algorithms designed for efficient error detection.

    CRC algorithms are commonly used where accidental transmission or storage errors need to be detected efficiently.

    Examples include:

    • Network frames
    • Storage devices
    • Compressed archives
    • Object-transfer verification
    • Protocol messages

    CRC limitation: CRC is designed for error detection, not for resisting deliberate manipulation by an attacker.

    Cryptographic Hash Functions

    A cryptographic hash function maps input of arbitrary length to a fixed-size digest and is designed to make certain forms of reverse computation and collision construction difficult.

    Cryptographic hashes are useful for:

    • File-integrity verification
    • Software-artifact verification
    • Backup verification
    • Object-integrity metadata
    • Content-addressable storage
    • Digital-signature workflows
    • Deduplication candidates

    Common Algorithms

    Algorithm Output General Use Security Consideration
    CRC32 32 bits Efficient accidental-error detection Not a cryptographic integrity mechanism
    MD5 128 bits Legacy compatibility and non-adversarial checks Collision weaknesses make it unsuitable for security-sensitive verification
    SHA-1 160 bits Legacy compatibility Should not be selected for new security-sensitive designs
    SHA-256 256 bits General-purpose cryptographic content verification Widely supported, but authenticity still requires trusted checksum distribution or authentication
    SHA-512 512 bits Cryptographic content verification Produces a larger digest and still requires trusted verification context

    Algorithm support and performance depend on the storage service, runtime, hardware, interoperability requirements, and threat model.

    Avalanche Behaviour

    A strong cryptographic hash is designed so that a small change in the input produces a substantially different digest.

    Input A:
    
    Course lesson version 1
    
    
    Input B:
    
    Course lesson version 2
    
    
    Only one character changed,
    but the resulting cryptographic
    digests are substantially different.

    This helps detect small changes without comparing every byte manually.

    Checksum Comparison

    Let \(C_e\) be the expected checksum and \(C_a\) be the checksum calculated from the available content.

    \[ Verified = \begin{cases} True, & C_e = C_a \\ False, & C_e \neq C_a \end{cases} \]

    A matching checksum means that the verification calculation produced the expected value. It does not prove that the expected value was obtained from a trusted source.

    Calculate a SHA-256 Checksum on Linux

    sha256sum course-video.mp4

    Example output format:

    expected-sha256-value  course-video.mp4

    Verify with a Checksum File

    sha256sum -c course-video.sha256

    The checksum file must use the format expected by the selected command-line utility.

    Calculate a Checksum in PowerShell

    Get-FileHash `
        -Path "course-video.mp4" `
        -Algorithm SHA256

    Compare the returned value with the expected SHA-256 value from a trusted source.

    Calculate a File Checksum in PHP

    <?php
    
    declare(strict_types=1);
    
    function calculateSha256(
        string $filePath
    ): string {
        if (!is_file($filePath) ||
            !is_readable($filePath)) {
    
            throw new RuntimeException(
                'The file is not available for checksum calculation.'
            );
        }
    
        $checksum =
            hash_file(
                'sha256',
                $filePath
            );
    
        if ($checksum === false) {
            throw new RuntimeException(
                'The checksum could not be calculated.'
            );
        }
    
        return strtolower(
            $checksum
        );
    }

    hash_file() calculates the digest from the file without requiring the application to load the complete file into one PHP string.

    Verify a File in PHP

    <?php
    
    declare(strict_types=1);
    
    function verifyFileChecksum(
        string $filePath,
        string $expectedChecksum
    ): bool {
        $actualChecksum =
            calculateSha256(
                $filePath
            );
    
        return hash_equals(
            strtolower(
                trim(
                    $expectedChecksum
                )
            ),
            $actualChecksum
        );
    }

    The expected checksum should come from a trusted source and use the same algorithm and encoding as the actual value.

    Upload Integrity

    An upload workflow can calculate a checksum before transmission and verify it after the storage service receives the content.

    Client file
        |
        v
    Client calculates checksum
        |
        v
    Client uploads file and checksum
        |
        v
    Storage calculates checksum
        |
        v
    Storage compares values
        |
        +-- Match:
        |      accept object
        |
        +-- Mismatch:
               reject upload

    This detects content that changed during transmission or storage processing.

    Conceptual HTTP Upload

    PUT /course-assets/video.mp4 HTTP/1.1
    Host: storage.example.com
    Content-Type: video/mp4
    Content-Length: 52428800
    X-Checksum-Algorithm: SHA-256
    X-Checksum-Value: expected-checksum-value
    
    [BINARY CONTENT]

    Header names and supported algorithms depend on the selected API. Follow the storage provider's documented checksum contract.

    Download Integrity

    Download object
          |
          v
    Receive expected stored checksum
          |
          v
    Calculate checksum from downloaded bytes
          |
          v
    Compare values
          |
          +-- Match:
          |      use content
          |
          +-- Mismatch:
                 discard or quarantine
                 retry from trusted source
                 record integrity failure

    The application should not publish or process content after a failed integrity check.

    Object-storage Checksums

    An object-storage service can accept a client-provided checksum, calculate its own checksum, compare the values, and store the verified checksum as object metadata.

    {
      "objectKey": "courses/42/lessons/7/video.mp4",
      "checksumAlgorithm": "SHA-256",
      "checksumValue": "expected-checksum-value",
      "contentLength": 52428800,
      "verificationStatus": "verified"
    }

    The exact algorithm names, checksum encodings, metadata fields, and multipart behaviour are provider-specific.

    Multipart-upload Checksums

    A large object can be split into parts and uploaded independently.

    Large file
        |
        +-- Part 1 -> checksum 1
        |
        +-- Part 2 -> checksum 2
        |
        +-- Part 3 -> checksum 3
        |
        +-- Part 4 -> checksum 4
        |
        v
    Complete multipart upload
        |
        v
    Verify final object integrity

    Multipart verification can involve:

    • A checksum for every uploaded part
    • A final full-object checksum
    • A provider-defined composite checksum
    • Verification during finalization

    Multipart rule: A multipart object's identifier or ETag must not be assumed to equal the checksum of the complete object. Verify the exact provider algorithm and multipart semantics.

    Store Checksum Metadata

    CREATE TABLE course_assets
    (
        asset_id BIGINT PRIMARY KEY,
        tenant_id BIGINT NOT NULL,
        object_key VARCHAR(500) NOT NULL,
        original_file_name VARCHAR(255) NOT NULL,
        content_type VARCHAR(100) NOT NULL,
        content_length BIGINT NOT NULL,
        checksum_algorithm VARCHAR(30) NOT NULL,
        checksum_value VARCHAR(200) NOT NULL,
        verification_status VARCHAR(30) NOT NULL,
        verified_at TIMESTAMP NULL,
        created_at TIMESTAMP NOT NULL,
    
        CONSTRAINT uq_course_assets_object
            UNIQUE
            (
                tenant_id,
                object_key
            ),
    
        CONSTRAINT ck_asset_length
            CHECK (content_length >= 0),
    
        CONSTRAINT ck_checksum_algorithm
            CHECK
            (
                checksum_algorithm IN
                (
                    'CRC32',
                    'SHA-256',
                    'SHA-512'
                )
            ),
    
        CONSTRAINT ck_verification_status
            CHECK
            (
                verification_status IN
                (
                    'pending',
                    'verified',
                    'failed',
                    'quarantined'
                )
            )
    );

    Store both the algorithm and value. A checksum value without its algorithm, byte scope, and encoding is ambiguous.

    Checksum Metadata Requirements

    Useful checksum metadata includes:

    • Algorithm name
    • Checksum value
    • Encoding format
    • Content length
    • Calculation scope
    • Part number for multipart data
    • Verification time
    • Verification result
    • Object version
    • Source of the expected checksum

    Streaming Large Files

    Large files should be processed incrementally instead of being loaded completely into memory.

    Open file stream
          |
          v
    Read chunk
          |
          v
    Update checksum state
          |
          v
    Repeat until end of file
          |
          v
    Finalize checksum

    Chunked PHP Hashing

    <?php
    
    declare(strict_types=1);
    
    function calculateSha256Streamed(
        string $filePath
    ): string {
        $stream =
            fopen(
                $filePath,
                'rb'
            );
    
        if ($stream === false) {
            throw new RuntimeException(
                'The file could not be opened.'
            );
        }
    
        $context =
            hash_init(
                'sha256'
            );
    
        try {
            while (!feof($stream)) {
                $chunk =
                    fread(
                        $stream,
                        1024 * 1024
                    );
    
                if ($chunk === false) {
                    throw new RuntimeException(
                        'The file could not be read.'
                    );
                }
    
                if ($chunk !== '') {
                    hash_update(
                        $context,
                        $chunk
                    );
                }
            }
    
            return hash_final(
                $context
            );
        } finally {
            fclose(
                $stream
            );
        }
    }

    The chunk size affects application memory and I/O behaviour but does not change the final hash when the same bytes are processed in the same order.

    Text Encoding Matters

    Two text values can look identical to a person while containing different byte sequences.

    Possible differences:
    
    - UTF-8 vs UTF-16
    - CRLF vs LF line endings
    - Unicode normalization
    - Byte-order marker
    - Trailing whitespace
    - Final newline
    - Character case
    - JSON property order

    Checksums operate on bytes. If two systems need the same checksum for a logical document, both must define a canonical byte representation.

    Canonicalization

    Canonicalization converts logically equivalent data into one defined byte representation before hashing.

    Logical data
          |
          v
    Apply canonical representation
          |
          +-- Defined character encoding
          +-- Defined line endings
          +-- Defined whitespace rules
          +-- Defined property ordering
          +-- Defined numeric formatting
          |
          v
    Calculate checksum

    Do not canonicalize arbitrary binary files unless the format and contract explicitly require it.

    JSON Checksum Example

    These JSON documents can represent similar information:

    {
      "courseId": 42,
      "title": "System Design"
    }
    {"title":"System Design","courseId":42}

    Their raw byte sequences differ, so their raw checksums differ. If a stable logical-data checksum is required, define a canonical JSON representation.

    Checksums and Deduplication

    A checksum can identify candidate duplicate content.

    New upload
        |
        v
    Calculate checksum
        |
        v
    Search for matching checksum
        |
        +-- No match:
        |      store new content
        |
        +-- Match:
               compare size and,
               when required,
               verify complete content

    A matching checksum should not automatically be treated as proof that two untrusted files are identical. Collision risk, tenant boundaries, privacy, authorization, encryption, and content ownership must be considered.

    Privacy rule: Cross-tenant deduplication can reveal whether another tenant stores the same content. Apply tenant isolation and privacy requirements before using checksums for global deduplication.

    Content-addressable Storage

    In content-addressable storage, content can be identified by a digest calculated from its bytes.

    Content bytes
          |
          v
    Cryptographic hash
          |
          v
    Content identifier
          |
          v
    Store or retrieve content by digest

    This can support immutable artifacts and duplicate detection, but the system still needs authorization, metadata, lifecycle, and collision-handling policies.

    Replica Integrity

    Checksums can help compare replicated blocks or objects.

    Replica A checksum
            |
            v
    Compare
            ^
            |
    Replica B checksum
    
    
    Match:
    
    Replicas contain the same verified bytes.
    
    
    Mismatch:
    
    At least one copy differs.
    Locate a trusted valid copy
    before repairing data.

    The system must not repair a good copy from a corrupted copy merely because the corrupted copy is newer or closer.

    Data Scrubbing

    Data scrubbing periodically reads stored data, recalculates checksums, and repairs corrupted content using verified redundant data.

    Scheduled integrity scan
          |
          v
    Read stored content
          |
          v
    Calculate checksum
          |
          v
    Compare with expected value
          |
          +-- Match:
          |      mark healthy
          |
          +-- Mismatch:
                 quarantine bad copy
                 find healthy replica
                 repair content
                 record incident

    Without periodic verification, silent corruption can remain undiscovered until the content is needed or every redundant copy has degraded.

    Backup Verification

    A successful backup operation does not prove that every backed-up item can be restored correctly.

    Create backup
          |
          v
    Store expected checksums
          |
          v
    Read backup periodically
          |
          v
    Recalculate checksums
          |
          v
    Compare values
          |
          v
    Perform restore test

    Backup verification should include both byte-level integrity and application-level usability.

    Checksum Mismatch Handling

    A checksum mismatch should trigger a controlled response.

    Checksum mismatch
          |
          v
    Stop processing content
          |
          v
    Record algorithm, source,
    object version, and trace ID
          |
          v
    Classify probable cause
          |
          +-- Transfer issue:
          |      retry from trusted source
          |
          +-- Damaged replica:
          |      repair from healthy copy
          |
          +-- Unknown or suspicious:
                 quarantine
                 investigate
                 preserve evidence

    Do not silently replace the expected checksum with the newly calculated value. Doing so hides the integrity failure.

    Intentional Modification

    An ordinary unkeyed checksum detects modification only when the expected checksum itself is trusted.

    Attacker changes file
          |
          v
    Attacker calculates new checksum
          |
          v
    Attacker replaces stored checksum
          |
          v
    Ordinary comparison succeeds

    Security-sensitive verification can use:

    • A checksum delivered through a trusted independent channel
    • An HMAC calculated with a protected shared secret
    • A digital signature verified with a trusted public key
    • Protected immutable metadata

    HMAC

    A Hash-based Message Authentication Code combines data with a secret key to provide integrity and authenticity between trusted parties sharing the key.

    Conceptually:

    \[ Tag = HMAC(SecretKey, Message) \]

    A party without the secret key should not be able to generate a valid tag for modified content.

    Key-management rule: HMAC security depends on protecting, rotating, and scoping the secret key appropriately.

    Digital Signatures

    A digital-signature workflow calculates a digest and signs it using a private key. A verifier uses the corresponding trusted public key to verify the signature.

    Publisher:
    
    Content
      |
      v
    Hash
      |
      v
    Sign digest with private key
      |
      v
    Publish content and signature
    
    
    Verifier:
    
    Download content and signature
      |
      v
    Calculate content hash
      |
      v
    Verify signature with trusted public key

    Digital signatures can provide stronger publisher-authenticity guarantees than publishing an unsigned checksum beside a downloadable file.

    Performance Considerations

    Checksum calculation consumes CPU and requires reading the protected bytes.

    Performance depends on:

    • Algorithm
    • File or object size
    • Storage throughput
    • Memory and chunk size
    • Hardware acceleration
    • Number of concurrent calculations
    • Whether the checksum is calculated during transfer

    Avoid reading a very large file an additional time when the checksum can be calculated incrementally during the existing upload or download stream.

    One-pass Processing

    Input stream
          |
          +-- Write content to storage
          |
          +-- Update checksum state
          |
          +-- Count bytes
          |
          v
    Finalize storage and checksum together

    One-pass processing can reduce repeated I/O, provided that failures and finalization are handled correctly.

    Checksum Observability

    Useful checksum metrics include:

    • Checksum calculations by algorithm
    • Verification success count
    • Verification failure count
    • Upload checksum mismatch count
    • Download checksum mismatch count
    • Replica-integrity mismatch count
    • Backup-verification failure count
    • Scrubbing coverage
    • Repair count
    • Quarantined-object count
    • Checksum-calculation duration
    • Bytes verified
    • Objects missing checksum metadata

    Alert Conditions

    Alert when:

    • Checksum mismatches increase unexpectedly
    • A storage replica fails integrity verification
    • Backups cannot be verified
    • Objects are stored without required checksum metadata
    • The repair backlog grows
    • Multipart completion fails integrity verification
    • One storage device or path produces repeated mismatches
    • Scrubbing does not complete within the approved cycle

    Checksum Troubleshooting Workflow

    1. Identify the affected file, object, block, or upload part.
    2. Identify the checksum algorithm.
    3. Confirm checksum encoding and letter-case handling.
    4. Confirm the calculation scope.
    5. Confirm the expected content length.
    6. Calculate the checksum from the available bytes.
    7. Compare the expected and actual values exactly.
    8. Check whether text encoding or line endings changed.
    9. Check whether content was compressed, encrypted, or transformed.
    10. Check multipart checksum semantics.
    11. Retry from a trusted source when appropriate.
    12. Compare redundant copies.
    13. Quarantine unresolved mismatches.
    14. Record and investigate repeated failures.

    Common Checksum Mistakes

    1

    Using Different Algorithms

    A SHA-256 digest cannot be compared with a CRC32, MD5, SHA-1, or SHA-512 value.

    2

    Comparing Different Byte Representations

    Encoding, line endings, compression, encryption, whitespace, or serialization differences produce different checksum inputs.

    3

    Storing the Value without the Algorithm

    The checksum value is ambiguous when the algorithm and encoding are unknown.

    4

    Using CRC for Adversarial Integrity

    CRC is designed for efficient error detection rather than protection from intentional manipulation.

    5

    Using MD5 or SHA-1 for New Security-sensitive Verification

    Known collision weaknesses make these algorithms unsuitable for new security-sensitive verification designs.

    6

    Trusting an Unauthenticated Checksum

    An attacker who replaces both content and checksum can make an ordinary integrity comparison succeed.

    7

    Assuming a Matching Checksum Proves Identity

    A checksum comparison does not prove who created or authorized the content.

    8

    Assuming an Object ETag Is Always a File Checksum

    Encryption, multipart upload, and provider-specific behaviour can give an ETag semantics different from a complete-object checksum.

    9

    Loading Large Files Fully into Memory

    Use file or stream hashing to process large content incrementally.

    10

    Replacing the Expected Checksum after a Mismatch

    Updating metadata with the newly calculated value hides the integrity failure instead of resolving it.

    11

    Verifying Uploads but Not Stored Data

    Content can become corrupted after upload, so important stored data can require periodic verification.

    12

    Using Checksums without a Repair Strategy

    Detection alone does not recover damaged data. Maintain verified redundant or backup copies where recovery is required.

    Recommended Test Cases

    Test Expected Evidence
    Unchanged file The expected and actual checksums match
    One-byte modification The checksum comparison fails
    Truncated file Length and checksum verification fail
    Wrong algorithm The verification contract rejects the comparison
    Different line endings Raw-byte checksums differ
    Large streamed file The digest matches without loading the complete file into memory
    Upload corruption The storage service or application rejects the upload
    Download corruption The client rejects or quarantines the downloaded content
    Multipart part corruption The affected part fails verification and can be retried
    Replica mismatch The damaged copy is identified and repaired from verified data
    Backup verification The backup content matches retained checksums
    Tampered checksum metadata Authenticated verification detects or prevents unauthorized replacement

    Checksum Best Practices

    Recommended Practices

    • Define the purpose and threat model before choosing an algorithm.
    • Use CRC for appropriate accidental-error detection.
    • Use a modern cryptographic hash for strong content fingerprinting.
    • Use HMAC or signatures when authenticity is required.
    • Store the algorithm name with every checksum value.
    • Define checksum encoding and calculation scope.
    • Calculate checksums from the exact byte representation.
    • Use canonicalization only under a clearly defined data contract.
    • Hash large files incrementally.
    • Calculate checksums during existing transfer streams where practical.
    • Verify both content length and checksum.
    • Verify uploaded data before marking business metadata active.
    • Verify downloaded data before processing or publishing it.
    • Follow provider-specific multipart checksum semantics.
    • Do not assume that an ETag is a complete-object checksum.
    • Keep the expected checksum in a trusted location.
    • Quarantine unresolved integrity failures.
    • Maintain verified redundant or backup copies for repair.
    • Periodically scrub important stored data.
    • Monitor mismatches, verification coverage, repairs, and missing metadata.

    Practice Exercise

    Add checksum verification to the course-asset upload workflow for your online learning platform.

    Requirements

    1. Accept a video, PDF, image, or source-code asset.
    2. Validate the file size and content type.
    3. Generate a stable object key.
    4. Calculate a SHA-256 checksum while streaming the upload.
    5. Store the checksum algorithm and value.
    6. Verify the stored object's content length.
    7. Verify the storage-service checksum where supported.
    8. Mark the asset verified only after successful comparison.
    9. Quarantine a mismatched upload.
    10. Support checksums for multipart upload parts.
    11. Store a final complete-object checksum.
    12. Verify the checksum during download testing.
    13. Run a scheduled integrity scan.
    14. Repair a corrupted copy from a verified source.
    15. Record verification failures and corrective actions.

    Verification-state Template

    State Meaning Allowed Action
    Pending Upload or checksum verification is incomplete Continue or retry verification
    Verified Expected and actual checksum values match Publish according to authorization
    Failed Checksum, length, or upload verification failed Reject and investigate
    Quarantined The content is isolated pending investigation Restricted administrative access only
    Repaired A corrupted copy was replaced from verified content Reverify before normal use

    Suggested Finalization Logic

    Receive upload completion
            |
            v
    Retrieve stored object metadata
            |
            v
    Compare expected content length
            |
            v
    Compare expected checksum
            |
            +-- Both match:
            |       mark asset verified
            |
            +-- Any mismatch:
                    mark asset failed
                    quarantine object
                    record integrity event

    Frequently Asked Questions

    1

    What is a checksum?

    A checksum is a fixed-size value calculated from data and used to verify whether the available bytes match the expected content.

    2

    Does a checksum prevent corruption?

    No. A checksum detects a mismatch. Redundancy, retry, repair, or recovery is required to restore damaged data.

    3

    What is CRC?

    CRC is a family of efficient error-detection algorithms commonly used in storage and communication systems.

    4

    What is the difference between CRC and SHA-256?

    CRC is primarily designed for efficient accidental-error detection. SHA-256 is a cryptographic hash suitable for stronger content fingerprinting.

    5

    Can MD5 be used for security-sensitive verification?

    It should not be selected for new security-sensitive verification because of known collision weaknesses.

    6

    Why can two visually identical text files have different checksums?

    Their byte representation can differ because of encoding, line endings, normalization, whitespace, or hidden markers.

    7

    Does a matching checksum prove who created a file?

    No. Publisher authenticity requires a trusted checksum channel, HMAC, or digital-signature mechanism.

    8

    How should large files be checksummed?

    Process the file incrementally as a stream so the complete file does not need to be loaded into memory.

    9

    Is an object-storage ETag always a checksum?

    No. Its meaning can depend on provider behaviour, encryption, and multipart-upload processing.

    10

    Should checksum values be stored in a database?

    They can be stored with the algorithm, object identifier, content length, object version, and verification state as part of the business metadata.

    11

    What should happen after a checksum mismatch?

    Stop processing the content, record the mismatch, retry or repair from a trusted source where appropriate, and quarantine unresolved content.

    12

    What comes after checksums?

    The next topic is metadata, followed by uploads and multipart transfer, inverted indexes, and full-text search.

    Key Takeaway

    Checksums provide a compact method for detecting whether stored or transferred bytes differ from the expected content. Use CRC for suitable accidental-error detection and a modern cryptographic hash such as SHA-256 for stronger content fingerprinting. Store the algorithm, value, byte scope, content length, and verification state together. Hash large content incrementally, verify uploads and downloads, follow provider- specific multipart rules, and never assume that an ETag is a complete- object checksum. For adversarial integrity, protect the expected digest through HMAC, signatures, or another authenticated channel. When a mismatch occurs, reject or quarantine the content and recover from a verified redundant or backup copy.