Table of Contents

    polyglot persistence

    NOSQL & ACCESS-PATTERN DESIGN

    Polyglot Persistence

    Learn how one application can use relational, document, key-value, wide-column, graph, search, object-storage, and analytical systems for different workloads, while managing ownership, consistency, synchronization, transactions, security, observability, backup, recovery, migration, and operational complexity.

    Introduction

    Applications commonly manage several forms of data with substantially different access patterns.

    An online learning platform can require:

    • Transactional enrollment and payment records
    • Flexible course-catalog documents
    • Fast session and rate-limit lookups
    • Large video and document storage
    • Full-text article and course search
    • High-volume activity-event ingestion
    • Course-prerequisite and topic-relationship traversal
    • Historical reporting and analytics

    One database can sometimes support all these requirements. However, forcing every workload into one storage model can create inefficient queries, complicated schemas, scaling limitations, or unnecessary operational compromises.

    Polyglot persistence is the deliberate use of multiple data-storage technologies within one application or system, with each datastore selected for a specific data model and access pattern.

    Core idea: Use different datastores only when different workloads have materially different requirements. Choose each datastore for a defined purpose, assign clear data ownership, and accept the resulting consistency and operational responsibilities.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Relational databases Relational storage remains a strong choice for transactional data, constraints, and complex relationships.
    2 NoSQL data models Key-value, document, wide-column, and graph databases solve different access-pattern problems.
    3 Denormalization Derived stores and read models commonly duplicate authoritative data.
    4 Secondary indexes An additional index can sometimes solve a query without introducing another datastore.
    5 Transactions and consistency Independent datastores do not normally share one ordinary local transaction.
    6 Event-driven architecture Events frequently synchronize data between authoritative and derived stores.
    7 Observability and recovery Every datastore adds monitoring, backup, restoration, and incident-response requirements.

    Why Is It Called Polyglot?

    A polyglot person can use several languages and select the appropriate language for a particular situation.

    Polyglot persistence applies the same idea to data storage. Instead of requiring one database technology to solve every storage problem, the architecture can select a suitable datastore for each workload.

    Application workloads
    
    +-- Transaction processing
    |      -> Relational database
    |
    +-- Flexible content
    |      -> Document database
    |
    +-- Sessions and caching
    |      -> Key-value store
    |
    +-- Activity timelines
    |      -> Wide-column database
    |
    +-- Relationship traversal
    |      -> Graph database
    |
    +-- Keyword retrieval
    |      -> Search index
    |
    +-- Videos and documents
           -> Object storage
    Polyglot Design Flow
    identify workloads → define access patterns → choose the simplest valid store → assign ownership → design synchronization → operate and recover every store

    Single-database vs Polyglot Persistence

    Area Single-database Approach Polyglot Persistence
    Technology count One primary database technology Several purpose-specific datastores
    Operational complexity Generally lower Higher because every store must be operated
    Transaction management Can remain inside one database transaction Cross-store operations require distributed workflow design
    Workload specialization One engine handles several workload types Each engine can focus on a dominant access pattern
    Skills and tooling Smaller technology surface More expertise, clients, monitoring, and automation required
    Data duplication Can remain lower Derived and query-specific copies are common
    Failure modes Fewer integration boundaries Partial failures and synchronization lag must be managed

    Workload-first Selection

    Polyglot persistence should begin with workloads, not product names.

    For each workload, identify:

    • Data structure
    • Primary reads
    • Primary writes
    • Expected scale
    • Ordering and range requirements
    • Transaction boundaries
    • Consistency requirements
    • Latency and availability objectives
    • Retention and recovery requirements
    • Team operational capability

    Example Workloads

    Workload Dominant Requirement Candidate Datastore
    Enrollment Constraints and transactional updates Relational database
    Course catalog Flexible nested course descriptions Document database
    User session Fast retrieval by exact session key Key-value store
    Activity stream High-volume time-ordered writes Wide-column or event-oriented store
    Course prerequisites Multi-hop relationship traversal Graph database
    Article search Keyword retrieval, ranking, and highlighting Full-text search engine
    Course video Large durable binary objects Object storage
    Historical reporting Large scans and aggregation Analytical warehouse or lakehouse

    Selection rule: A different data model does not automatically require a different datastore. First verify whether an existing approved database can satisfy the workload through suitable schema, indexes, partitioning, or built-in capabilities.

    Relational Database Role

    A relational database is commonly used for data requiring constraints, relationships, transactional updates, and flexible structured queries.

    CREATE TABLE enrollments
    (
        tenant_id BIGINT NOT NULL,
        learner_id BIGINT NOT NULL,
        course_id BIGINT NOT NULL,
        enrollment_status VARCHAR(30) NOT NULL,
        enrolled_at TIMESTAMP NOT NULL,
    
        PRIMARY KEY
        (
            tenant_id,
            learner_id,
            course_id
        ),
    
        CONSTRAINT fk_enrollment_course
            FOREIGN KEY (course_id)
            REFERENCES courses (course_id)
    );

    Suitable responsibilities can include:

    • Accounts
    • Enrollments
    • Quiz attempts
    • Certificates
    • Orders and payments
    • Authoritative workflow state

    Document Database Role

    A document database can store flexible, self-contained aggregates.

    {
      "tenantId": 17,
      "courseId": 42,
      "title": "System Design",
      "difficulty": "beginner",
      "instructor": {
        "instructorId": 81,
        "displayName": "Course Instructor"
      },
      "tags": [
        "database",
        "architecture",
        "scalability"
      ],
      "presentation": {
        "thumbnailKey": "thumbnail-object-key",
        "summary": "Learn practical system design."
      }
    }

    Use bounded documents and define ownership for embedded and duplicated fields.

    Key-Value Store Role

    A key-value store is suitable when the application knows the exact key.

    session:8f4a2
        -> session data
    
    
    rate-limit:tenant-17:user-1042
        -> request counter
    
    
    course-card:tenant-17:course-42
        -> cached course card

    Typical uses include:

    • Sessions
    • Caching
    • Rate-limit state
    • Idempotency records
    • Temporary workflow state
    • Short-lived tokens

    Wide-Column Store Role

    A wide-column database can serve high-volume access patterns based on known partition and ordering keys.

    Partition key:
    
    tenantId
    +
    courseId
    +
    activityDate
    +
    writeBucket
    
    
    Clustering order:
    
    activityTime
    activityId

    This supports retrieving course activity within a bounded date partition, while controlled write buckets can distribute unusually heavy writes.

    Graph Database Role

    A graph database can model entities and their explicit relationships.

    [Course]
        |
        | REQUIRES
        v
    [Course]
        |
        | COVERS
        v
    [Topic]
        |
        | RELATED_TO
        v
    [Topic]

    Graph-oriented access patterns can include:

    • Course prerequisites
    • Topic relationships
    • Learning paths
    • Knowledge graphs
    • Dependency analysis
    • Bounded recommendation traversal

    Search Engine Role

    A search engine maintains query-optimized documents and inverted indexes.

    {
      "documentId": "tenant-17:article-981",
      "title": "Polyglot Persistence",
      "description": "Learn how applications use multiple datastores.",
      "body": "Searchable article content",
      "course": "System Design",
      "status": "published",
      "sourceVersion": 12
    }

    The search index is generally a derived representation, not the authoritative source of article or publication state.

    Object Storage Role

    Object storage is suitable for large files and unstructured binary content.

    Object storage:
    
    tenants/17/courses/42/videos/lesson-7.mp4
    
    tenants/17/courses/42/documents/chapter-3.pdf
    
    tenants/17/courses/42/images/thumbnail.webp

    Application metadata should connect each object key and version to its business asset, lifecycle state, ownership, checksum, and authorization rules.

    Analytical Store Role

    An analytical datastore can hold prepared historical data for reporting and large aggregations.

    Operational events
          |
          v
    Validated ingestion pipeline
          |
          v
    Prepared analytical tables
          |
          v
    Dashboards and reports

    Analytical copies should have defined lineage, freshness, retention, quality, and access rules.

    Example Architecture

    Clients
       |
       v
    Application APIs
       |
       +-- Enrollment Service
       |      -> Relational database
       |
       +-- Catalog Service
       |      -> Document database
       |
       +-- Session Service
       |      -> Key-value store
       |
       +-- Activity Service
       |      -> Wide-column store
       |
       +-- Learning Graph Service
       |      -> Graph database
       |
       +-- Search Service
       |      -> Search index
       |
       +-- Content Service
              -> Object storage
    
    Domain events
       |
       +-- Update search projections
       +-- Update caches
       +-- Update activity summaries
       +-- Feed analytics platform

    Authoritative Ownership

    Every business fact should have one authoritative owner.

    Fact Authoritative Owner Derived Copies
    Enrollment status Enrollment service and relational database Cache, search filter, analytics dataset
    Course title Course-management domain Catalog document, search document, course-card cache
    Video object Content service and object metadata Delivery-cache entries and search metadata
    Article-search terms Derived from authoritative article content Search index only
    Daily activity total Derived from authoritative activity events Dashboard and analytical summaries

    Ownership rule: Polyglot persistence must not create several independently editable sources of truth for the same business fact. Define ownership and update direction explicitly.

    Avoid Shared-database Ownership

    Services should not casually update one another's datastore.

    Unsafe coupling
    Catalog Service
        directly updates
    Enrollment Service tables
    
    
    Search Service
        directly changes
    Course Service documents
    Explicit ownership
    Service owns its data
          |
          v
    Other services use:
    
    - Published API
    - Approved event
    - Replicated read model

    Direct database access can bypass validation, authorization, invariants, auditing, and lifecycle logic.

    Data Synchronization

    Data can move between datastores through synchronous calls, durable events, change-data capture, batch pipelines, or controlled rebuild processes.

    Method Suitable Direction Main Consideration
    Synchronous API call Immediate dependency between services Availability and latency become coupled
    Domain event Update derived stores asynchronously Temporary staleness and delivery handling
    Change-data capture Publish committed database changes Schema interpretation and operational tooling
    Batch pipeline Historical analytics and periodic summaries Longer freshness delay
    Full rebuild Recreate derived projections Time, capacity, and cutover strategy

    Event-driven Synchronization

    {
      "eventType": "CoursePublished",
      "eventVersion": 1,
      "eventId": "generated-event-id",
      "tenantId": 17,
      "courseId": 42,
      "sourceVersion": 12,
      "occurredAt": "event-time"
    }
    Course transaction commits
          |
          v
    Outbox event becomes available
          |
          v
    Event is published
          |
          +-- Search consumer updates index
          +-- Cache consumer invalidates course card
          +-- Analytics consumer records publication event
          +-- Recommendation consumer updates relationship data

    Transactional Outbox

    The transactional outbox records a business change and its event in the same local database transaction.

    BEGIN;
    
    UPDATE courses
    SET
        course_status = 'published',
        source_version = source_version + 1
    WHERE tenant_id = :tenant_id
      AND course_id = :course_id;
    
    INSERT INTO outbox_events
    (
        event_id,
        event_type,
        aggregate_id,
        source_version,
        payload,
        created_at
    )
    VALUES
    (
        :event_id,
        'CoursePublished',
        :course_id,
        :source_version,
        :payload,
        CURRENT_TIMESTAMP
    );
    
    COMMIT;

    A publisher subsequently delivers unsent outbox records. Consumers should still be idempotent because delivery can occur more than once.

    Idempotent Consumers

    Event received
          |
          v
    Check event ID and source version
          |
          +-- Already processed:
          |      return previous outcome
          |
          +-- Older than current:
          |      reject stale event
          |
          +-- New event:
                 update projection
                 record processed event

    Stable event IDs, aggregate IDs, source versions, and projection versions help prevent duplicated or stale updates.

    Cross-database Transactions

    Independent datastores do not normally share one ordinary local ACID transaction.

    Relational update succeeds
          |
          v
    Graph update fails
          |
          v
    Search update remains pending
    
    
    Result:
    
    The system is temporarily inconsistent.

    Possible architectural responses include:

    • Keep the business transaction in one authoritative datastore
    • Update other stores asynchronously
    • Use idempotent retries
    • Use compensating business actions where appropriate
    • Maintain workflow state
    • Reconcile derived stores

    Transaction rule: Do not split one strongly consistent business invariant across several databases unless the distributed coordination and failure behaviour are explicitly designed and justified.

    Saga-style Workflow

    A saga coordinates a multi-step business process through local transactions and compensating actions.

    Create enrollment
          |
          v
    Reserve payment
          |
          v
    Provision course access
          |
          v
    Send confirmation
    
    
    If provisioning fails:
    
    Release payment reservation
    Mark enrollment failed
    Record workflow outcome

    Compensation is a business action, not a guaranteed reversal of every technical side effect.

    Eventual Consistency

    Derived datastores can temporarily display older information after the authoritative transaction commits.

    Time A:
    
    Course title updated
    in authoritative store
    
    
    Time B:
    
    Search document updated
    
    
    Between A and B:
    
    Direct course read:
    New title
    
    
    Search result:
    Old title

    Define acceptable freshness separately for each projection.

    Data Possible Consistency Direction
    Search title Short documented indexing delay can be acceptable
    Course-view counter Eventually consistent aggregate can be acceptable
    Payment balance Authoritative transactional consistency is required
    Publication and authorization Requires strongly bounded or authoritative enforcement

    Authorization across Datastores

    Security must apply consistently across primary and derived stores.

    Search request
          |
          v
    Apply trusted tenant scope
          |
          v
    Apply current publication and visibility rules
          |
          v
    Return authorized search results

    A stale search index or cache must not expose content after access is revoked.

    Protect:

    • Database records
    • Document copies
    • Cache entries
    • Search snippets
    • Graph relationships
    • Analytical datasets
    • Object metadata
    • Events and dead-letter records
    • Backups and exports

    Tenant Isolation

    Relational key:
    
    tenant_id + enrollment_id
    
    
    Document identity:
    
    tenant-17:course-42
    
    
    Cache key:
    
    tenant-17:course-card:42
    
    
    Search filter:
    
    trustedTenantId = 17
    
    
    Graph property:
    
    tenantId = 17

    Tenant scope should come from authenticated server-side context and remain present in every applicable data access path.

    Reconciliation

    Reconciliation detects differences between authoritative and derived stores.

    Read authoritative course
          |
          v
    Calculate expected projections
          |
          +-- Catalog document
          +-- Search document
          +-- Cache version
          +-- Graph relationships
          |
          v
    Compare IDs and versions
          |
          +-- Match:
          |      mark healthy
          |
          +-- Mismatch:
                 repair or rebuild
                 record diagnostic evidence

    Reconciliation should use stable identifiers and source versions instead of comparing only display values.

    Rebuildable Derived Stores

    Search indexes, caches, recommendation projections, and analytical summaries should be rebuildable from authoritative data wherever practical.

    Read authoritative records
          |
          v
    Create replacement projection
          |
          v
    Validate counts, versions, and sample queries
          |
          v
    Switch reads to replacement
          |
          v
    Retire old projection safely

    Rebuild capability reduces the long-term risk of missed events, mapping defects, and schema changes.

    Deletion Propagation

    Deleting authoritative data does not automatically delete its copies from every datastore.

    Deletion approved
          |
          v
    Check retention and hold requirements
          |
          v
    Delete or mark authoritative record
          |
          +-- Remove catalog projection
          +-- Remove search document
          +-- Invalidate cache
          +-- Remove graph relationships
          +-- Apply analytics-retention rule
          +-- Delete objects according to lifecycle policy
          |
          v
    Reconcile deletion outcome

    Backups and historical snapshots can follow separate approved retention processes.

    Historical Snapshots

    Some copies intentionally preserve the value valid at an earlier business event.

    Current course price:
    
    ₹2,000
    
    
    Purchased-course snapshot:
    
    ₹1,500
    
    
    Expected:
    
    The completed purchase retains
    the historical price of ₹1,500.

    A historical snapshot should not be synchronized with later master-data changes.

    Backup and Recovery

    Every authoritative datastore needs recovery planning. Derived stores need a restore or rebuild strategy.

    Store Recovery Direction
    Relational source of truth Backup, transaction-log recovery, and restore testing
    Document source of truth Database backup and restore with index validation
    Cache Repopulate from authoritative sources where appropriate
    Search index Restore or rebuild from authoritative content
    Graph projection Restore or regenerate relationships from approved source records
    Object storage Versioning, replication, backup, or recovery according to asset requirements
    Analytical projection Reload from retained source data and transformation pipelines

    Recovery rule: Restoring each database independently to a different point in time can create an inconsistent system. Define authoritative recovery points and how derived stores will be replayed, reconciled, or rebuilt.

    Schema Evolution

    The same business entity can have different schemas in different stores.

    Authoritative course schema:
    
    Version 12
    
    
    Catalog document schema:
    
    Version 5
    
    
    Search document schema:
    
    Version 7
    
    
    Analytics schema:
    
    Version 3

    Plan for:

    • Versioned events
    • Backward-compatible consumers
    • Optional fields during migration
    • Projection backfills
    • Index rebuilds
    • Dual-read or dual-write transition where justified
    • Retirement of obsolete fields and versions

    Dual Writes

    A dual write occurs when application code attempts to update two datastores directly for one business operation.

    Application
        |
        +-- Write relational database
        |
        +-- Write search index

    One write can succeed while the other fails.

    Uncoordinated dual write
    Update source
    Update projection
    
    No durable event
    No retry state
    No reconciliation
    Authoritative write plus projection workflow
    Commit authoritative change
    and outbox record
          |
          v
    Publish durable event
          |
          v
    Update projection idempotently
          |
          v
    Reconcile failures

    PHP Event-consumer Example

    <?php
    
    declare(strict_types=1);
    
    final class CoursePublishedHandler
    {
        public function __construct(
            private CourseRepository $courses,
            private SearchProjectionRepository $search,
            private CacheRepository $cache
        ) {
        }
    
        public function handle(
            int $tenantId,
            int $courseId,
            int $sourceVersion
        ): void {
            $course =
                $this->courses->findById(
                    $tenantId,
                    $courseId
                );
    
            if ($course === null) {
                return;
            }
    
            $indexedVersion =
                $this->search->getSourceVersion(
                    $tenantId,
                    $courseId
                );
    
            if ($indexedVersion !== null &&
                $indexedVersion >= $sourceVersion) {
    
                return;
            }
    
            $this->search->upsertCourse(
                tenantId:
                    $tenantId,
    
                courseId:
                    $courseId,
    
                title:
                    $course['title'],
    
                description:
                    $course['description'],
    
                sourceVersion:
                    $sourceVersion
            );
    
            $this->cache->remove(
                "tenant-{$tenantId}:course-{$courseId}"
            );
        }
    }

    Production processing also needs durable event handling, retry limits, transaction boundaries, dead-letter treatment, alerting, and reconciliation.

    Data-access Abstraction

    Application domains should use repositories or service interfaces rather than exposing datastore-specific details throughout the codebase.

    Course API
        |
        v
    Course application service
        |
        +-- Course repository
        +-- Search gateway
        +-- Cache gateway
        +-- Content gateway

    Abstraction should not hide important consistency, transaction, pagination, filtering, or failure semantics. A generic interface that pretends every datastore behaves identically can become misleading.

    Cost Model

    Polyglot persistence cost includes more than database licensing or storage.

    Consider:

    • Compute and storage
    • Indexes and replicas
    • Cross-service network traffic
    • Backup and retention
    • Monitoring and alerting
    • Data-transfer pipelines
    • Engineering skills
    • Operational support
    • Testing environments
    • Security reviews
    • Schema and client upgrades
    • Incident investigation

    A useful datastore should provide enough measurable workload value to justify these recurring costs.

    Complexity Budget

    Every additional datastore consumes part of the system's complexity budget.

    A conceptual decision score can be considered as:

    \[ NetValue = WorkloadBenefit - OperationalComplexity - ConsistencyRisk - MigrationCost \]

    This is not a universal mathematical formula. It provides a structured way to evaluate whether a specialized datastore creates more value than cost.

    Adoption Criteria

    Consider another datastore when:

    • The workload has a distinct data model or access pattern
    • The current platform cannot meet a verified requirement efficiently
    • The benefit is measurable
    • Data ownership is clear
    • Consistency requirements are understood
    • Synchronization and deletion are designed
    • The team can operate the datastore safely
    • Backup, recovery, and migration are supported
    • The datastore can be secured and monitored

    When to Avoid Polyglot Persistence

    Avoid adding another datastore when:

    • The existing database already meets the requirements
    • A suitable schema or index solves the problem
    • The workload is small
    • The system is still early and access patterns are uncertain
    • The team cannot support another production technology
    • Cross-store consistency would be unsafe
    • There is no recovery or reconciliation design
    • The technology is being selected only because it is popular
    • The expected performance benefit has not been measured

    Observability

    Monitor each datastore individually and the flows between them.

    • Request rate and latency by store
    • Error and timeout rate
    • Connection-pool health
    • Storage growth
    • Replication lag
    • Cache hit rate
    • Search-index freshness
    • Event-consumer lag
    • Projection failures
    • Dead-letter volume
    • Source-to-projection version lag
    • Reconciliation mismatches
    • Backup age and restore-test status
    • Deletion-propagation delay
    • Cross-store request count
    • Cost by datastore and workload

    Distributed Tracing

    Client request
          |
          v
    API gateway
          |
          v
    Course service
          |
          +-- Relational query
          +-- Cache lookup
          +-- Search request
          +-- Object-metadata lookup
          |
          v
    Response

    A shared correlation or trace identifier helps isolate which datastore or integration contributed to latency or failure.

    Alert Conditions

    Alert when:

    • A derived store exceeds its freshness objective
    • Event consumers stop processing
    • Cross-store retries increase
    • Authoritative and derived versions diverge
    • A deleted or restricted record remains visible
    • Backup or restore verification fails
    • One datastore becomes an availability bottleneck
    • Connection pools are exhausted
    • Dead-letter records increase
    • A schema migration leaves incompatible consumers
    • Cross-store network or request cost grows unexpectedly

    Troubleshooting Workflow

    1. Identify the user-visible incorrect or slow operation.
    2. Trace every datastore involved.
    3. Identify the authoritative source for the affected fact.
    4. Compare source and projection versions.
    5. Check whether the authoritative transaction committed.
    6. Check whether an event or change record was produced.
    7. Check delivery, retries, and dead-letter handling.
    8. Check consumer idempotency and stale-event rules.
    9. Check cache and search freshness.
    10. Check tenant and authorization filters.
    11. Check connection, timeout, and rate-limit behaviour.
    12. Repair or rebuild the derived store.
    13. Run reconciliation for potentially affected records.

    Common Polyglot-persistence Mistakes

    1

    Using a Different Database for Every Service

    Service independence does not require every service to introduce a new database technology.

    2

    Choosing Technology before Access Patterns

    The solution can become optimized for technology features instead of the actual workload.

    3

    Keeping Multiple Sources of Truth

    Independent edits in several stores create conflicting business facts.

    4

    Using Uncoordinated Dual Writes

    One datastore can update successfully while another remains stale.

    5

    Expecting One Transaction across Every Store

    Cross-database atomicity requires explicit distributed coordination and failure handling.

    6

    Ignoring Eventual Consistency

    Search, cache, graph, and analytical views can temporarily show different versions.

    7

    Using Stale Copies for Authorization

    Delayed publication or permission updates can expose restricted content.

    8

    Having No Reconciliation Process

    Missed events and partial failures can leave stores permanently inconsistent.

    9

    Having No Projection-rebuild Process

    Search and read models become difficult to repair after schema or mapping defects.

    10

    Backing Up Stores Independently without Recovery Coordination

    Restoring each store to a different logical point can create inconsistent system state.

    11

    Ignoring Deletion Propagation

    Deleted data can remain in caches, search indexes, graphs, analytics, objects, exports, or event records.

    12

    Underestimating Operational Cost

    Every datastore adds clients, credentials, upgrades, monitoring, recovery, skills, testing, and incident-response work.

    Recommended Test Cases

    Test Expected Evidence
    Authoritative write The business transaction commits in its owning datastore
    Projection update Derived stores receive the new source version
    Duplicate event The consumer remains idempotent
    Out-of-order event An older version does not replace newer data
    Consumer outage The projection catches up after recovery
    Partial failure Workflow state, retry, or compensation handles the outcome
    Search lag The delay remains within its defined objective
    Cache invalidation Changed content does not remain stale beyond its permitted duration
    Authorization change Restricted content stops appearing across every access path
    Deletion propagation Active and derived stores converge to the approved lifecycle state
    Reconciliation Injected differences are detected and repaired
    Projection rebuild A derived store is recreated from authoritative data
    Backup restoration Authoritative stores recover and derived stores are replayed or rebuilt
    Schema evolution Old and new producers and consumers remain compatible during migration
    Tenant isolation No datastore, cache, index, graph, or event path leaks cross-tenant data

    Polyglot-persistence Best Practices

    Recommended Practices

    • Use the smallest practical number of datastore technologies.
    • Select datastores from measured access patterns.
    • Confirm that the existing database cannot satisfy the requirement first.
    • Assign one authoritative owner to every business fact.
    • Keep strongly consistent invariants inside one transactional boundary where practical.
    • Treat caches, search indexes, and analytical stores as derived representations.
    • Use durable outbox or change-publication patterns.
    • Make projection consumers idempotent.
    • Use stable IDs and source versions.
    • Reject stale out-of-order changes.
    • Define acceptable consistency and freshness for every derived store.
    • Do not use unsafe stale projections for authorization.
    • Apply trusted tenant scope consistently across stores.
    • Maintain reconciliation and rebuild processes.
    • Coordinate backup and recovery around authoritative sources.
    • Propagate deletion and retention changes.
    • Version events, schemas, and projections.
    • Trace requests across datastore boundaries.
    • Measure operational and financial cost.
    • Document why every datastore exists and which workload it owns.

    Practice Exercise

    Design a polyglot-persistence architecture for your online learning platform.

    Requirements

    1. Store accounts, enrollments, quiz attempts, and payments transactionally.
    2. Store flexible course-catalog content.
    3. Store videos, PDFs, and images.
    4. Support session and rate-limit lookups.
    5. Support article and course full-text search.
    6. Store high-volume learner activity.
    7. Traverse course-prerequisite and topic relationships.
    8. Produce historical learning analytics.
    9. Assign one owner to every business fact.
    10. Define consistency and freshness for every projection.
    11. Design event-driven synchronization.
    12. Handle duplicate and out-of-order events.
    13. Protect tenant and authorization boundaries.
    14. Design deletion propagation.
    15. Design reconciliation and rebuild procedures.
    16. Design coordinated recovery.
    17. Estimate operational complexity and cost.

    Architecture-decision Template

    Workload Selected Store Ownership Consistency and Recovery
    Enrollment transactions Relational database Enrollment domain Transactional source of truth with tested restore
    Course catalog Relational or document database based on verified requirements Course-management domain Authoritative or derived status must be explicit
    Sessions Key-value store Identity or session domain Expiring state with documented failure behaviour
    Course media Object storage Content domain Version, checksum, backup, and lifecycle controls
    Full-text search Search index Derived from content domains Eventual consistency with complete rebuild
    Activity timeline Wide-column or event-oriented datastore Activity domain Partitioned retention and replay strategy
    Prerequisite graph Graph database or graph projection Learning-path domain Versioned relationships and rebuild procedure
    Analytics Warehouse or lakehouse Analytics domain using approved source data Documented lineage, freshness, and reload process

    Frequently Asked Questions

    1

    What is polyglot persistence?

    Polyglot persistence is the deliberate use of multiple data-storage technologies in one system, with each selected for a specific workload.

    2

    Does every microservice need a different database?

    No. Services can have separate ownership while using the same approved database technology.

    3

    Why not use one database for everything?

    One database can be the best choice when it satisfies all requirements. Separate stores are justified only when specialized access patterns create measurable value.

    4

    What is the most important design rule?

    Assign one authoritative owner to every business fact and treat other representations as controlled projections or historical snapshots.

    5

    How is data synchronized between stores?

    Common methods include APIs, durable events, change-data capture, batch pipelines, and complete rebuilds.

    6

    Can several databases share one transaction?

    Independent stores do not normally share one ordinary local transaction. Cross-store workflows require explicit coordination and failure handling.

    7

    What is an uncoordinated dual write?

    It occurs when application code updates two stores directly without a durable synchronization or recovery mechanism.

    8

    What is a derived store?

    A derived store contains data copied or calculated from an authoritative source, such as a cache, search index, graph projection, or analytical summary.

    9

    How are derived stores repaired?

    Use retries, version comparisons, reconciliation, targeted repair, and complete rebuilds from authoritative data.

    10

    How should deletion work?

    Deletion should follow an approved workflow that reaches active, projected, cached, indexed, analytical, and object copies while respecting retention rules.

    11

    What is the main disadvantage?

    The main disadvantage is increased operational and consistency complexity across several technologies and data copies.

    12

    When is polyglot persistence justified?

    It is justified when a distinct workload benefit is proven and the team can safely manage ownership, synchronization, security, recovery, observability, migration, and cost.

    Key Takeaway

    Polyglot persistence uses multiple purpose-specific datastores within one system. A relational database can own transactional records, a document database can hold flexible aggregates, a key-value store can serve sessions and caches, a wide-column store can handle partition-oriented event access, a graph database can support relationship traversal, a search engine can provide full-text retrieval, and object storage can hold large content. This specialization creates value only when it is driven by measured workloads. Every additional store introduces data ownership, synchronization, consistency, security, observability, migration, deletion, backup, recovery, skill, and cost responsibilities. Maintain one authoritative source for every current fact, make projection updates durable and idempotent, reject stale events, define freshness explicitly, reconcile derived copies, and preserve complete rebuild and recovery procedures.