polyglot persistence
Polyglot Persistence
Learn how one application can use relational, document, key-value, wide-column, graph, search, object-storage, and analytical systems for different workloads, while managing ownership, consistency, synchronization, transactions, security, observability, backup, recovery, migration, and operational complexity.
Introduction
Applications commonly manage several forms of data with substantially different access patterns.
An online learning platform can require:
- Transactional enrollment and payment records
- Flexible course-catalog documents
- Fast session and rate-limit lookups
- Large video and document storage
- Full-text article and course search
- High-volume activity-event ingestion
- Course-prerequisite and topic-relationship traversal
- Historical reporting and analytics
One database can sometimes support all these requirements. However, forcing every workload into one storage model can create inefficient queries, complicated schemas, scaling limitations, or unnecessary operational compromises.
Polyglot persistence is the deliberate use of multiple data-storage technologies within one application or system, with each datastore selected for a specific data model and access pattern.
Core idea: Use different datastores only when different workloads have materially different requirements. Choose each datastore for a defined purpose, assign clear data ownership, and accept the resulting consistency and operational responsibilities.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Relational databases | Relational storage remains a strong choice for transactional data, constraints, and complex relationships. |
| 2 | NoSQL data models | Key-value, document, wide-column, and graph databases solve different access-pattern problems. |
| 3 | Denormalization | Derived stores and read models commonly duplicate authoritative data. |
| 4 | Secondary indexes | An additional index can sometimes solve a query without introducing another datastore. |
| 5 | Transactions and consistency | Independent datastores do not normally share one ordinary local transaction. |
| 6 | Event-driven architecture | Events frequently synchronize data between authoritative and derived stores. |
| 7 | Observability and recovery | Every datastore adds monitoring, backup, restoration, and incident-response requirements. |
Why Is It Called Polyglot?
A polyglot person can use several languages and select the appropriate language for a particular situation.
Polyglot persistence applies the same idea to data storage. Instead of requiring one database technology to solve every storage problem, the architecture can select a suitable datastore for each workload.
Application workloads
+-- Transaction processing
| -> Relational database
|
+-- Flexible content
| -> Document database
|
+-- Sessions and caching
| -> Key-value store
|
+-- Activity timelines
| -> Wide-column database
|
+-- Relationship traversal
| -> Graph database
|
+-- Keyword retrieval
| -> Search index
|
+-- Videos and documents
-> Object storage
Single-database vs Polyglot Persistence
| Area | Single-database Approach | Polyglot Persistence |
|---|---|---|
| Technology count | One primary database technology | Several purpose-specific datastores |
| Operational complexity | Generally lower | Higher because every store must be operated |
| Transaction management | Can remain inside one database transaction | Cross-store operations require distributed workflow design |
| Workload specialization | One engine handles several workload types | Each engine can focus on a dominant access pattern |
| Skills and tooling | Smaller technology surface | More expertise, clients, monitoring, and automation required |
| Data duplication | Can remain lower | Derived and query-specific copies are common |
| Failure modes | Fewer integration boundaries | Partial failures and synchronization lag must be managed |
Workload-first Selection
Polyglot persistence should begin with workloads, not product names.
For each workload, identify:
- Data structure
- Primary reads
- Primary writes
- Expected scale
- Ordering and range requirements
- Transaction boundaries
- Consistency requirements
- Latency and availability objectives
- Retention and recovery requirements
- Team operational capability
Example Workloads
| Workload | Dominant Requirement | Candidate Datastore |
|---|---|---|
| Enrollment | Constraints and transactional updates | Relational database |
| Course catalog | Flexible nested course descriptions | Document database |
| User session | Fast retrieval by exact session key | Key-value store |
| Activity stream | High-volume time-ordered writes | Wide-column or event-oriented store |
| Course prerequisites | Multi-hop relationship traversal | Graph database |
| Article search | Keyword retrieval, ranking, and highlighting | Full-text search engine |
| Course video | Large durable binary objects | Object storage |
| Historical reporting | Large scans and aggregation | Analytical warehouse or lakehouse |
Selection rule: A different data model does not automatically require a different datastore. First verify whether an existing approved database can satisfy the workload through suitable schema, indexes, partitioning, or built-in capabilities.
Relational Database Role
A relational database is commonly used for data requiring constraints, relationships, transactional updates, and flexible structured queries.
CREATE TABLE enrollments
(
tenant_id BIGINT NOT NULL,
learner_id BIGINT NOT NULL,
course_id BIGINT NOT NULL,
enrollment_status VARCHAR(30) NOT NULL,
enrolled_at TIMESTAMP NOT NULL,
PRIMARY KEY
(
tenant_id,
learner_id,
course_id
),
CONSTRAINT fk_enrollment_course
FOREIGN KEY (course_id)
REFERENCES courses (course_id)
);
Suitable responsibilities can include:
- Accounts
- Enrollments
- Quiz attempts
- Certificates
- Orders and payments
- Authoritative workflow state
Document Database Role
A document database can store flexible, self-contained aggregates.
{
"tenantId": 17,
"courseId": 42,
"title": "System Design",
"difficulty": "beginner",
"instructor": {
"instructorId": 81,
"displayName": "Course Instructor"
},
"tags": [
"database",
"architecture",
"scalability"
],
"presentation": {
"thumbnailKey": "thumbnail-object-key",
"summary": "Learn practical system design."
}
}
Use bounded documents and define ownership for embedded and duplicated fields.
Key-Value Store Role
A key-value store is suitable when the application knows the exact key.
session:8f4a2
-> session data
rate-limit:tenant-17:user-1042
-> request counter
course-card:tenant-17:course-42
-> cached course card
Typical uses include:
- Sessions
- Caching
- Rate-limit state
- Idempotency records
- Temporary workflow state
- Short-lived tokens
Wide-Column Store Role
A wide-column database can serve high-volume access patterns based on known partition and ordering keys.
Partition key:
tenantId
+
courseId
+
activityDate
+
writeBucket
Clustering order:
activityTime
activityId
This supports retrieving course activity within a bounded date partition, while controlled write buckets can distribute unusually heavy writes.
Graph Database Role
A graph database can model entities and their explicit relationships.
[Course]
|
| REQUIRES
v
[Course]
|
| COVERS
v
[Topic]
|
| RELATED_TO
v
[Topic]
Graph-oriented access patterns can include:
- Course prerequisites
- Topic relationships
- Learning paths
- Knowledge graphs
- Dependency analysis
- Bounded recommendation traversal
Search Engine Role
A search engine maintains query-optimized documents and inverted indexes.
{
"documentId": "tenant-17:article-981",
"title": "Polyglot Persistence",
"description": "Learn how applications use multiple datastores.",
"body": "Searchable article content",
"course": "System Design",
"status": "published",
"sourceVersion": 12
}
The search index is generally a derived representation, not the authoritative source of article or publication state.
Object Storage Role
Object storage is suitable for large files and unstructured binary content.
Object storage:
tenants/17/courses/42/videos/lesson-7.mp4
tenants/17/courses/42/documents/chapter-3.pdf
tenants/17/courses/42/images/thumbnail.webp
Application metadata should connect each object key and version to its business asset, lifecycle state, ownership, checksum, and authorization rules.
Analytical Store Role
An analytical datastore can hold prepared historical data for reporting and large aggregations.
Operational events
|
v
Validated ingestion pipeline
|
v
Prepared analytical tables
|
v
Dashboards and reports
Analytical copies should have defined lineage, freshness, retention, quality, and access rules.
Example Architecture
Clients
|
v
Application APIs
|
+-- Enrollment Service
| -> Relational database
|
+-- Catalog Service
| -> Document database
|
+-- Session Service
| -> Key-value store
|
+-- Activity Service
| -> Wide-column store
|
+-- Learning Graph Service
| -> Graph database
|
+-- Search Service
| -> Search index
|
+-- Content Service
-> Object storage
Domain events
|
+-- Update search projections
+-- Update caches
+-- Update activity summaries
+-- Feed analytics platform
Authoritative Ownership
Every business fact should have one authoritative owner.
| Fact | Authoritative Owner | Derived Copies |
|---|---|---|
| Enrollment status | Enrollment service and relational database | Cache, search filter, analytics dataset |
| Course title | Course-management domain | Catalog document, search document, course-card cache |
| Video object | Content service and object metadata | Delivery-cache entries and search metadata |
| Article-search terms | Derived from authoritative article content | Search index only |
| Daily activity total | Derived from authoritative activity events | Dashboard and analytical summaries |
Ownership rule: Polyglot persistence must not create several independently editable sources of truth for the same business fact. Define ownership and update direction explicitly.
Avoid Shared-database Ownership
Services should not casually update one another's datastore.
Catalog Service
directly updates
Enrollment Service tables
Search Service
directly changes
Course Service documents
Service owns its data
|
v
Other services use:
- Published API
- Approved event
- Replicated read model
Direct database access can bypass validation, authorization, invariants, auditing, and lifecycle logic.
Data Synchronization
Data can move between datastores through synchronous calls, durable events, change-data capture, batch pipelines, or controlled rebuild processes.
| Method | Suitable Direction | Main Consideration |
|---|---|---|
| Synchronous API call | Immediate dependency between services | Availability and latency become coupled |
| Domain event | Update derived stores asynchronously | Temporary staleness and delivery handling |
| Change-data capture | Publish committed database changes | Schema interpretation and operational tooling |
| Batch pipeline | Historical analytics and periodic summaries | Longer freshness delay |
| Full rebuild | Recreate derived projections | Time, capacity, and cutover strategy |
Event-driven Synchronization
{
"eventType": "CoursePublished",
"eventVersion": 1,
"eventId": "generated-event-id",
"tenantId": 17,
"courseId": 42,
"sourceVersion": 12,
"occurredAt": "event-time"
}
Course transaction commits
|
v
Outbox event becomes available
|
v
Event is published
|
+-- Search consumer updates index
+-- Cache consumer invalidates course card
+-- Analytics consumer records publication event
+-- Recommendation consumer updates relationship data
Transactional Outbox
The transactional outbox records a business change and its event in the same local database transaction.
BEGIN;
UPDATE courses
SET
course_status = 'published',
source_version = source_version + 1
WHERE tenant_id = :tenant_id
AND course_id = :course_id;
INSERT INTO outbox_events
(
event_id,
event_type,
aggregate_id,
source_version,
payload,
created_at
)
VALUES
(
:event_id,
'CoursePublished',
:course_id,
:source_version,
:payload,
CURRENT_TIMESTAMP
);
COMMIT;
A publisher subsequently delivers unsent outbox records. Consumers should still be idempotent because delivery can occur more than once.
Idempotent Consumers
Event received
|
v
Check event ID and source version
|
+-- Already processed:
| return previous outcome
|
+-- Older than current:
| reject stale event
|
+-- New event:
update projection
record processed event
Stable event IDs, aggregate IDs, source versions, and projection versions help prevent duplicated or stale updates.
Cross-database Transactions
Independent datastores do not normally share one ordinary local ACID transaction.
Relational update succeeds
|
v
Graph update fails
|
v
Search update remains pending
Result:
The system is temporarily inconsistent.
Possible architectural responses include:
- Keep the business transaction in one authoritative datastore
- Update other stores asynchronously
- Use idempotent retries
- Use compensating business actions where appropriate
- Maintain workflow state
- Reconcile derived stores
Transaction rule: Do not split one strongly consistent business invariant across several databases unless the distributed coordination and failure behaviour are explicitly designed and justified.
Saga-style Workflow
A saga coordinates a multi-step business process through local transactions and compensating actions.
Create enrollment
|
v
Reserve payment
|
v
Provision course access
|
v
Send confirmation
If provisioning fails:
Release payment reservation
Mark enrollment failed
Record workflow outcome
Compensation is a business action, not a guaranteed reversal of every technical side effect.
Eventual Consistency
Derived datastores can temporarily display older information after the authoritative transaction commits.
Time A:
Course title updated
in authoritative store
Time B:
Search document updated
Between A and B:
Direct course read:
New title
Search result:
Old title
Define acceptable freshness separately for each projection.
| Data | Possible Consistency Direction |
|---|---|
| Search title | Short documented indexing delay can be acceptable |
| Course-view counter | Eventually consistent aggregate can be acceptable |
| Payment balance | Authoritative transactional consistency is required |
| Publication and authorization | Requires strongly bounded or authoritative enforcement |
Authorization across Datastores
Security must apply consistently across primary and derived stores.
Search request
|
v
Apply trusted tenant scope
|
v
Apply current publication and visibility rules
|
v
Return authorized search results
A stale search index or cache must not expose content after access is revoked.
Protect:
- Database records
- Document copies
- Cache entries
- Search snippets
- Graph relationships
- Analytical datasets
- Object metadata
- Events and dead-letter records
- Backups and exports
Tenant Isolation
Relational key:
tenant_id + enrollment_id
Document identity:
tenant-17:course-42
Cache key:
tenant-17:course-card:42
Search filter:
trustedTenantId = 17
Graph property:
tenantId = 17
Tenant scope should come from authenticated server-side context and remain present in every applicable data access path.
Reconciliation
Reconciliation detects differences between authoritative and derived stores.
Read authoritative course
|
v
Calculate expected projections
|
+-- Catalog document
+-- Search document
+-- Cache version
+-- Graph relationships
|
v
Compare IDs and versions
|
+-- Match:
| mark healthy
|
+-- Mismatch:
repair or rebuild
record diagnostic evidence
Reconciliation should use stable identifiers and source versions instead of comparing only display values.
Rebuildable Derived Stores
Search indexes, caches, recommendation projections, and analytical summaries should be rebuildable from authoritative data wherever practical.
Read authoritative records
|
v
Create replacement projection
|
v
Validate counts, versions, and sample queries
|
v
Switch reads to replacement
|
v
Retire old projection safely
Rebuild capability reduces the long-term risk of missed events, mapping defects, and schema changes.
Deletion Propagation
Deleting authoritative data does not automatically delete its copies from every datastore.
Deletion approved
|
v
Check retention and hold requirements
|
v
Delete or mark authoritative record
|
+-- Remove catalog projection
+-- Remove search document
+-- Invalidate cache
+-- Remove graph relationships
+-- Apply analytics-retention rule
+-- Delete objects according to lifecycle policy
|
v
Reconcile deletion outcome
Backups and historical snapshots can follow separate approved retention processes.
Historical Snapshots
Some copies intentionally preserve the value valid at an earlier business event.
Current course price:
₹2,000
Purchased-course snapshot:
₹1,500
Expected:
The completed purchase retains
the historical price of ₹1,500.
A historical snapshot should not be synchronized with later master-data changes.
Backup and Recovery
Every authoritative datastore needs recovery planning. Derived stores need a restore or rebuild strategy.
| Store | Recovery Direction |
|---|---|
| Relational source of truth | Backup, transaction-log recovery, and restore testing |
| Document source of truth | Database backup and restore with index validation |
| Cache | Repopulate from authoritative sources where appropriate |
| Search index | Restore or rebuild from authoritative content |
| Graph projection | Restore or regenerate relationships from approved source records |
| Object storage | Versioning, replication, backup, or recovery according to asset requirements |
| Analytical projection | Reload from retained source data and transformation pipelines |
Recovery rule: Restoring each database independently to a different point in time can create an inconsistent system. Define authoritative recovery points and how derived stores will be replayed, reconciled, or rebuilt.
Schema Evolution
The same business entity can have different schemas in different stores.
Authoritative course schema:
Version 12
Catalog document schema:
Version 5
Search document schema:
Version 7
Analytics schema:
Version 3
Plan for:
- Versioned events
- Backward-compatible consumers
- Optional fields during migration
- Projection backfills
- Index rebuilds
- Dual-read or dual-write transition where justified
- Retirement of obsolete fields and versions
Dual Writes
A dual write occurs when application code attempts to update two datastores directly for one business operation.
Application
|
+-- Write relational database
|
+-- Write search index
One write can succeed while the other fails.
Update source
Update projection
No durable event
No retry state
No reconciliation
Commit authoritative change
and outbox record
|
v
Publish durable event
|
v
Update projection idempotently
|
v
Reconcile failures
PHP Event-consumer Example
<?php
declare(strict_types=1);
final class CoursePublishedHandler
{
public function __construct(
private CourseRepository $courses,
private SearchProjectionRepository $search,
private CacheRepository $cache
) {
}
public function handle(
int $tenantId,
int $courseId,
int $sourceVersion
): void {
$course =
$this->courses->findById(
$tenantId,
$courseId
);
if ($course === null) {
return;
}
$indexedVersion =
$this->search->getSourceVersion(
$tenantId,
$courseId
);
if ($indexedVersion !== null &&
$indexedVersion >= $sourceVersion) {
return;
}
$this->search->upsertCourse(
tenantId:
$tenantId,
courseId:
$courseId,
title:
$course['title'],
description:
$course['description'],
sourceVersion:
$sourceVersion
);
$this->cache->remove(
"tenant-{$tenantId}:course-{$courseId}"
);
}
}
Production processing also needs durable event handling, retry limits, transaction boundaries, dead-letter treatment, alerting, and reconciliation.
Data-access Abstraction
Application domains should use repositories or service interfaces rather than exposing datastore-specific details throughout the codebase.
Course API
|
v
Course application service
|
+-- Course repository
+-- Search gateway
+-- Cache gateway
+-- Content gateway
Abstraction should not hide important consistency, transaction, pagination, filtering, or failure semantics. A generic interface that pretends every datastore behaves identically can become misleading.
Cost Model
Polyglot persistence cost includes more than database licensing or storage.
Consider:
- Compute and storage
- Indexes and replicas
- Cross-service network traffic
- Backup and retention
- Monitoring and alerting
- Data-transfer pipelines
- Engineering skills
- Operational support
- Testing environments
- Security reviews
- Schema and client upgrades
- Incident investigation
A useful datastore should provide enough measurable workload value to justify these recurring costs.
Complexity Budget
Every additional datastore consumes part of the system's complexity budget.
A conceptual decision score can be considered as:
\[ NetValue = WorkloadBenefit - OperationalComplexity - ConsistencyRisk - MigrationCost \]
This is not a universal mathematical formula. It provides a structured way to evaluate whether a specialized datastore creates more value than cost.
Adoption Criteria
Consider another datastore when:
- The workload has a distinct data model or access pattern
- The current platform cannot meet a verified requirement efficiently
- The benefit is measurable
- Data ownership is clear
- Consistency requirements are understood
- Synchronization and deletion are designed
- The team can operate the datastore safely
- Backup, recovery, and migration are supported
- The datastore can be secured and monitored
When to Avoid Polyglot Persistence
Avoid adding another datastore when:
- The existing database already meets the requirements
- A suitable schema or index solves the problem
- The workload is small
- The system is still early and access patterns are uncertain
- The team cannot support another production technology
- Cross-store consistency would be unsafe
- There is no recovery or reconciliation design
- The technology is being selected only because it is popular
- The expected performance benefit has not been measured
Observability
Monitor each datastore individually and the flows between them.
- Request rate and latency by store
- Error and timeout rate
- Connection-pool health
- Storage growth
- Replication lag
- Cache hit rate
- Search-index freshness
- Event-consumer lag
- Projection failures
- Dead-letter volume
- Source-to-projection version lag
- Reconciliation mismatches
- Backup age and restore-test status
- Deletion-propagation delay
- Cross-store request count
- Cost by datastore and workload
Distributed Tracing
Client request
|
v
API gateway
|
v
Course service
|
+-- Relational query
+-- Cache lookup
+-- Search request
+-- Object-metadata lookup
|
v
Response
A shared correlation or trace identifier helps isolate which datastore or integration contributed to latency or failure.
Alert Conditions
Alert when:
- A derived store exceeds its freshness objective
- Event consumers stop processing
- Cross-store retries increase
- Authoritative and derived versions diverge
- A deleted or restricted record remains visible
- Backup or restore verification fails
- One datastore becomes an availability bottleneck
- Connection pools are exhausted
- Dead-letter records increase
- A schema migration leaves incompatible consumers
- Cross-store network or request cost grows unexpectedly
Troubleshooting Workflow
- Identify the user-visible incorrect or slow operation.
- Trace every datastore involved.
- Identify the authoritative source for the affected fact.
- Compare source and projection versions.
- Check whether the authoritative transaction committed.
- Check whether an event or change record was produced.
- Check delivery, retries, and dead-letter handling.
- Check consumer idempotency and stale-event rules.
- Check cache and search freshness.
- Check tenant and authorization filters.
- Check connection, timeout, and rate-limit behaviour.
- Repair or rebuild the derived store.
- Run reconciliation for potentially affected records.
Common Polyglot-persistence Mistakes
Using a Different Database for Every Service
Service independence does not require every service to introduce a new database technology.
Choosing Technology before Access Patterns
The solution can become optimized for technology features instead of the actual workload.
Keeping Multiple Sources of Truth
Independent edits in several stores create conflicting business facts.
Using Uncoordinated Dual Writes
One datastore can update successfully while another remains stale.
Expecting One Transaction across Every Store
Cross-database atomicity requires explicit distributed coordination and failure handling.
Ignoring Eventual Consistency
Search, cache, graph, and analytical views can temporarily show different versions.
Using Stale Copies for Authorization
Delayed publication or permission updates can expose restricted content.
Having No Reconciliation Process
Missed events and partial failures can leave stores permanently inconsistent.
Having No Projection-rebuild Process
Search and read models become difficult to repair after schema or mapping defects.
Backing Up Stores Independently without Recovery Coordination
Restoring each store to a different logical point can create inconsistent system state.
Ignoring Deletion Propagation
Deleted data can remain in caches, search indexes, graphs, analytics, objects, exports, or event records.
Underestimating Operational Cost
Every datastore adds clients, credentials, upgrades, monitoring, recovery, skills, testing, and incident-response work.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Authoritative write | The business transaction commits in its owning datastore |
| Projection update | Derived stores receive the new source version |
| Duplicate event | The consumer remains idempotent |
| Out-of-order event | An older version does not replace newer data |
| Consumer outage | The projection catches up after recovery |
| Partial failure | Workflow state, retry, or compensation handles the outcome |
| Search lag | The delay remains within its defined objective |
| Cache invalidation | Changed content does not remain stale beyond its permitted duration |
| Authorization change | Restricted content stops appearing across every access path |
| Deletion propagation | Active and derived stores converge to the approved lifecycle state |
| Reconciliation | Injected differences are detected and repaired |
| Projection rebuild | A derived store is recreated from authoritative data |
| Backup restoration | Authoritative stores recover and derived stores are replayed or rebuilt |
| Schema evolution | Old and new producers and consumers remain compatible during migration |
| Tenant isolation | No datastore, cache, index, graph, or event path leaks cross-tenant data |
Polyglot-persistence Best Practices
Recommended Practices
- Use the smallest practical number of datastore technologies.
- Select datastores from measured access patterns.
- Confirm that the existing database cannot satisfy the requirement first.
- Assign one authoritative owner to every business fact.
- Keep strongly consistent invariants inside one transactional boundary where practical.
- Treat caches, search indexes, and analytical stores as derived representations.
- Use durable outbox or change-publication patterns.
- Make projection consumers idempotent.
- Use stable IDs and source versions.
- Reject stale out-of-order changes.
- Define acceptable consistency and freshness for every derived store.
- Do not use unsafe stale projections for authorization.
- Apply trusted tenant scope consistently across stores.
- Maintain reconciliation and rebuild processes.
- Coordinate backup and recovery around authoritative sources.
- Propagate deletion and retention changes.
- Version events, schemas, and projections.
- Trace requests across datastore boundaries.
- Measure operational and financial cost.
- Document why every datastore exists and which workload it owns.
Practice Exercise
Design a polyglot-persistence architecture for your online learning platform.
Requirements
- Store accounts, enrollments, quiz attempts, and payments transactionally.
- Store flexible course-catalog content.
- Store videos, PDFs, and images.
- Support session and rate-limit lookups.
- Support article and course full-text search.
- Store high-volume learner activity.
- Traverse course-prerequisite and topic relationships.
- Produce historical learning analytics.
- Assign one owner to every business fact.
- Define consistency and freshness for every projection.
- Design event-driven synchronization.
- Handle duplicate and out-of-order events.
- Protect tenant and authorization boundaries.
- Design deletion propagation.
- Design reconciliation and rebuild procedures.
- Design coordinated recovery.
- Estimate operational complexity and cost.
Architecture-decision Template
| Workload | Selected Store | Ownership | Consistency and Recovery |
|---|---|---|---|
| Enrollment transactions | Relational database | Enrollment domain | Transactional source of truth with tested restore |
| Course catalog | Relational or document database based on verified requirements | Course-management domain | Authoritative or derived status must be explicit |
| Sessions | Key-value store | Identity or session domain | Expiring state with documented failure behaviour |
| Course media | Object storage | Content domain | Version, checksum, backup, and lifecycle controls |
| Full-text search | Search index | Derived from content domains | Eventual consistency with complete rebuild |
| Activity timeline | Wide-column or event-oriented datastore | Activity domain | Partitioned retention and replay strategy |
| Prerequisite graph | Graph database or graph projection | Learning-path domain | Versioned relationships and rebuild procedure |
| Analytics | Warehouse or lakehouse | Analytics domain using approved source data | Documented lineage, freshness, and reload process |
Frequently Asked Questions
What is polyglot persistence?
Polyglot persistence is the deliberate use of multiple data-storage technologies in one system, with each selected for a specific workload.
Does every microservice need a different database?
No. Services can have separate ownership while using the same approved database technology.
Why not use one database for everything?
One database can be the best choice when it satisfies all requirements. Separate stores are justified only when specialized access patterns create measurable value.
What is the most important design rule?
Assign one authoritative owner to every business fact and treat other representations as controlled projections or historical snapshots.
How is data synchronized between stores?
Common methods include APIs, durable events, change-data capture, batch pipelines, and complete rebuilds.
Can several databases share one transaction?
Independent stores do not normally share one ordinary local transaction. Cross-store workflows require explicit coordination and failure handling.
What is an uncoordinated dual write?
It occurs when application code updates two stores directly without a durable synchronization or recovery mechanism.
What is a derived store?
A derived store contains data copied or calculated from an authoritative source, such as a cache, search index, graph projection, or analytical summary.
How are derived stores repaired?
Use retries, version comparisons, reconciliation, targeted repair, and complete rebuilds from authoritative data.
How should deletion work?
Deletion should follow an approved workflow that reaches active, projected, cached, indexed, analytical, and object copies while respecting retention rules.
What is the main disadvantage?
The main disadvantage is increased operational and consistency complexity across several technologies and data copies.
When is polyglot persistence justified?
It is justified when a distinct workload benefit is proven and the team can safely manage ownership, synchronization, security, recovery, observability, migration, and cost.
Key Takeaway
Polyglot persistence uses multiple purpose-specific datastores within one system. A relational database can own transactional records, a document database can hold flexible aggregates, a key-value store can serve sessions and caches, a wide-column store can handle partition-oriented event access, a graph database can support relationship traversal, a search engine can provide full-text retrieval, and object storage can hold large content. This specialization creates value only when it is driven by measured workloads. Every additional store introduces data ownership, synchronization, consistency, security, observability, migration, deletion, backup, recovery, skill, and cost responsibilities. Maintain one authoritative source for every current fact, make projection updates durable and idempotent, reject stale events, define freshness explicitly, reconcile derived copies, and preserve complete rebuild and recovery procedures.