Key-value, document, wide-column and graph models
Key-Value, Document, Wide-Column, and Graph Models
Learn how the four major NoSQL data models organize information, which access patterns each model supports well, how keys and partitions affect performance, and how to select a database from query, consistency, transaction, scale, and operational requirements.
Introduction
Relational databases organize data into tables with defined columns, relationships, constraints, and SQL queries. They are a strong choice for many transactional systems, but they are not the only available data model.
NoSQL databases provide alternative ways to organize and retrieve data. Different NoSQL systems are optimized for different access patterns, distribution strategies, consistency requirements, and data structures.
The four commonly discussed NoSQL models are:
- Key-value: Retrieve a value using a unique key.
- Document: Store and query self-contained documents.
- Wide-column: Organize rows by partition and clustering keys across flexible columns.
- Graph: Represent entities and relationships as nodes and edges.
Core idea: Select a NoSQL model from the application's access patterns. First identify how data will be written, retrieved, updated, traversed, filtered, and scaled. Then design keys and records that make those operations direct and predictable.
This lesson begins the NoSQL & Access-Pattern Design module and introduces the four major NoSQL models before deeper lessons on partition keys, denormalization, consistency, secondary indexes, hot partitions, and datastore selection.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Entities and relationships | NoSQL designs still represent business entities and their relationships. |
| 2 | Keys and constraints | Record identity and partition-key design determine how data is located. |
| 3 | Normalization and denormalization | NoSQL models commonly duplicate data to support predefined queries. |
| 4 | Transactions and consistency | Atomic and consistency guarantees vary between databases and operation scopes. |
| 5 | Query patterns | NoSQL schemas are frequently designed around known reads and writes. |
| 6 | Basic distributed systems | Partitioning, replication, failures, and network communication affect distributed databases. |
What Does NoSQL Mean?
NoSQL is a broad category of database technologies that use data models other than, or in addition to, the traditional relational table model.
The term does not identify one database architecture or one consistency model. Two NoSQL databases can provide very different query languages, transaction scopes, indexing features, and scaling behaviour.
NoSQL database families
+-- Key-value
|
+-- Document
|
+-- Wide-column
|
+-- Graph
Important: NoSQL does not automatically mean schemaless, eventually consistent, faster, or infinitely scalable. The exact guarantees and limitations must be verified for the selected database.
Access-pattern-first Design
Relational modeling often starts with entities and their normalized relationships. NoSQL design frequently starts with the operations the application must perform.
For every important operation, ask:
- What information does the request already know?
- Which key can locate the required data?
- Does the request retrieve one item or many items?
- What ordering is required?
- What filters are required?
- How many records can the request examine?
- Which fields change together?
- Must several records update atomically?
- How frequently is the operation executed?
- How evenly will traffic be distributed?
Key-Value Model
A key-value database stores a value under a unique key. The application retrieves or modifies the value by supplying that key.
Key Value
session:8f4a2 { ...session data... }
course-progress:1042:42 { ...progress data... }
rate-limit:user:1042 { ...counter state... }
feature:search-v2 { ...feature settings... }
The database can treat the value as an opaque byte sequence or expose additional operations for supported data structures.
Basic Operation
Application knows key
|
v
GET value by key
|
v
Database locates entry
|
v
Return value
Common Operations
GET key
PUT key value
DELETE key
COMPARE-AND-SET key expected-version new-value
INCREMENT key
EXPIRE key after-duration
Supported operations depend on the selected database.
Course-progress Example
{
"key": "course-progress:learner-1042:course-42",
"value": {
"completionPercentage": 72.5,
"lastLessonId": 107,
"lastActivityAt": "activity-time",
"version": 8
}
}
Suitable Access Patterns
- Retrieve one value by exact key
- Store session data
- Store temporary tokens
- Maintain rate-limit counters
- Store feature configuration
- Cache calculated results
- Maintain shopping-cart or draft state
Less-direct Patterns
- Ad hoc filtering across arbitrary value fields
- Joins between several entity types
- Complex aggregation without additional structures
- Relationship traversal
- Queries when the required key is unknown
Key-value rule: A key-value store is strongest when the application already knows the exact key required to locate the value.
Document Model
A document database stores records as self-contained documents containing scalar values, nested objects, and arrays.
{
"courseId": 42,
"tenantId": 17,
"title": "System Design",
"difficulty": "beginner",
"instructor": {
"instructorId": 81,
"displayName": "Course Instructor"
},
"tags": [
"architecture",
"database",
"scalability"
],
"statistics": {
"chapterCount": 20,
"estimatedHours": 196
},
"status": "published"
}
Documents commonly use JSON-like structures, although the database can store them in another internal binary representation.
Aggregate-oriented Design
A document often represents an aggregate that is commonly read and updated as one unit.
Course document
|
+-- Core course fields
+-- Instructor summary
+-- Tags
+-- Publication state
+-- Small embedded settings
Embedding related data can avoid additional reads, but it also duplicates information and requires an update strategy.
Document Queries
Find course by course ID
Find published courses by tenant
Find beginner courses containing a tag
Find courses by embedded instructor ID
Update one nested setting
Suitable Access Patterns
- Retrieve one aggregate by document ID
- Store flexible product, course, profile, or catalog records
- Query indexed document properties
- Store nested data normally read together
- Support records whose optional fields vary by type
- Evolve document structure gradually
Design Considerations
- Maximum document size
- Unbounded embedded arrays
- Duplication of commonly changing information
- Cross-document transactions
- Secondary-index cost
- Schema-version migration
- Concurrent updates to one large document
Embed or Reference?
Document designs must decide whether related data belongs inside the same document or in a separate document.
Embedded Lessons
{
"courseId": 42,
"title": "System Design",
"lessons": [
{
"lessonId": 101,
"title": "Requirements Discovery",
"displayOrder": 1
},
{
"lessonId": 102,
"title": "Functional Requirements",
"displayOrder": 2
}
]
}
Embedding can be useful when the child collection is bounded, normally read with the parent, and updated within a manageable aggregate.
Referenced Lessons
{
"courseId": 42,
"title": "System Design",
"lessonIds": [
101,
102
]
}
Referencing can be preferable when child records are large, unbounded, accessed independently, or updated at a different frequency.
| Consideration | Embed | Reference |
|---|---|---|
| Read together | Usually suitable | Can require additional reads |
| Atomic update | Can remain within one document | Can require multi-document coordination |
| Collection growth | Suitable when bounded | Safer for large or unbounded collections |
| Independent access | Less convenient | Each referenced record can be retrieved independently |
| Shared child | Creates duplication | One child can be referenced by several parents |
Wide-Column Model
A wide-column database organizes data around partition keys and ordered rows or columns within each partition.
The logical model varies between products, but it commonly emphasizes:
- Partition key
- Clustering or sort key
- Efficient retrieval within one partition
- Sparse or flexible columns
- Large distributed datasets
Simplified Representation
Partition key:
course-42
Rows ordered by event time:
2026-09-23T08:00 | learner-1042 | lesson-started
2026-09-23T08:10 | learner-1042 | lesson-completed
2026-09-23T08:20 | learner-2048 | quiz-started
2026-09-23T08:28 | learner-2048 | quiz-completed
Composite Primary Key
Partition key:
course_id
Clustering keys:
event_date
event_time
event_id
Records sharing the partition key are colocated logically, while clustering keys determine their order inside the partition.
Conceptual Table
CREATE TABLE course_activity_by_course_day
(
tenant_id BIGINT,
course_id BIGINT,
activity_date DATE,
activity_time TIMESTAMP,
activity_id VARCHAR(100),
learner_id BIGINT,
activity_type VARCHAR(50),
lesson_id BIGINT,
activity_payload TEXT,
PRIMARY KEY
(
(tenant_id, course_id, activity_date),
activity_time,
activity_id
)
);
This educational syntax illustrates a composite partition key followed by clustering columns. Exact syntax and guarantees depend on the selected wide-column database.
Suitable Access Patterns
- Time-ordered events for one entity or partition
- High-volume write workloads
- Telemetry and activity histories
- Queries scoped to a known partition key
- Large sparse datasets
- Precomputed query-specific tables
Less-direct Patterns
- Ad hoc queries without the partition key
- Cross-partition joins
- Global ordering across many partitions
- Cross-partition aggregation without a separate design
- Frequent access patterns not represented by an existing table
Wide-column rule: Model each important query around a partition key and clustering order. A table can be designed specifically for one access pattern.
Hot Partitions
A hot partition occurs when a disproportionate amount of traffic targets one partition key.
Good distribution:
course-1 -> moderate traffic
course-2 -> moderate traffic
course-3 -> moderate traffic
course-4 -> moderate traffic
Hot distribution:
global-activity -> nearly all traffic
A low-cardinality or monotonically concentrated partition key can route too much work to one partition.
Possible design responses include:
- Add tenant or entity identity to the partition key
- Add time buckets
- Add a controlled write shard
- Separate high-traffic entities
- Pre-aggregate data
- Redesign the access pattern
Every sharding technique increases read, aggregation, or operational complexity and should be validated with representative traffic.
Graph Model
A graph database represents data as nodes and relationships, commonly called edges.
[Learner]
|
| ENROLLED_IN
v
[Course]
|
| CONTAINS
v
[Lesson]
|
| COVERS
v
[Topic]
Nodes represent entities. Edges represent relationships between entities. Both can have properties.
Node Example
{
"label": "Course",
"properties": {
"courseId": 42,
"title": "System Design",
"status": "published"
}
}
Relationship Example
{
"type": "ENROLLED_IN",
"from": "learner-1042",
"to": "course-42",
"properties": {
"enrolledAt": "enrollment-time",
"status": "active"
}
}
Graph Traversal
Graph queries follow relationships from one node to connected nodes.
Starting node:
Learner 1042
Traversal:
Learner
-> enrolled courses
-> lessons in those courses
-> topics covered by those lessons
Result:
Topics accessible through
the learner's current enrollments
Conceptual Cypher Query
MATCH
(learner:Learner {learnerId: $learnerId})
-[:ENROLLED_IN]->
(course:Course)
-[:CONTAINS]->
(lesson:Lesson)
-[:COVERS]->
(topic:Topic)
RETURN DISTINCT
topic.topicId,
topic.name;
Query language, syntax, indexing, and transaction behaviour depend on the selected graph database.
Suitable Access Patterns
- Multi-hop relationship traversal
- Social or collaboration networks
- Knowledge graphs
- Dependency analysis
- Fraud-pattern exploration
- Recommendation relationships
- Identity and entitlement relationships
- Network and topology analysis
Design Considerations
- High-degree nodes
- Unbounded traversals
- Traversal depth
- Cycle handling
- Relationship direction
- Graph partitioning
- Authorization during traversal
- Bulk analytical workloads
Four-model Comparison
| Area | Key-Value | Document | Wide-Column | Graph |
|---|---|---|---|---|
| Primary structure | Key and value | Nested document | Partitioned, clustered rows | Nodes and relationships |
| Primary access | Exact key | Document ID or indexed field | Partition key and ordered range | Relationship traversal |
| Typical strength | Simple fast lookup | Flexible aggregates | Distributed high-volume query-specific access | Connected-data exploration |
| Joins | Usually application-managed | Embedding or references | Usually modeled out of the read path | Relationships are central |
| Common risk | Unknown-key queries | Unbounded documents | Hot or oversized partitions | Unbounded traversal |
| Example use | Session | Course catalog record | Course activity by day | Learning knowledge graph |
NoSQL vs Relational Database
| Requirement | Possible Direction |
|---|---|
| Strong relational integrity and multi-table transactions | Relational database |
| Exact retrieval by a known key | Key-value database |
| Flexible aggregate containing nested fields | Document database |
| Large partition-oriented time-series access | Wide-column database |
| Variable-depth relationship traversal | Graph database |
These are architectural directions, not automatic decisions. Modern databases can overlap in capability, and the exact product must be evaluated against the required workload and guarantees.
Denormalization
NoSQL access-pattern design commonly places data together according to how it will be read.
Authoritative instructor:
instructor_id = 81
display_name = Course Instructor
Copied course summary:
course_id = 42
instructor_id = 81
instructor_display_name = Course Instructor
Denormalization can reduce read operations but creates duplicated data that must be updated, tolerated temporarily, or rebuilt.
Denormalization Questions
- Which copy is authoritative?
- How frequently does the duplicated field change?
- How many copies must be updated?
- Is temporary inconsistency acceptable?
- Can stale copies be detected?
- Can the projection be rebuilt?
Primary-key Design
NoSQL keys frequently do more than identify a record. They can determine partition placement, ordering, locality, and supported queries.
Key design:
TENANT#17#COURSE#42
Possible sort keys:
COURSE
CHAPTER#001
CHAPTER#002
LESSON#001#001
LESSON#001#002
A key should be:
- Stable
- Unique within the required scope
- Compatible with the main access patterns
- Able to distribute traffic appropriately
- Safe to store and log under the applicable policy
- Independent of mutable display names where practical
Data Locality
Data locality places records commonly accessed together in the same document, partition, aggregate, or nearby key range.
Access pattern:
Load course summary and its first
20 lesson summaries.
Locality-oriented design:
Course summary and bounded lesson
summaries stored together
or under one partition.
Result:
Fewer distributed reads.
Excessive colocation can create oversized documents, hot partitions, contention, or expensive updates.
Secondary Access Patterns
One primary key rarely supports every query.
Additional patterns can be supported through:
- Secondary indexes
- Additional query-specific tables
- Materialized projections
- Search indexes
- Reverse-lookup records
- Event-driven derived views
Primary access:
Course by course ID
Secondary access:
Published courses by category
Possible projection:
Category
-> publication date
-> course ID
-> course summary
Consistency
NoSQL systems can expose different consistency choices for reads and writes.
Write accepted
|
v
Primary copy updated
|
v
Replicas or projections update
|
v
Readers observe new value
according to consistency contract
Review:
- Read-after-write behaviour
- Replica consistency
- Conditional updates
- Conflict handling
- Transaction scope
- Secondary-index freshness
- Cross-region replication lag
Do not infer the consistency behaviour of one NoSQL database from another.
Conditional Updates
A conditional update can prevent one writer from silently overwriting a newer value.
{
"courseId": 42,
"title": "System Design",
"version": 8
}
Update request:
Set title to new value
Condition:
Current version must equal 8
If condition succeeds:
Store new value
Set version to 9
If condition fails:
Return concurrency conflict
Exact conditional-write syntax depends on the selected database.
Transaction Boundaries
The safest aggregate often matches the data that must change atomically.
One document or item:
Can often update atomically
under the database contract.
Several documents or partitions:
Can require a database-supported
transaction, workflow, compensation,
or eventual-consistency design.
Verify transaction limits, item counts, partition scope, latency, and failure behaviour before depending on multi-record transactions.
Capacity Estimation
Estimate both data volume and request distribution.
A simplified storage estimate is:
\[ Storage = ItemCount \times AverageItemSize \times ReplicationAndIndexFactor \]
A simplified hot-key estimate is:
\[ RequestsPerHotKey = TotalRequestRate \times HotKeyTrafficShare \]
Average traffic alone can hide a partition receiving a disproportionate share of requests.
Choosing a Model by Access Pattern
| Access Pattern | Candidate Model |
|---|---|
| Retrieve session by opaque session ID | Key-value |
| Retrieve a flexible course catalog record | Document |
| Retrieve learner activity for a course and date range | Wide-column |
| Traverse prerequisites through several course relationships | Graph |
| Execute transactional enrollment and payment updates | Relational can remain the stronger starting point |
| Search article text by relevance | Full-text search index |
Polyglot Persistence
One application can use several data technologies when different components have materially different access patterns.
Online learning platform
+-- Relational database
| Enrollments, payments, quiz attempts
|
+-- Key-value store
| Sessions, rate limits, temporary state
|
+-- Document database
| Flexible course catalog documents
|
+-- Wide-column store
| High-volume learner activity timelines
|
+-- Graph database
| Topic, prerequisite, and recommendation relationships
|
+-- Search index
Article and course full-text search
Every additional datastore increases operational, security, backup, monitoring, skill, migration, and consistency complexity.
Technology rule: Use another datastore only when its workload benefit justifies the additional operational complexity. Do not adopt several databases merely because each supports a different model.
Security Considerations
NoSQL security should address:
- Authentication
- Least-privilege authorization
- Tenant isolation
- Encryption in transit
- Encryption at rest
- Credential rotation
- Backup protection
- Audit logging
- Query and traversal limits
- Protection of indexes and replicas
Keys, document properties, graph relationships, and partition names can expose sensitive business information even when values are encrypted.
Tenant-aware Keys
Tenant-scoped key:
TENANT#17#COURSE#42
Tenant-scoped graph property:
tenantId = 17
Tenant-scoped document:
{
"tenantId": 17,
"courseId": 42
}
Tenant scope must come from trusted authentication and authorization context. A client-provided tenant ID alone does not prove access.
Observability
Useful NoSQL metrics include:
- Read and write request rate
- Latency by percentile
- Error and timeout rate
- Throttled request count
- Partition-size distribution
- Hot-key or hot-partition activity
- Item or document size
- Secondary-index usage
- Replication lag
- Conflict count
- Conditional-write failure count
- Graph traversal depth and result size
- Backup and restore status
- Storage and request cost
Common NoSQL Modeling Mistakes
Choosing a Database before Listing Access Patterns
The resulting model can make important queries expensive or impossible without redesign.
Copying a Normalized Relational Schema Directly
Numerous separate records can require application-managed joins that the selected NoSQL model was not designed to perform efficiently.
Using One Giant Document
The document can exceed size limits, create write contention, and make small updates expensive.
Embedding Unbounded Collections
Comments, events, followers, or activity histories can grow without a safe document limit.
Using a Low-cardinality Partition Key
A small set of values can concentrate traffic and storage into hot partitions.
Using Random Keys When Ordered Retrieval Is Required
The design can distribute data well but lose the locality needed by the required range query.
Expecting Ad Hoc Queries Automatically
Important queries can require secondary indexes, projections, or separate query-specific tables.
Duplicating Data without an Update Strategy
Denormalized copies can remain stale or contradictory.
Assuming Every NoSQL Database Is Eventually Consistent
Consistency options and transaction guarantees differ between products and operation types.
Using Graph Traversal without Boundaries
Unrestricted depth and high-degree nodes can produce expensive searches and expose unintended relationships.
Ignoring Secondary-index Cost
Every additional index can increase write, storage, replication, and maintenance work.
Using Several Datastores without Operational Readiness
Each database adds monitoring, backup, recovery, security, migration, and specialist-support requirements.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Exact key lookup | The key-value item is retrieved directly |
| Missing key | The API returns the documented not-found result |
| Document retrieval | The complete required aggregate is returned |
| Document growth | The model remains within approved item-size limits |
| Conditional update | A stale version does not overwrite a newer value |
| Partition-range query | Rows are retrieved within one known partition and expected order |
| Hot-partition test | Traffic distribution and throttling behaviour are measured |
| Graph traversal | The expected bounded relationship path is returned |
| Traversal limit | Excessive depth or result size is rejected or bounded |
| Denormalized update | Required copies converge according to the consistency objective |
| Tenant isolation | Keys, queries, and traversals cannot expose another tenant's data |
| Backup restore | Data and required indexes are restored and verified |
NoSQL Model-selection Best Practices
Recommended Practices
- List and prioritize access patterns before selecting a database.
- Choose the simplest datastore that satisfies the requirements.
- Use key-value storage when the application knows the exact key.
- Use document storage for bounded self-contained aggregates.
- Use wide-column storage for partition-oriented ordered access at scale.
- Use graph storage for relationship-centric traversal.
- Keep strongly relational transactional workloads in a relational database when appropriate.
- Design partition keys for both locality and traffic distribution.
- Avoid unbounded documents, partitions, arrays, and traversals.
- Use stable identities rather than mutable display values.
- Document authoritative and denormalized copies.
- Define consistency and transaction requirements explicitly.
- Use conditional updates to protect concurrent writes.
- Create query-specific projections deliberately.
- Measure secondary-index write and storage cost.
- Apply trusted tenant scope to keys, queries, and traversals.
- Load-test realistic key and traffic distributions.
- Test backup and restoration.
- Monitor hot partitions, throttling, replication lag, and conflicts.
- Adopt polyglot persistence only when the benefit justifies its complexity.
Practice Exercise
Select data models for the major access patterns of your online learning platform.
Requirements
- Retrieve a session by session token.
- Retrieve one flexible course catalog record.
- List a learner's recent activity in time order.
- List activity for one course and one day.
- Traverse course prerequisites through several levels.
- Find topics related to a lesson.
- Update learner progress safely under concurrent requests.
- Search article content by relevance.
- Maintain tenant isolation.
- Identify the authoritative source for every derived record.
- Estimate item, document, and partition size.
- Estimate read and write traffic by key.
- Identify possible hot partitions.
- Define consistency requirements.
- Define backup, restore, and deletion behaviour.
Model-selection Template
| Access Pattern | Candidate Model | Key or Starting Point | Reason |
|---|---|---|---|
| Session by token | Key-value | Session token | Direct exact-key retrieval |
| Course catalog record | Document | Tenant and course ID | Flexible bounded aggregate |
| Course activity by day | Wide-column | Tenant, course, and date partition | Ordered range retrieval within a known partition |
| Prerequisite traversal | Graph | Course node | Variable-depth relationship traversal |
| Enrollment transaction | Relational | Learner and course keys | Constraints and transactional relationship updates |
| Article keyword search | Search index | Analyzed search query | Full-text retrieval and ranking |
Frequently Asked Questions
What is NoSQL?
NoSQL is a broad category of databases using non-relational or multi-model approaches such as key-value, document, wide-column, and graph models.
What is a key-value database?
It stores a value under a unique key and is strongest when the application already knows the required key.
What is a document database?
It stores self-contained documents containing scalar fields, nested objects, and arrays.
What is a wide-column database?
It organizes data around partition-oriented records and commonly supports ordered retrieval through clustering or sort keys.
What is a graph database?
It represents entities as nodes and their connections as relationships or edges.
What is access-pattern-first design?
It means designing keys, records, partitions, and indexes around the operations the application must perform.
Should related data be embedded in one document?
Embed when the data is bounded, owned by the aggregate, and usually read or updated together. Reference data that is large, shared, unbounded, or accessed independently.
What is a hot partition?
A hot partition receives a disproportionate share of storage operations because of an unsuitable or highly concentrated partition-key pattern.
Does NoSQL mean no schema?
No. Applications still need field definitions, validation, identifiers, versions, and migration strategies even when the database allows flexible records.
Are NoSQL databases always eventually consistent?
No. Consistency and transaction options depend on the database, configuration, operation, and deployment model.
Can one application use several database models?
Yes. This is commonly called polyglot persistence, but every additional datastore increases operational and consistency complexity.
How should a NoSQL model be selected?
Select it from access patterns, data shape, scale, key distribution, consistency, transaction, security, recovery, and operational requirements.
Key Takeaway
NoSQL is not one data model. Key-value databases provide direct retrieval when the key is known. Document databases store flexible, bounded aggregates. Wide-column databases organize large distributed datasets around partition keys and ordered access. Graph databases make relationship traversal central to the model. Begin with the application's reads, writes, ranges, traversals, transaction boundaries, and traffic distribution. Then design stable keys, bounded records, query-specific projections, consistency behaviour, and recovery procedures. Avoid unbounded documents and traversals, hot partitions, uncontrolled denormalization, and unnecessary datastore variety. The best database is the simplest one that satisfies the verified access patterns and operational requirements.