basic data lifecycle
Basic Data Lifecycle
Learn how data moves from creation and collection through validation, storage, processing, use, sharing, archival, retention, and secure deletion, and how ownership, metadata, quality, security, recovery, and observability apply at every stage.
Introduction
Data does not remain in one place or one state forever. It is created, collected, validated, stored, transformed, queried, shared, archived, and eventually deleted.
This complete journey is called the data lifecycle.
Understanding the lifecycle helps system designers answer practical questions:
- Where did the data come from?
- Who owns it?
- Can the application trust it?
- Where should it be stored?
- Who may access it?
- How long should it remain?
- Can it be recovered after a failure?
- When should it be archived or deleted?
- How can deletion be verified?
Core idea: Data lifecycle management is not only about storage. It coordinates data quality, ownership, security, access, processing, retention, recovery, cost, and deletion throughout the useful life of the data.
Different frameworks divide the lifecycle into different numbers of stages. This article uses a practical system-design flow:
Plan
|
v
Create or Collect
|
v
Validate and Classify
|
v
Store and Protect
|
v
Process and Transform
|
v
Use and Share
|
v
Monitor and Maintain
|
v
Archive or Retain
|
v
Delete or Destroy
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Block, file, and object storage | Lifecycle stages can place data in different storage models. |
| 2 | Durability | Important data must survive the failures covered by its storage contract. |
| 3 | Checksums | Integrity verification helps detect corrupted stored or transferred data. |
| 4 | Metadata | Ownership, classification, retention, provenance, and lifecycle state are metadata. |
| 5 | Databases and transactions | Lifecycle operations frequently update authoritative records and related data atomically. |
| 6 | Security fundamentals | Authentication, authorization, encryption, audit, and deletion apply throughout the lifecycle. |
What Is a Data Lifecycle?
A data lifecycle is the sequence of stages through which data passes from its initial creation or collection to its final deletion or destruction.
Beginning of lifecycle:
A learner creates an account.
A course instructor uploads a video.
An application records a quiz attempt.
A service receives an event.
Middle of lifecycle:
Data is validated.
Data is stored.
Data is transformed.
Data is searched.
Data is shared with authorized services.
Data supports application decisions.
End of lifecycle:
Data becomes inactive.
Data moves to archive.
Required retention expires.
Data is deleted according to policy.
Not every dataset follows exactly the same path. Some temporary data can be deleted soon after processing. Financial or audit data can require a longer retention period. A derived search index can be rebuilt instead of backed up like its authoritative source.
Data Lifecycle vs Data Pipeline
| Concept | Primary Focus | Example |
|---|---|---|
| Data lifecycle | The complete existence of data from creation to deletion | Collect, store, use, retain, archive, and delete learner data |
| Data pipeline | The movement and transformation of data between systems | Extract quiz attempts, transform the records, and load an analytics table |
| Data workflow | Coordinated operations that produce a business outcome | Upload, verify, scan, transcode, and publish a lesson video |
A pipeline can operate during one or more lifecycle stages. The lifecycle is the broader management model.
Stage 1: Plan
Lifecycle management should begin before data is collected.
Planning defines:
- The business purpose
- The authoritative owner
- The minimum required fields
- The expected data volume
- The required quality
- The classification and sensitivity
- The permitted uses
- The storage model
- The retention period
- The recovery requirements
- The deletion process
Learning-platform Example
Data:
Learner quiz attempts
Purpose:
Display results,
calculate progress,
and support approved learning analytics.
Owner:
Learning-platform domain
Authoritative store:
Relational database
Retention:
Defined by approved business
and organizational requirements.
Derived copies:
Progress dashboard
and analytics dataset
Minimization rule: Do not collect data simply because the application might use it someday. Every collected field should have a defined purpose, owner, protection level, and lifecycle.
Stage 2: Create or Collect
Data enters the system through user actions, application events, sensors, uploaded files, APIs, database transactions, imports, or external sources.
| Source | Example |
|---|---|
| User input | Learner profile or course registration |
| Application transaction | Enrollment or quiz-attempt record |
| File upload | Course video, PDF, image, or source-code file |
| Service event | CoursePublished or EnrollmentCompleted event |
| External integration | Approved data received from another business system |
| Derived generation | Search document, report, thumbnail, or aggregate |
Collection Flow
Source produces data
|
v
Authenticate source
|
v
Authorize operation
|
v
Validate format and size
|
v
Attach provenance metadata
|
v
Accept or reject data
Collection should record enough provenance to explain where the data came from and when it entered the system.
Stage 3: Validate and Classify
Newly collected data should not automatically be treated as complete, accurate, safe, or trustworthy.
Validation can include:
- Required-field checks
- Data-type checks
- Length and range checks
- Allowed-value checks
- Uniqueness checks
- Referential-integrity checks
- File-size and content-type checks
- Checksum verification
- Schema validation
- Duplicate detection
- Business-rule validation
SQL Validation Example
CREATE TABLE learner_progress
(
learner_id BIGINT NOT NULL,
course_id BIGINT NOT NULL,
completion_percentage DECIMAL(5, 2) NOT NULL,
progress_status VARCHAR(30) NOT NULL,
last_activity_at TIMESTAMP NOT NULL,
PRIMARY KEY
(
learner_id,
course_id
),
CONSTRAINT ck_completion_percentage
CHECK
(
completion_percentage
BETWEEN 0 AND 100
),
CONSTRAINT ck_progress_status
CHECK
(
progress_status IN
(
'not-started',
'in-progress',
'completed'
)
)
);
Classification
Classification identifies how the data should be handled.
Possible classifications:
Public
Internal
Confidential
Restricted
Possible data categories:
Business data
Account data
Course content
Operational logs
Audit data
Temporary processing data
Use the classifications and handling rules approved by the relevant organization. The example labels above are illustrative.
Stage 4: Store and Protect
Validated data is stored using an appropriate storage model.
| Data | Candidate Storage | Reason |
|---|---|---|
| Enrollments and quiz attempts | Relational database | Structured relationships, transactions, constraints, and queries |
| Lesson videos and documents | Object storage | Large independently addressable content with metadata |
| Shared legacy files | File storage | Applications require shared paths and file-system semantics |
| Database volume | Block or managed database storage | Supports database-engine storage operations |
| Search projection | Search index | Optimized for textual retrieval and ranking |
Protection can include:
- Least-privilege access
- Encryption in transit
- Encryption at rest
- Checksums
- Replication or erasure coding
- Snapshots and backups
- Versioning
- Audit records
- Recovery testing
Storage rule: Select storage according to the required access pattern, update behaviour, durability, recovery, security, scalability, and cost, not only according to data size.
Stage 5: Process and Transform
Stored data can be cleaned, normalized, enriched, aggregated, converted, or indexed for downstream use.
Raw data
|
v
Validate
|
v
Clean
|
v
Standardize
|
v
Transform
|
v
Enrich
|
v
Publish prepared data
Media Example
Original lesson video
|
v
Integrity verification
|
v
Security analysis
|
v
Metadata extraction
|
v
Video transcoding
|
v
Thumbnail generation
|
v
Approved renditions published
Search Example
Published article
|
v
Extract searchable fields
|
v
Tokenize and normalize text
|
v
Build search document
|
v
Update inverted index
Raw, Prepared, and Derived Data
| Category | Meaning | Example |
|---|---|---|
| Raw data | Data close to the received source representation | Original uploaded CSV or source event |
| Prepared data | Validated and transformed data ready for a defined use | Normalized quiz-attempt dataset |
| Derived data | Data calculated from other data | Course-completion percentage or search index |
Derived data should retain lineage to the authoritative inputs, processing version, and generation time when reproducibility matters.
Data Lineage
Data lineage records where data originated, how it moved, and which transformations were applied.
Source quiz attempts
|
v
Validation pipeline
|
v
Normalized attempt records
|
v
Course-progress calculation
|
v
Progress dashboard
|
v
Approved analytics report
Useful lineage metadata includes:
- Source system
- Source record or object identifier
- Pipeline name
- Pipeline version
- Transformation time
- Input and output schema versions
- Quality-check results
- Responsible owner
Stage 6: Use and Analyze
Authorized applications, users, reports, search systems, and analytical processes consume the prepared data.
Uses can include:
- Application transactions
- Operational dashboards
- Business reports
- Full-text search
- Notifications
- Data analysis
- Machine-learning features
- Audit and compliance review
Access Flow
User or service requests data
|
v
Authenticate identity
|
v
Authorize purpose and resource
|
v
Apply tenant and data filters
|
v
Return minimum required fields
|
v
Record audit evidence where required
Access should follow least privilege and purpose limitation. A service should receive only the data necessary for its approved responsibility.
Stage 7: Share and Distribute
Data can be shared between internal services, teams, applications, partners, reporting systems, or approved external recipients.
A sharing contract should define:
- Producer and consumer
- Purpose
- Schema
- Data owner
- Permitted fields
- Freshness
- Quality expectations
- Security requirements
- Retention
- Error handling
- Deletion propagation
Event-sharing Example
{
"eventType": "EnrollmentCompleted",
"eventVersion": 1,
"eventId": "generated-event-id",
"tenantId": "tenant-17",
"learnerId": 1042,
"courseId": 42,
"occurredAt": "event-time"
}
Avoid placing unnecessary personal or confidential fields into broadly distributed events.
Stage 8: Monitor and Maintain
Active data needs continuous management. Data can become stale, duplicated, inconsistent, corrupted, inaccessible, or incorrectly classified.
Monitor:
- Data freshness
- Completeness
- Validity
- Duplicate rate
- Integrity failures
- Schema changes
- Pipeline failures
- Access anomalies
- Storage growth
- Backup age
- Retention exceptions
- Deletion backlog
Data-quality Dimensions
| Dimension | Question |
|---|---|
| Accuracy | Does the data correctly represent the real entity or event? |
| Completeness | Are required values present? |
| Consistency | Do related systems and fields agree? |
| Validity | Does the data follow its schema and business rules? |
| Uniqueness | Are duplicate representations controlled? |
| Timeliness | Is the data current enough for its intended use? |
| Integrity | Has the data remained complete and uncorrupted? |
Stage 9: Archive and Retain
Data that is no longer needed for frequent operational access can move to an archive or lower-access storage tier.
Frequently used data
|
v
Infrequently used data
|
v
Archive storage
|
v
Retention period expires
|
v
Approved deletion
Archiving is not the same as deletion. Archived data remains stored and must still be protected, discoverable when authorized, included in retention management, and recoverable according to its requirements.
Retention Metadata
{
"retentionClass": "completed-course-record",
"retentionStart": "approved-start-time",
"retainUntil": "approved-end-time",
"legalHold": false,
"lifecycleState": "archived"
}
Retention periods should come from approved legal, regulatory, organizational, contractual, and business requirements rather than from an arbitrary technical default.
Retention vs Backup
| Concept | Purpose |
|---|---|
| Retention | Determines how long data should remain |
| Backup | Maintains a recoverable copy after loss or corruption |
| Archive | Stores inactive data for longer-term authorized access |
| Legal hold | Suspends normal deletion under an approved requirement |
Backup copies must also follow approved retention and deletion requirements.
Stage 10: Delete or Destroy
At the end of the approved lifecycle, data should be deleted or destroyed according to policy.
Deletion becomes eligible
|
v
Check retention requirement
|
v
Check legal or investigation hold
|
v
Authorize deletion
|
v
Delete authoritative data
|
v
Delete or expire derived copies
|
v
Update search, cache, replicas,
exports, and processing systems
|
v
Record deletion outcome
Deletion can involve:
- Primary database records
- Object versions
- File-system copies
- Search documents
- Cache entries
- Analytics projections
- Temporary files
- Queued messages
- Exports
- Backups according to their retention process
Deletion rule: Deleting the primary record does not automatically remove every derived, cached, replicated, exported, indexed, or backed-up copy. Maintain a data inventory and deletion-propagation design.
Soft Delete vs Hard Delete
| Method | Meaning | Consideration |
|---|---|---|
| Soft delete | Mark the record deleted while retaining it temporarily | Queries and authorization must consistently exclude deleted data |
| Hard delete | Remove the active record or stored object | Recovery depends on versions, backups, or another approved mechanism |
| Anonymization | Remove or transform identifying values according to an approved design | The result must be evaluated against re-identification risk and requirements |
Soft-delete Schema Example
ALTER TABLE course_assets
ADD
deletion_status VARCHAR(30) NOT NULL
DEFAULT 'active',
deletion_requested_at TIMESTAMP NULL,
deleted_at TIMESTAMP NULL;
A soft-delete marker is a lifecycle state, not proof that the underlying content has been permanently destroyed.
Lifecycle State Machine
Created
|
v
Validated
|
v
Active
|
+-- Quality failure --> Quarantined
|
+-- Update --> Active new version
|
v
Inactive
|
v
Archived
|
+-- Legal hold --> Retained
|
v
Deletion pending
|
v
Deleted
Explicit state transitions make lifecycle behaviour easier to test, audit, retry, and reconcile.
Lifecycle Metadata Model
CREATE TABLE data_assets
(
data_asset_id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
asset_name VARCHAR(255) NOT NULL,
asset_type VARCHAR(50) NOT NULL,
lifecycle_status VARCHAR(30) NOT NULL,
classification VARCHAR(30) NOT NULL,
owner_id BIGINT NOT NULL,
source_system VARCHAR(100) NOT NULL,
source_reference VARCHAR(500) NULL,
schema_version INT NOT NULL,
retention_class VARCHAR(100) NOT NULL,
retain_until TIMESTAMP NULL,
legal_hold BOOLEAN NOT NULL
DEFAULT FALSE,
created_at TIMESTAMP NOT NULL,
last_verified_at TIMESTAMP NULL,
archived_at TIMESTAMP NULL,
deletion_requested_at TIMESTAMP NULL,
deleted_at TIMESTAMP NULL,
CONSTRAINT ck_data_asset_status
CHECK
(
lifecycle_status IN
(
'created',
'validated',
'active',
'quarantined',
'inactive',
'archived',
'deletion-pending',
'deleted'
)
)
);
Actual classification, retention, and status values should follow approved organizational policies and domain requirements.
Lifecycle Automation
Lifecycle rules can automate selected transitions.
lifecycle:
inactive-content:
condition:
status: inactive
age: approved-inactive-period
action:
move-to: archive
expired-content:
condition:
retention-expired: true
legal-hold: false
action:
queue-deletion: true
abandoned-uploads:
condition:
status: pending
session-expired: true
action:
abort-and-clean: true
This is conceptual configuration. Lifecycle automation must use approved rules, protected permissions, audit evidence, and safe retry behaviour.
Lifecycle Events
{
"eventType": "DataAssetArchived",
"eventVersion": 1,
"eventId": "generated-event-id",
"tenantId": "tenant-17",
"dataAssetId": 981,
"sourceVersion": 14,
"occurredAt": "event-time"
}
Consumers can update search indexes, caches, analytics, and storage tiers in response to lifecycle events.
Event consumers should be idempotent and reject stale lifecycle versions.
Authoritative and Derived Copies
One logical dataset can exist in several forms.
Authoritative learner record
|
+-- Search projection
+-- Cache entry
+-- Analytics dataset
+-- Export
+-- Backup
+-- Audit record
For every copy, document:
- Purpose
- Owner
- Source of truth
- Freshness
- Access rules
- Retention
- Deletion behaviour
- Rebuild or recovery process
Lifecycle Reconciliation
Partial failures can leave systems in different lifecycle states.
Database:
asset = deleted
Object storage:
object still exists
Search index:
document remains searchable
Cache:
download link remains cached
Result:
Lifecycle state is inconsistent.
A reconciliation job compares authoritative state with derived systems and repairs mismatches.
Security throughout the Lifecycle
| Lifecycle Stage | Security Focus |
|---|---|
| Collection | Authentication, authorization, minimization, and input validation |
| Transfer | Encrypted transport, integrity verification, and bounded endpoints |
| Storage | Least privilege, encryption, durability, backup, and key management |
| Processing | Approved environments, temporary-data controls, and protected credentials |
| Use and sharing | Purpose limitation, tenant isolation, access control, and auditability |
| Archive | Restricted access, retention enforcement, integrity, and key availability |
| Deletion | Authorization, hold checks, propagation, and evidence of completion |
Ownership and Responsibilities
Every important dataset should have clear ownership.
| Responsibility | Question |
|---|---|
| Business ownership | Who defines the purpose and acceptable use? |
| Technical ownership | Who operates the storage, pipelines, APIs, and recovery process? |
| Quality ownership | Who defines and reviews quality rules? |
| Security ownership | Who approves protection, classification, and access patterns? |
| Retention ownership | Who approves retention and deletion requirements? |
| Incident ownership | Who responds when data is lost, corrupted, exposed, or unavailable? |
Lifecycle Observability
Useful metrics include:
- Records or objects created
- Validation failure rate
- Quarantined-data count
- Data-quality rule failures
- Pipeline processing delay
- Replication and backup health
- Data freshness
- Storage growth by classification
- Inactive-data volume
- Archive-transition count
- Retention exceptions
- Legal-hold count
- Deletion requests
- Deletion backlog
- Deletion failures
- Orphaned-data count
- Derived-copy synchronization delay
Alert Conditions
Alert when:
- Required validation stops running
- Data-quality failures increase unexpectedly
- Backups or integrity checks fail
- Authoritative and derived copies diverge
- Restricted data is accessed unexpectedly
- Retention-eligible data is not archived or deleted
- Lifecycle automation fails repeatedly
- Deletion remains incomplete across downstream systems
- A legal hold is bypassed or changed unexpectedly
- Storage growth exceeds the approved forecast
Lifecycle-review Workflow
- Identify the dataset and business purpose.
- Identify the owner and authoritative source.
- Inventory all stored and derived copies.
- Document collection and validation rules.
- Document classification and access requirements.
- Document storage, durability, and recovery requirements.
- Document processing and lineage.
- Document permitted consumers and sharing contracts.
- Define data-quality and freshness objectives.
- Define archive and retention requirements.
- Define deletion and legal-hold behaviour.
- Test recovery, archival, and deletion.
- Monitor lifecycle events and reconciliation results.
Common Data-lifecycle Mistakes
Collecting Data without a Defined Purpose
Unnecessary collection increases storage, security, privacy, governance, and deletion responsibilities.
Keeping No Authoritative Source
Database, search, cache, analytics, and export copies can contradict one another.
Trusting Data before Validation
Invalid, incomplete, duplicated, or unsafe data can spread to downstream systems.
Storing Raw and Derived Data without Lineage
Teams cannot explain which source or transformation produced a result.
Applying the Same Access throughout the Lifecycle
Raw, active, archived, quarantined, and deletion-pending data can require different permissions.
Keeping Data Forever
Unlimited retention increases cost and can conflict with approved privacy, contractual, or organizational requirements.
Archiving without a Retrieval Plan
Archived data can become inaccessible when applications, credentials, formats, or encryption keys are unavailable.
Assuming Replication Is Backup
Deletion and corruption can propagate to replicated copies.
Deleting Only the Primary Record
Search, cache, objects, exports, events, analytics, and backups can retain additional copies.
Using Soft Delete as Permanent Destruction
A deletion flag hides the record but does not remove the underlying data.
Automating Deletion without Hold Checks
Lifecycle automation can remove records that must remain retained under an approved exception.
Never Testing the Lifecycle
Collection and use can work while archival, restoration, or deletion fails when first needed.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Valid data collection | The record is accepted with required provenance and metadata |
| Invalid input | The data is rejected or quarantined safely |
| Duplicate data | The uniqueness or reconciliation policy is enforced |
| Classification | The approved handling rules are applied |
| Derived-data lineage | The output can be traced to its source and transformation version |
| Unauthorized access | Data, metadata, counts, and derived results remain protected |
| Backup restore | The required data is restored and passes integrity checks |
| Archive transition | Eligible data moves to the approved storage tier |
| Archived-data retrieval | Authorized retrieval succeeds within its service objective |
| Retention expiry | Data becomes eligible for the approved deletion workflow |
| Legal hold | Normal deletion is prevented while the hold remains active |
| Deletion propagation | Authoritative, search, cache, storage, and derived copies converge |
| Partial deletion failure | Reconciliation retries or escalates the incomplete deletion |
| Audit verification | Lifecycle transitions have the required evidence |
Data-lifecycle Best Practices
Recommended Practices
- Define the business purpose before collecting data.
- Collect only the fields required for approved uses.
- Assign business and technical ownership.
- Identify one authoritative source for each important data element.
- Record provenance, classification, and lifecycle metadata.
- Validate data before active use.
- Use constraints and controlled vocabularies for critical fields.
- Select storage according to access, durability, recovery, and cost requirements.
- Encrypt and authorize data according to its classification.
- Maintain lineage for transformed and derived data.
- Share only the minimum required fields.
- Define quality and freshness expectations.
- Back up authoritative data and test restoration.
- Archive inactive data according to approved policy.
- Define retention at the dataset or record category level.
- Check holds and exceptions before deletion.
- Propagate lifecycle changes to derived systems.
- Make lifecycle events and consumers idempotent.
- Reconcile authoritative and derived copies.
- Test creation, recovery, archival, and deletion end to end.
Practice Exercise
Create a data-lifecycle design for your online learning platform.
Requirements
- Inventory learner accounts, enrollments, quiz attempts, course videos, certificates, logs, and search documents.
- Define the purpose of each dataset.
- Identify the authoritative source.
- Assign the data owner and technical owner.
- Define validation and quality rules.
- Define classification and access requirements.
- Select the appropriate storage model.
- Define backup and restore requirements.
- Document transformations and derived copies.
- Define data-sharing contracts.
- Define active, inactive, archived, and deleted states.
- Define approved retention requirements.
- Define legal-hold handling.
- Design deletion propagation.
- Create a reconciliation job.
- Test one complete lifecycle from creation to deletion.
Lifecycle-design Template
| Data Asset | Authoritative Source | Active Use | Archive and Deletion |
|---|---|---|---|
| Learner account | Application database | Authentication, profile, and learning access | Apply approved account-retention and deletion workflow |
| Course video | Application metadata and object storage | Authorized lesson streaming | Version, archive, or delete according to course-content policy |
| Quiz attempt | Relational database | Scoring and learner progress | Retain or delete according to approved learning-record requirements |
| Search document | Derived from published content | Full-text retrieval | Remove and rebuild from authoritative content |
| Operational log | Logging platform | Troubleshooting and monitoring | Expire according to approved operational and security requirements |
| Backup | Protected backup system | Recovery after loss or corruption | Expire using the approved backup-retention schedule |
Frequently Asked Questions
What is a data lifecycle?
A data lifecycle is the complete journey data follows from creation or collection through storage, use, retention, archival, and deletion.
Why is data-lifecycle management important?
It helps coordinate data quality, ownership, security, access, storage, recovery, retention, cost, and deletion.
Does every dataset follow the same lifecycle?
No. The stages and controls depend on the dataset's purpose, classification, value, storage model, recovery requirements, and approved retention rules.
What is the authoritative source?
It is the system designated as the trusted source for a particular data element or business record.
What is data lineage?
Data lineage records where data originated, how it moved, and which transformations produced its current form.
What is derived data?
Derived data is calculated, transformed, aggregated, or indexed from another data source.
What is data retention?
Data retention defines how long a dataset or record category should remain before archival or deletion.
Is archiving the same as deleting?
No. Archived data remains stored for authorized future access, while deletion removes data according to the approved lifecycle process.
Is soft-deleted data permanently removed?
No. Soft deletion normally changes the lifecycle status while retaining the underlying record.
Why is deletion propagation necessary?
The same logical data can exist in databases, objects, search indexes, caches, analytics datasets, exports, replicas, and backups.
What is lifecycle reconciliation?
Reconciliation compares authoritative lifecycle state with derived systems and repairs or reports inconsistencies.
Should data be kept forever?
No. Retain data only for an approved business, legal, regulatory, contractual, operational, or recovery purpose.
Key Takeaway
The data lifecycle describes how data is planned, created, validated, classified, stored, protected, processed, used, shared, monitored, archived, retained, and deleted. Begin with a defined purpose and owner, collect only required data, identify the authoritative source, validate before use, and maintain metadata and lineage for every important copy. Protect data according to its classification throughout transfer, processing, storage, sharing, and archival. Define quality, freshness, durability, backup, retention, and recovery requirements explicitly. Finally, treat deletion as a distributed workflow that must reach search indexes, caches, objects, analytics, exports, and other derived copies, while respecting approved holds and maintaining evidence of completion.