✏️ Explanatory Question

How do you estimate the read to write ratio and use it to drive design decisions?

👁 1 Views
📘 Detailed Answer
🟢 Easy
No previous question
No next question
💡

Answer with Explanation

Estimating it: reason from user behaviour, not from a guess. For each core action, ask how many people consume what one person produces.

  • Social feed: a user posts 0.5 times a day and views 100 posts. Additionally one post reaches many followers, so the ratio compounds - typically 100:1 to 1000:1.
  • E-commerce: hundreds of product views per order placed - roughly 200:1.
  • Chat: one message sent is read once or a few times in a group - close to 1:1 to 5:1.
  • Analytics ingestion: millions of events written, queried by a handful of dashboards - inverted, maybe 1:1000 in favour of writes.

How each ratio drives design:

Very read-heavy (100:1+): justifies caching as the primary lever, read replicas, denormalisation, and precomputation - do expensive work once on write so reads are trivial, such as fan-out-on-write for timelines. Accept slower, more expensive writes; they are rare.

Balanced (near 1:1): caching helps far less because entries are invalidated almost as fast as they are populated. Focus on efficient direct storage access and partitioning.

Write-heavy (inverted): caching is nearly useless. Go to LSM-tree stores, batching, partitioning, and pre-aggregation. Read replicas do not help, since each replica absorbs every write anyway.

The sentence to say: "With a 200:1 read ratio, I will happily make writes 10x more expensive to make reads 10x cheaper - that is a 20x net win."

No previous question
No next question