A realistic on-call scenario. Work along the path rather than guessing.
1. Establish the facts. Which user, which entity, what exact time, what did they expect? Get a request ID or trace ID if possible - this makes everything else easy.
2. Check for the usual suspects, in order of likelihood:
- Stale cache. The most common cause. Compare the cached value with the database value directly. If the database is right and the cache is wrong, invalidation failed.
- Replication lag. The write went to the primary but the read hit a lagging replica. Check the replica lag metric at that timestamp. Fix with read-your-own-writes: route a user to the primary for a few seconds after their write, or pin them via a session token.
- The write silently failed. Check whether the row actually exists. A swallowed exception or a full queue can drop it.
- The async worker is backed up. If the update only becomes visible through a queue consumer, check queue depth and consumer lag - the write may be fine but the derived view is minutes behind.
- CDN or browser caching. Check Cache-Control headers; the request may never have reached your servers.
- Sharded to the wrong place or a partial failure in a fan-out.
3. Voice the systemic fix. "The immediate fix is to purge that cache key, but the real fix is that our invalidation is best-effort. I would add a short TTL as a backstop so any missed invalidation self-heals within 60 seconds."