Work through four buckets. The goal is not to ask many questions, it is to ask the few whose answers change the architecture.
Users
- Who uses this, and how many daily active users?
- Global or single region? (Global implies replication, data residency and latency trade-offs.)
- Is there a mobile client, an offline mode, or a public API?
Features
- What are the two or three core actions? What can we defer?
- Is there an existing system this must interoperate with?
Scale
- Read heavy or write heavy, and roughly what ratio?
- What is the peak to average ratio? A 10x daily spike changes your capacity planning far more than the daily average does.
- How large is a single item - a tweet is bytes, a video is gigabytes.
- How fast is it growing, and how long must data be retained?
Quality (non-functional)
- What latency is acceptable, expressed at p99 rather than as an average?
- How much downtime is tolerable? "Highly available" is not a requirement; 99.9% is.
- Is stale data acceptable, and for how long? This is the consistency question, and it is usually the single most design-shaping answer you will get.
- What is the cost of losing data versus the cost of being unavailable?
How to ask them. Prefer proposing an assumption over asking an open question: "I will assume 10 million DAU and a 100:1 read to write ratio" moves faster than "how many users are there?" and still invites a correction.
Write the answers down where everyone can see them. Unstated assumptions are the main cause of wrong designs, and a visible list is what you return to when the interviewer changes something later.