Test types, each answering a different question:
Metrics, in two groups. Watch latency percentiles (p50, p95, p99, p99.9 - never the mean), throughput (successful requests per second, not attempted), and error rate by type, since timeouts and 5xx mean different things. Alongside those, watch saturation: CPU, memory, disk I/O, network, and critically queue depths and connection pool utilisation, which saturate before CPU does.
The key reading: find the knee of the curve. As load rises, throughput increases while latency stays flat - until a point where throughput plateaus and latency climbs sharply. That knee is your real capacity. Beyond it, throughput often decreases as the system spends its time on timeouts and retries.
Common mistakes: testing with an unrealistically small dataset so everything fits in cache; using uniform random keys when production is heavily skewed; ignoring think time so the test is nothing like real user behaviour; and load testing from one machine, making the client itself the bottleneck.