Averages lie about latency, so we moved to percentiles. Then we learned that percentiles lie too — just more quietly.
Why averages fail first
A mean latency blends the fast majority with the slow tail into a single reassuring number. Ten thousand snappy requests hide a thousand terrible ones. Nobody experiences the average; they experience their own request.
Why p99 is not enough either
The p99 is better, but it has two traps. First, at scale, one percent of requests is an enormous number of unhappy people. Second, percentiles cannot be averaged — averaging the p99 across ten servers gives you a number that means nothing.
- Watch p99.9 and p99.99, not just p99, if your traffic is high.
- Compute percentiles from raw histograms, never by averaging pre-computed percentiles.
- Slice by route and customer — a healthy overall p99 can hide one endpoint that is on fire.
Compute it correctly
traces
| where service == "api"
| stats p99 = percentile(duration, 99),
p999 = percentile(duration, 99.9)
by route
| sort p999 desc
The tail is the product
For most users, your slowest ten percent of requests is your reputation. Chasing the tail is not a vanity exercise — it is where churn hides. Escape the trap by refusing to let a single number stand in for a distribution.