How trace-driven debugging found a retry storm behind a healthy-looking dashboard.
The dashboard was green. The p50 was fine. And yet checkout felt sluggish, and a stubborn slice of users were timing out. The averages were lying, and traces told the truth.
The symptom
Our headline metrics looked healthy because they were dominated by the fast majority. The pain lived in the p95, where a meaningful fraction of checkouts were taking three times longer than they should.
What the traces showed
Opening the slow traces revealed the culprit immediately: a downstream inventory call was intermittently failing, and our client was silently retrying it three times with no backoff. The retries succeeded, so the request eventually completed — slowly, invisibly.
- The metric hid the retries because the final result was a success.
- The trace exposed them because it showed the same span three times.
- The fix was backoff plus a tighter timeout, not more capacity.
The result
Adding jittered backoff and a circuit breaker cut p95 by forty percent within a day. No new hardware, no rewrite — just the ability to see the shape of a request instead of a summary statistic about it.