We started a monthly retro-of-retros just for the on-call rotation itself, separate from incident postmortems. Trust, it turns out, compounds slowly and erodes fast.

Incident postmortems fix individual failures. They rarely surface the slow drift in how a team feels about carrying the pager at all. So we started a separate, monthly, thirty-minute retro whose only subject is the rotation itself — not any specific incident.

The first meeting told us more than we expected

In month one, the strongest theme wasn't about tooling or alert quality. It was about response time to Slack messages from the on-call engineer during an active incident — specifically, how long it took a non-on-call teammate to acknowledge a request for help. Median time was 11 minutes. Nobody had ever measured it because it wasn't in any SLO. But it was the single thing engineers cited most often when asked what made a shift feel lonely versus supported.

rotation_retro_template.md (monthly, 30 min, no incident review)
1. What made a shift this month feel supported?
2. What made a shift feel isolating?
3. Any pattern in who responds when the on-call engineer asks for help?
4. One thing to try differently next month (max 1 — protect focus)
"I wasn't afraid of the outage. I was afraid of paging someone at 2am and getting silence for fifteen minutes while the customer impact grew."

What we changed because of it

We introduced a lightweight "buddy" system: every on-call shift has a named, non-paged buddy from the same team, expected to respond to a help request within 5 minutes during their normal waking hours, and explicitly told it's fine to be unavailable outside those hours. It's not a second on-call rotation — it's a social safety net layered on top of the technical one. Median help-response time dropped to 4 minutes within two months.

Why we keep this retro separate

Mixing it into incident postmortems would bury the trust signal under technical detail every time. A dedicated, low-stakes, recurring space for "how did carrying the pager feel" surfaces things that no incident timeline ever will, because loneliness during a shift isn't a root cause of anything — it's just true, and worth hearing on its own.

  • Run a separate, recurring retro for the rotation itself, distinct from incident postmortems — different questions surface different truths.
  • Response time from teammates during a shift is a trust signal worth measuring, even if it never shows up in an SLA.
  • A lightweight buddy system, with honest limits on availability, can rebuild support faster than expanding the formal rotation.