Our strongest on-call engineer quietly stopped picking up shifts. It took us three months to notice, and one honest conversation to understand why.

She was the person everyone wanted paged. She solved things in ten minutes that took the rest of us an hour. Which is exactly why she kept getting paged, on shifts that weren't even hers, until she quietly stopped answering her phone at all.

The rotation that punished competence

Our old schedule was a simple weekly round-robin across seven engineers. What it didn't account for was secondary escalation — when the primary didn't respond fast enough, PagerDuty would fall through to whoever had the best resolution history, regardless of whose week it was. Over six months, one engineer absorbed 31% of all pages while carrying only 14% of primary shifts. Being good at the job was being punished with more of the job.

What we changed

We rebuilt the escalation policy around fairness, not speed of resolution, and added a hard cap on pages-per-person-per-month that pages a manager, not the engineer, when it's exceeded.

escalation_policy: helix-core-prod
primary: rotation(weekly, 7-person pool)
secondary_fallback: next_in_rotation   # not "best resolver"
overflow_after_min: 8
monthly_page_cap: 12
on_cap_exceeded:
  action: alert_manager
  remove_from_secondary_pool: true
  duration: rest_of_month
"I didn't want to say I was drowning. I just wanted the pages to stop making sense as a compliment."

The uncomfortable part

Fixing the schedule was easy. The harder fix was cultural: we had to stop praising fast individual resolution in all-hands and start praising boring, well-distributed weeks. We now report rotation fairness — the ratio between a person's highest and lowest page-count month — as a team-health metric next to MTTR. A ratio above 2.5x triggers a schedule review, no questions asked.

It also meant retraining managers who'd been quietly relying on her as an informal safety net. More than one had gotten into the habit of privately pinging her directly when a page looked scary, bypassing the rotation entirely, because they trusted her judgment more than whoever was actually on shift. That habit doesn't show up in any on-call metric, but it was arguably worse than the escalation-policy problem, because it couldn't be fixed with a config change — it had to be named out loud, in a room, as a behavior that needed to stop.

  • Escalation policies optimized for speed will quietly funnel load onto your best people until they burn out.
  • A hard monthly page cap that removes someone from the pool is more effective than asking them to say no.
  • Track load distribution, not just MTTR — a great average can hide a terrible outlier.