We spent a year buying every remote-collaboration tool on the market before realizing our fully distributed reliability team's real problem wasn't tooling at all.
New incident dashboard, new async standup app, new virtual whiteboard, new presence indicator — we bought all of it, and our distributed reliability team's incident response times barely moved. The tools weren't wrong. They just weren't the bottleneck.
What was actually slowing incidents down
We dug into transcripts of a dozen incident channels and found a pattern that no tool could fix: engineers on the team hesitated to jump in and help on an incident owned by someone they'd never met in person, even when they clearly had relevant context. In a colocated team, someone would just walk over. Remotely, without an established relationship, people defaulted to staying out of each other's way.
"I saw the incident channel, I knew exactly what was wrong, and I still typed and deleted my message twice before sending it, because I didn't know if jumping into someone else's incident would read as helpful or as overstepping."
What we built instead of another tool
We started pairing engineers across regions for low-stakes, non-incident work — joint runbook writing, shared chaos-engineering exercises — specifically so the first time two people worked together wasn't during a real outage.
cross_region_pairing_program
cadence: 1 pairing session per engineer per month, 60 min
work_type: NON-incident only — runbook authoring, chaos drills,
alert tuning. Deliberately low stakes.
matching_rule: >
Prioritize pairs who have never worked an incident together.
Rotate so every engineer has met, worked with, and can put a
voice to at least 6 teammates outside their immediate pod
within 2 quarters.
What moved because of it
Cross-region "I saw this and jumped in" incident contributions rose by roughly 60% over two quarters, without any change to tooling, dashboards, or escalation policy. The bottleneck had never been visibility into what was happening. It was the social permission to act on it with someone you'd never met.
- Before buying more remote-collaboration tooling, check whether the bottleneck is actually visibility, or whether it's an unestablished relationship.
- Pair distributed engineers on low-stakes, non-incident work specifically so incidents aren't the first time they collaborate.
- Voluntary cross-team incident help is a trust metric — track it, it moves faster with relationships than with dashboards.