Kinetic observability, reimagined
  • Home

    Home Styles

    • Terminal
    • Blueprint
    • Signal
  • Product

    Platform

    • Features
    • Integrations
    • Changelog

    Content

    • Blog
    • Article Archive

    Discover

    • All Categories
    • Our Writers
  • Pricing
  • Company

    Company

    • About
    • Team
    • Contact
  • Pages

    Account

    • Login
    • Register
    • Password Reset
    • Username Reminder

    Content

    • Engineering
    • Single Article
    • Author

    Discover

    • All Tags
    • Search
    • Landing
Sign in Start free
Kinetic
  • Home
    • Terminal
    • Blueprint
    • Signal
  • Product
    • Features
    • Integrations
    • Changelog
    • Blog
    • FAQ
    • All Categories
    • Article Archive
    • Our Writers
  • Pricing
  • Company
    • About
    • Team
    • Contact
  • Pages
    • Landing
    • Single Article
    • All Tags
    • Search
    • Login
    • Register
    • Password Reset
    • Username Reminder
    • Engineering
    • Author
Sign in Start free
BLOG

Field notes from the on-call

Engineering deep-dives, incident retros and observability practice from the Helix team.

Browse
  • All posts
  • Engineering
  • Incident retros
  • Product
  • Tutorials
  • Practice
  • Culture
  • incident
  • metrics
  • slo
  • alerting
  • cost
  • latency
Topics All posts 241 Engineering Incident retros Product Tutorials Practice Culture
All articles // 61 posts · newest first
11 Mar
FIELD NOTES FROM THE ON-CALL

How we survived a retry storm (and rewrote our backoff)

A brief database hiccup turned into a 40-minute outage because every client retried at once. Here is the anatomy of a retry...

Tomás Rivera Mar 11 2026 2 min
05 Mar
INCIDENT

The Night Our Runbook Lied to Us

Step 4 said restart the ingest workers. Step 4 had been wrong for three months, and nobody had run it since the architecture...

Lena Park Mar 5 2026 3 min
02 Mar
FIELD NOTES FROM THE ON-CALL

Exemplars: the missing link between metrics and traces

A spike on a latency chart tells you something is wrong. An exemplar tells you exactly which request to open. It is the shortest...

Lena Park Mar 2 2026 2 min
26 Feb
INCIDENT

Six Months of On-Call Taught Me to Trust Dashboards Over Gut Feeling

I used to open five tabs and guess. Now I open one dashboard and let the data argue with my instincts, which turn out to be...

Priya Raman Feb 26 2026 3 min
19 Feb
INCIDENT

How a Single Dropped Index Took Down Checkout for 40 Minutes

A cleanup migration dropped an index nobody remembered was load-bearing. Checkout latency went from 90ms to 12 seconds in under...

Marco Vidal Feb 19 2026 3 min
18 Feb
FIELD NOTES FROM THE ON-CALL

The four golden signals, and the fifth everyone forgets

Latency, traffic, errors, saturation — the four golden signals are a great start. But the signal that actually predicts your...

Arjun Mehta Feb 18 2026 2 min
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • 11

Page 9 of 11

// get started in minutes

Your next incident is coming. Be ready for it.

Free for 14 days. No credit card. Pipe your first logs in under five minutes.

Start free $helix init
Helix

Observability without the overhead. Logs, metrics and traces on one fast timeline.

Product

  • Features
  • Pricing
  • Integrations
  • Changelog

Developers

  • Blog
  • HelixQL
  • Search
  • All Tags

Company

  • About
  • Team
  • FAQ
  • Contact

Monthly dispatch

Incident retros, query tips and product news. No spam, unsubscribe anytime.

© 2026 Helix Labs, Inc. All rights reserved.
  • Privacy
  • Terms
  • Status