Standardizing on OpenTelemetry was one of our best decisions. It was also, for one memorable quarter, one of our most painful. Both things are true.
What it gets right
A single, vendor-neutral way to describe metrics, logs, and traces is genuinely transformative. Instrument once, send anywhere. The semantic conventions mean a http.route attribute means the same thing across every service and language you run.
- One SDK surface instead of one per backend.
- Auto-instrumentation that gives you traces before you write a line.
- A collector that lets you reshape and route telemetry without redeploying apps.
Where it bites
The collector is powerful and, configured carelessly, a foot-gun. A misplaced processor can double your data volume or silently drop spans. The SDK's defaults are conservative; the interesting features hide behind flags you have to know exist.
The context-propagation cliff
The single sharpest edge is context propagation across async boundaries. Miss it and your traces fragment into disconnected pieces that look fine until you need them whole.
Advice from the far side
Adopt the collector early and treat its config like production code — reviewed, versioned, tested. Turn on the features you need deliberately rather than trusting defaults. And validate your traces are connected end to end before you rely on them, not during your first incident.