The observability stack usually looks healthy. Metrics are collected, traces are sampled, logs are searchable, dashboards exist for every service.
Then an incident starts at three in the morning and the responder spends the first twenty minutes deciding which of forty dashboards is relevant. The tooling was never the constraint.
Alert on symptoms, investigate with causes
Pages fired on causes produce noise, because most causes are survivable. A node going away, a queue growing briefly, a pod restarting: none of these necessarily mean anyone is having a bad time.
Alerting on what a user experiences produces far fewer pages and each one deserves a human. The cause metrics remain valuable, but as the material you consult after being paged, not as the thing that wakes you.
Every alert needs a first move
The gap between being paged and doing something useful is where incidents get long. An alert that arrives with a link to the relevant dashboard, the last deploy, and one sentence about what it means removes most of that gap.
This is unglamorous work and it pays back on the first incident. If an alert cannot be given a first move, that is usually evidence it should not be a page.
- What this means, in one sentence
- The dashboard scoped to the affected service
- Recent deploys and configuration changes
- The first thing to check, and what to do if it looks fine
Cardinality is a budget
High-cardinality labels are the most useful thing in an observability stack and the fastest way to make it unaffordable. A user identifier on a metric is a bill; the same identifier on a trace or a log line is a diagnosis.
The workable division is to keep metrics low-cardinality and aggregate, and let traces and logs carry the identifiers. Then dashboards stay fast, the bill stays predictable, and the detail is still there when a specific case has to be explained.
Rehearse it, because the first time should not be real
Teams that run game days find broken alert routing, stale runbooks, dashboards that query a decommissioned source and permissions nobody has. All of these are cheap to find on a Wednesday afternoon and expensive to find at 3am.
The point is not to prove the system is resilient. It is to find the parts of the response that only exist in one person's memory, and to write them down before that person is on holiday.



