Skip to content
Gach léargas
  • Observability
  • Grafana
  • SLO

Observability that isn't decoration: dashboards people actually open during an incident

Every company has a metrics store and a stack of dashboards. The real question isn't whether you have observability — it's which screen you actually open during a real incident.

Foilsithe

8 nóim. léitheoireachta

Metrics, traces, logs: each earns its cost differently

We collect three kinds of data about a running system, and each answers a different question. Metrics say "something is wrong," but not where. Traces say "here's where this specific request spent its time," but only for requests that got sampled. Logs say "here's exactly what happened," but only if you already know where to look.

The common mistake is treating all three as if they cost the same. Metrics are cheap — Prometheus makes maintaining a counter or a histogram nearly free, so it's reasonable to collect anything that might one day be useful. Traces are expensive: OpenTelemetry makes full instrumentation possible, but storage and sampling carry a real cost, so you have to choose which paths get fully traced. Logs are the most expensive once volume climbs — and the least useful when written without structure.

The rule that holds up: metrics tell you that something broke, traces tell you where in a request it broke, logs tell you why. If diagnosing a failure starts with grepping logs, your metrics weren't good enough.

Alert thresholds that aren't guesses

Most alert thresholds in the world start as a round number: "page if latency exceeds 500ms." Where did that number come from? Usually nowhere — it just felt reasonable at the moment someone wrote the alert.

A defensible threshold comes from the system's actual behavioral distribution, not a guess. That means looking at historical percentiles (never the mean — a mean hides the tail), understanding what normal actually looks like for this system, and setting a threshold that flags a real deviation from that pattern rather than ordinary daily variance. A static threshold that doesn't know the difference between peak-hour and off-hour traffic either pages constantly for nothing or buries a real failure in noise — and both outcomes end the same way: whoever sees the alert stops taking it seriously.

A good threshold is an engineering decision that should be written down and reviewed, not a number dropped into a config file once and never touched again.

An SLO is a conversation, not a number handed down

A service level objective works when it has been negotiated between whoever builds the system and whoever feels the consequences of its results — not when it's copied from a generic template. "99.9% availability" by itself means nothing until someone has said what that number implies in minutes of acceptable downtime per month, and what it costs to earn those minutes.

The SLO conversation needs to answer: what happens when the error budget runs out? Does feature release stop? Does the team shift focus to stability? If that question has no clear answer, the SLO is decoration on a dashboard, not a decision-making tool.

The discipline of removing graphs

Dashboards only grow unless someone deliberately prunes them. Every incident tempts a team to add one more graph "so we catch this sooner next time." A few years in, the main dashboard has a hundred panels and nobody opens the whole thing during a real incident — because finding the right panel among a hundred takes time on its own.

The discipline that works: every graph needs an owner who can explain why it's there and what decision it supports. If nobody has looked at a panel in six months, it either gets deleted or moved to an archive dashboard reserved for deep investigation — not the incident dashboard. The incident dashboard has to stay small enough that, under real stress, the eye lands on what matters without hunting for it.

The gap between good tools and good use

Having Prometheus, Grafana and OpenTelemetry in the stack doesn't produce observability by itself; it just supplies the capacity for it. Three teams can hold exactly the same three tools, and one of them finds the cause in three minutes mid-incident while the other two spend an hour in the dark. The difference isn't the tooling — it's what was already sampled, labeled and laid out ahead of time.

Trace sampling is a decision that has to be made on purpose, not left at the library default. Sampling everything on every path makes storage cost indefensible; sampling too little means the exact request you need mid-incident was never captured. A reasonable pattern is keeping a higher sampling rate on high-risk paths and a lower one on low-risk paths — and revisiting those rates like any other decision, rather than setting them once at rollout and forgetting about them.

Labeling matters just as much. A metric with no route or tenant label just says "something somewhere is wrong" when things break, without saying where. Adding the label after the incident is easy; but every time that happens after the fact, the next incident carries the same gap, just in a different dimension.

An incident dashboard is not a management dashboard

A common mistake is building one dashboard that serves both management reporting and mid-incident use. Those two audiences need different decision speeds. A management dashboard is built for a weekly or monthly look and can show trend and aggregation. An incident dashboard has to say, within seconds, where the system is currently hurting, and its panels should be ordered the way diagnosis actually proceeds — from general symptom toward the specific failing path — not by which team happened to build which panel.

When the two get merged onto one screen, the incident dashboard usually loses: aggregate, monthly panels crowd out real-time ones, and whoever opens the screen mid-incident has to search past what's irrelevant to reach what matters. Keeping the two separate costs one extra screen; the payoff is that, under real stress, the right screen is already open.

The three signals need to connect, not just coexist

Collecting all three kinds of data isn't enough if you can't jump from one to the other. A metric showing latency spiked should link directly to traces from the same window and the same route, and a trace should link to the logs for that specific request. If that link only exists in an engineer's head — who has to manually copy a request ID from one tool into another every time — an incident burns exactly the few minutes that mattered most.

That link has to be designed in from day one: the trace ID in the metric's label, that same ID in a structured log field. It's technically simple, and it gets skipped constantly, because none of the three tools demands this connection on its own — its absence only becomes obvious in the middle of a real incident, when it's too late to add.

A short checklist

  • Does this panel, this metric, this alert tie to an actual decision?
  • Did the threshold come from data or from a guess?
  • Was the SLO negotiated with whoever feels its outcome, or just copied?
  • When was the last time a graph got deleted?

If the answer to the last one is "can't remember," your dashboard is accumulating decoration, not signal.

Cardinality: where metric generosity runs out

We said metrics are cheap, but that cheapness has one exception that doesn't get taken seriously enough: label cardinality. A counter labeled by HTTP status code has a handful of possible values. The same counter labeled by user ID creates a separate time series per user — and that's exactly what buckles a metrics store like Prometheus under load, not request volume.

The rule that holds up: labels are for categorizing, not identifying. A specific entity's identifier — a user, a request, a session — belongs in a trace or a structured log, not in a metric label. That decision needs to be made before the first line of instrumentation is written, because unwinding a cardinality explosion after the metrics store is already buckling under it is far more expensive than preventing it.

An bhfuil rud éigin le tógáil agat?

Inis dúinn cad air a bhfuil tú ag obair. Inseoimid go macánta duit an muidne an fhoireann cheart dó.

Cuir comhrá ar bun

nó seol ríomhphost chugainn ag hello@larsima.com