# Why the on-call rotation matters more than the initial architecture

- Published: 2026-09-30
- Reading time: 8 min
- Tags: Operations, Architecture, On-call

A system "designed to scale" and a system "designed to be woken up for" are not the same thing. The difference is only visible to whoever actually answers the page.

### Two designs that look identical on paper

Every architecture document we have seen talks about scale: how many services, which database, where the queue sits. Few talk about two in the morning. Those are different questions with different answers. A system that scales well can misbehave quietly for hours under a partial failure, and nobody notices — because nothing in its design ever asked "when this piece dies, what does the person answering the page actually see?" That second question only makes it into the architecture when the team that builds the system is the same team that carries the pager. When the two are separate — one team ships, another team answers calls at 3 a.m. — the incentive to design for diagnosability erodes quietly, because the cost of skipping it is paid by someone who was not in the design meeting.

### What deserves a page, what deserves a dashboard

Every metric you collect is worth displaying. Not every metric is worth waking someone up for. Separating the two is one of the more consequential architectural decisions, and it is usually made after the system is built rather than before — which is itself a sign something was skipped. A workable rule: a page fires only when (1) a real user is feeling the effect right now, (2) human action right now changes the outcome, and (3) the problem will not resolve itself by morning. Anything that fails one of those three belongs on a dashboard, not in someone's pocket while they sleep. A number that "looks bad" with no available action just produces alert fatigue — and alert fatigue is exactly what makes the real alert ignorable. The practical consequence is that the list of pages that actually wake a human should stay short and get pruned on purpose. Any alert that has been dismissed three times in a row with no real action taken should either be deleted or downgraded to a next-business-day ticket, not a midnight call.

### Runbooks are part of the architecture, not an attachment

An architectural decision with no documented recovery path is an unfinished decision. Runbooks tend to get written after the system exists, and often only after the first incident — which is exactly when they are least useful. If the architecture is written from the start with the question "when this component fails, what is the first diagnostic step, and what is the first corrective step?", the system itself usually ends up simpler, because the part that is hard to explain is usually the part that was poorly designed to begin with. A good runbook is not a step-by-step script; it is a decision tree: this symptom implies this likelihood, this likelihood implies this check, this check implies this action. Written any other way, whoever reads it at 2 a.m. with no context just sees a long document.

### Why the builder should carry the pager

When the engineer who designed a scheduled job, an index or a queue is the same person who might be called about it at 3 a.m., design-time behavior changes — not out of fear, but because observability and recoverability stop being abstract and become personal. The difference between "this query is a little slow" and "this query wakes me up" is one only felt by someone who has experienced both. This is also why a small, deep team beats a large, layered one. In a large team, responsibility gets distributed until no one fully owns a decision, and as a result no one feels its full cost either. End-to-end ownership — architecture through deployment through the on-call rotation — is the only reliable way for the real cost of a design choice to reach someone who can actually change it, in time.

### Designing for graceful degradation, not just uptime

Most architecture discussions center on how the system stays "up." A less common question is what the rest of the system looks like when one piece actually fails. A recommendation service that goes down should show the home page without recommendations, not throw an error on the whole page. A message queue that fills up should drop low-priority messages, not lock out the writer that produces them too. Graceful degradation isn't a feature you can bolt on later; it has to be designed at the boundary of every external dependency from day one: what is the default behavior when this dependency doesn't respond? If the answer is "we don't know, we assumed it always responds," that boundary is a total-failure boundary — and every total-failure boundary guarantees a midnight page eventually. The goal isn't to never get paged; it's for a page to only fire for something that genuinely couldn't be degraded gracefully.

### The rotation itself is a feedback loop

The list of who gets called each week is one of the most honest health metrics an architecture has — more honest than any dashboard. If one specific person, week after week, gets woken up for one specific class of failure, the problem isn't that person; it's that the corresponding part of the system was never actually finished. The common and wrong fix is adding a second person to the rotation to spread the load. That only spreads the human cost of the problem around — it doesn't fix the problem itself. The right fix is to treat every repeat page as architectural feedback: why did this happen again, and why wasn't last time's fix enough? A team that takes its rotation seriously reviews the list of repeat pages every few weeks — not just to measure how often people got woken up, but to find the one architectural decision that's still overdue. Skip that review, and the rotation just records how fragile the system is, without the fragility ever getting treated.

### Blast radius: boundaries that have to be drawn in advance

Every architectural decision has a blast radius: when this piece breaks, what else goes down with it? The question that usually gets asked too late is whether that radius was deliberately limited, or just happened to stay small by accident. A message queue shared across several unrelated features is a classic case: the day one of those features misbehaves and fills the queue, every other feature that has nothing to do with the problem stalls too — not because the design was wrong exactly, but because nobody asked up front, "if this one breaks, what else falls with it?" Drawing those boundaries has an upfront cost: separate queues, a separate database, or a separate rate limit per feature is usually slower to build and more complex than a shared resource. But that cost is paid once, at design time; the cost of not drawing those boundaries gets paid repeatedly, every time a small failure in one corner of the system takes half the product down with it. "What's the blast radius of this decision?" deserves to be asked as seriously as "how well does this decision scale?" — and usually isn't, because the answer isn't interesting at design time. It only becomes interesting during the incident.

### A checklist before the code is written

Ask these before signing off on a new piece of architecture, not after the first incident: If this component fails, what signal reaches a human, and is that signal enough to point at the next step? Does this alert genuinely need a middle-of-the-night human, or does it resolve by morning? Can the runbook be read and understood in three minutes, half-asleep? Is the person building this the same person who answers when it rings? If the answer to any of those is "we don't know," the architecture is not finished — only the code is.

### This belongs in the timeline and scope estimate too

When a project's timeline gets set with a client, the cost of writing runbooks, defining alert thresholds and designing graceful degradation usually doesn't show up in the initial estimate, because these aren't "features" the client explicitly asked for. But they're exactly what separates a system that's still maintainable in its second year from one that turns into a permanent source of stress by then. Being honest about that cost at the moment scope gets defined is cheaper than hiding it until the first midnight page goes off.


---

Source: https://larsima.com/bg/insights/on-call-shapes-architecture
Organisation: LARSIMA (شرکت فن‌آوران توسعه لار سیما), registered in Iran, no. 318515, since 2007-12-02.
