- Real-Time Systems
- Reliability
- Observability
Designing for failure: real-time software that assumes the network will drop
Most real-time systems are designed for the happy path and treat network loss as the exception. In the live-call product we built, disconnection is the rule — and that assumption reshapes reconnection, graceful degradation and observability.
8 min qari

If you design a real-time system assuming the network usually works and occasionally drops, a network will eventually surprise you where dropping is the normal state: mobile devices, roaming between Wi-Fi and cellular data, carrier tunnels, crowded rooms. The right design takes the opposite assumption: network loss is a first-class part of the system's behavior, not an exception path tested "whenever there's time." In the live-call product we build, we've implemented this assumption in three distinct places: reconnection without losing state, predictable graceful degradation, and observability at the level of the session rather than the request.
Reconnection without losing state
The important question when a connection drops isn't "will we reconnect" but "once we reconnect, what has to be rebuilt and what has to stay exactly as it was." Two different layers are involved here, and they need to behave differently.
The media layer (audio and video) tries, on its own real-time transport engine, to re-establish the path without restarting the call from scratch. But rebuilding the path itself isn't enough: some settings (received quality, audio output route) can silently revert to the engine's defaults during that rebuild, rather than to whatever the user had actually chosen. The lesson we learned in our own product was not to assume "reconnected" means "everything is back to how it was" — we explicitly reassert those settings after every recovery, because the underlying engine has forgotten them.
The signaling layer (the WebSocket connection carrying session events) follows a different rule: it never panics on disconnect, reconnects with exponential backoff — starting from a short delay, doubling on each failed attempt, capped at a bound — and silently discards malformed frames rather than breaking the whole connection over one bad message. One subtlety matters here: the backoff counter should not be reset just because a socket opened; it should wait for the first real message to arrive, because a socket that opens and immediately drops again, if it resets the counter too early, effectively creates an unbounded rapid-retry loop. This is exactly the bug we once hit in production and fixed.
A smaller but equally real point: state recovery isn't only an infrastructure concern, it's a UI concern too. A user whose video freezes for a few seconds mid-session needs to understand that the app is "recovering," not conclude that it has crashed and close and reopen it themselves — which creates a real, entirely unnecessary disconnection. A UI that completely hides automatic background recovery looks cleaner at first glance, but in practice it pushes the user into making exactly the wrong call.
More important than either of these is session identity. When a connection re-establishes, the user rejoins the same room with the same participant identity and the same valid token — this is not a fresh join, it is a recovery of the same membership. The server side carries the same assumption: a participant is only marked as having "left" once their connection is genuinely and conclusively closed, not at the first sign of trouble — because a mobile network produces plenty of false signals that look like disconnection but aren't.
Graceful degradation, not collapse
When the network gets worse, the right question isn't "keep the call or drop it" but "what do we sacrifice first." We use a degradation ladder: video quality steps down first, then incoming video is dropped entirely, and in the worst case the outgoing camera is turned off too, leaving only voice. Voice is the floor, not something to be traded away — because a call without video is still a call, but a call without audio isn't a call at all.
What actually makes this ladder useful is its time asymmetry: the decision to step down happens fast (because an immediate bad experience is the worst outcome), while the decision to step back up happens slowly and cautiously. If both directions moved at the same speed, an unstable — not bad, just unstable — network would bounce quality up and down every few seconds, which is worse for the user than a steady degraded state. Stability, even at a lower level, matters more than oscillation.
Observability at the session level, not the request level
Ordinary systems observe HTTP around the request: one log line per request, one status code, one response time. A real-time session stays alive for hours, and "healthy" cannot be read off a single log line. What actually needs to be recorded is: when each session connected and disconnected, which attempt number a given reconnection was and how many attempts remain before giving up, and a coarse, periodic signal indicating whether real media is actually flowing — without inspecting its content.
The other layer that needs watching is backpressure. Each client's outbound queue must be bounded; when a client falls behind and can't consume events as fast as they're produced, the system should drop the oldest events and log that drop — rather than letting the queue grow without bound until memory runs out, and rather than letting the whole system wait on one slow client.
What "healthy session" should mean
Pulling this together into one definition: a session isn't healthy just because its socket is open. It's healthy if media is flowing continuously, if its most recent reconnection succeeded and its settings were reasserted, if its outbound queue isn't growing unboundedly, and if — even while degraded — that degradation was predictable and stepwise rather than sudden and total. This definition is deliberately more complicated than "returned 200" — because the underlying problem is more complicated than a single HTTP request.
A team that builds a real-time system and borrows its dashboard from HTTP metrics will miss exactly what its users experience: not a call that went bad, but a call that never properly connected, or one whose quality quietly degraded with no alarm ever firing. Designing for network loss means accepting from the start that "connected" is a spectrum, not a boolean.
Why a real-time connection can't be retried like an HTTP request
Engineers new to real-time systems are often tempted to apply the same simple retry pattern that works for an HTTP request to a real-time session: if it fails, try again. That pattern is correct for a stateless request, because each attempt is independent of the last. A real-time session, however, carries state — who's in the room, what language they're speaking, what quality settings were chosen — and a naive retry that ignores this state effectively ejects the user from a room and drops them into a "fresh" one with the same name. The difference may be invisible from the outside, but from the inside it means losing everything that had accumulated in that session. Correct design treats state recovery as part of the contract from the start, not a feature bolted onto a simple retry afterward.
Testing something that only breaks under pressure
The practical problem with this kind of design is that most of these paths only activate under bad network conditions — exactly the conditions a normal development environment and CI pipeline never reproduce. If a team leaves these paths untested on the assumption that they "should just work," the first time they're really exercised will be in the middle of a real customer's call. The practical fix is to deliberately simulate bad conditions in a test environment: artificial latency, artificial packet loss, repeatedly cutting and restoring a connection, and watching whether the degradation ladder and the reconnection logic behave exactly as they were designed to on paper. This kind of testing is usually slower and more tedious than an ordinary unit test, and that is exactly why most teams skip it — until a real outage shows them why they shouldn't have.
Beyond simply finding bugs, this kind of testing moves the degradation ladder and reconnection logic from "theoretical assumption" to "observed behavior." A team that only knows on paper that its ladder should drop video quality first may later discover that, under a specific pattern of packet loss, the implemented logic doesn't actually trigger in that order at all — and that's something only real simulation reveals, not reading the code.
Iktar x’taqra

- Vector Search
- Retrieval
Vector retrieval in production: when a vector database earns its cost
A vector database is an architectural decision, not an automatic upgrade to search. If you can't name exactly what you're missing from plain text search, you probably don't need o…
8 min qari
- API Design
- Type Safety
Typed contracts between backend and client: why shared schemas cut integration bugs
"It worked in Postman" is a sentence nearly every team has said at least once, right before discovering the real problem was somewhere else. A typed contract between backend and c…
9 min qari
- Trust
- News Systems
Scoring news trustworthiness: designing a system that doesn't claim to be neutral
A trust label is an editorial decision encoded in software, not a measurement. Every design choice downstream of that fact — from data model to what you show the reader — depends…
9 min qari
Għandek xi ħaġa x'tibni?
Għidilna fuq xiex qed taħdem. Ngħidulek onestament jekk aħniex it-tim it-tajjeb għalih.
jew ibgħatilna email fuq hello@larsima.com
