- AI Infrastructure
- Model Routing
- Reliability
One model is a single point of failure: routing between hosted and local AI models
Wiring a feature to a single model provider looks like the simple choice — until that provider gets slow, changes price, or fails on one kind of input. Routing across models isn't a cleverness upgrade; it's a decision made per feature, per request.
8 min lukuaika

The first time you build a feature with a direct call to one provider, it's the fastest path. The first time that provider gets slow, changes its pricing, or fails on a particular kind of input, you find out that simplicity was a hidden dependency: the entire feature was tied to the health of a single service. Routing across models — Gemini, OpenAI, and local models on Ollama or HuggingFace — isn't a way to be "smarter." It's a response to the fact that no single model or provider is the right choice for every request. The right question isn't "which model is better?" It's "for this specific request, under these cost, latency and privacy constraints, which path is correct?"
Routing per request, not per service
The common mistake is defining routing at the service level: "this feature uses Gemini, that one uses OpenAI." That's simple, but it just moves the single-provider problem down one layer. Real routing happens at the request level: each call should decide which model to invoke based on the properties of that particular request — data sensitivity, task difficulty, latency budget, acceptable cost — not based on which feature made the call.
That means the routing layer needs to be independent of business logic and take its inputs from the request itself, not from a fixed per-feature configuration. A short, simple request from a sensitive feature might need to go to a local model; a complex request from that same feature might need a stronger hosted model. The decision isn't tied to the feature; it's tied to the request.
Privacy-driven routing: data that never leaves
The strongest reason to keep a local path alive — with Ollama or HuggingFace models run on your own infrastructure — isn't that it's cheaper. It's that some data should never cross the boundary of your infrastructure at all. If a feature works with content that a privacy policy or legal requirement blocks from reaching an external API, routing needs to enforce that at the design level, not at the level of a team guideline.
The right design encodes this decision directly in the routing layer: requests carrying a specific sensitivity tag never go to an external provider, not even as a fallback option. This has to be a hard rule, not a preference that gets bypassed under pressure — for instance when the local model is slow or unavailable. If your fallback logic allows that boundary to be crossed, the boundary doesn't actually exist.
The economics of local versus cloud inference
Comparing the cost of a local model against a hosted one is rarely straightforward, because the two have fundamentally different cost structures. A hosted model has a variable, per-token cost: no load, no cost; heavy load, cost rises linearly (or worse). A local model has a mostly fixed cost: hardware (or GPU rental) has to be purchased or reserved regardless of how much it's actually used, and that cost keeps accruing as long as the machine stays on.
That means the local path only makes economic sense when load is predictable and steady enough to justify the fixed hardware cost — a break-even point below which paying per token to an external provider is cheaper, even accounting for the privacy benefit. The local path also carries hidden costs that don't show up in an external API's per-token price: engineering time for maintenance, model upgrades, and scaling when load outgrows the hardware you have on hand. A team that only compares the per-token price of the two paths and ignores the operational cost of maintaining local infrastructure usually ends up seeing local inference as cheaper than it actually is.
The practical rule: reserve the local path for tasks that are both sensitive and predictable enough in volume to plan capacity around; for low-frequency or irregular-load tasks, a hosted model's variable cost almost always wins.
Fallback: the second path has to actually work
Routing without fallback is just an extra selection layer, not an increase in reliability. When the first provider errors out, is slow, or returns an invalid response, the system needs to fail over to a second path without human intervention — but that fallback has to have been tested to the same quality bar as the primary path, not just accepted as "anything is better than an error." A fallback that suddenly returns a different response format or switches language creates a new problem the user sees, even if the system looks "up" from a monitoring dashboard.
The practical point here: test that your fallback path actually works, not just that it exists. A second path that hasn't been invoked in months is as dangerous as having no fallback at all, because its reliability has never been put to the test.
Disagreement between models: signal, not noise
When two different models run on the same input and return different answers, the temptation is to treat that as noise and pick one as "correct." But disagreement between models is usually a valuable signal about the uncertainty of the task itself. If two strong models disagree sharply on a classification, that input is probably borderline and error-prone — regardless of which model, or even a third model, is asked.
Systems that run multiple models on sensitive tasks — not as a replacement for each other, but in parallel — can use that disagreement as a trigger for human review: when the models agree, more confidence can be placed in the automated output; when they disagree, that case should go into a review queue, rather than one model being declared the arbitrary winner.
Cross-model evaluation
Evaluating one model in isolation isn't enough once your system routes across several in production. Your golden set needs to run against every possible path — Gemini, OpenAI, and whichever local model is in rotation — not just the path that was used during development. A model that excels at one type of task may be noticeably weaker on another, and the only way to find that out is to run both against the same real evaluation set, rather than trusting a provider's public benchmarks.
This evaluation needs to be recurring, not one-time, because providers update the models underneath their APIs without notice. A model that performed well on your evaluation set last month is not guaranteed to perform the same way this month.
A per-feature decision framework
In the end, this decision needs to be documented explicitly per feature, not for the system as a whole: data sensitivity (can this even go to an external API at all?), latency budget (is a local call faster or slower than a hosted one?), acceptable cost per request, and fault tolerance (if the primary path fails, what happens to the user?). Those four axes give a different answer for every feature, and no blanket rule ("always use Gemini") gets that difference right.
A system with both hosted and local models available needs this framework not as a luxury but as a condition of staying up — because any single provider, in the end, is a single point of failure.
A small example: routing for a hypothetical feature
To keep this framework from staying abstract, consider a hypothetical feature: automatically summarizing a news story's text for display in a mobile feed. The data entering this feature is public (the already-published story text), so there's no hard privacy constraint. But the latency budget is tight, because the summary needs to be ready by the time the user scrolls to it, and request volume is high and fairly steady, since each new story gets summarized once, not once per view.
That combination — non-sensitive data, high and predictable load, a tight latency budget — is exactly the profile that makes a local path economical: the fixed hardware cost gets justified by the high request volume, and not depending on an external network makes latency more predictable. Now picture the same framework applied to a different feature in the same system — say, answering a user's complex question about a story's background context, where repetition is low and reasoning quality matters more than speed. The same framework, the same four axes, gives a completely different answer: a stronger hosted model, even at a higher per-request cost. The gap between these two hypothetical features within the same system is exactly why routing makes sense at the feature level, not the system level.
Lisää luettavaa

- Vector Search
- Retrieval
Vector retrieval in production: when a vector database earns its cost
A vector database is an architectural decision, not an automatic upgrade to search. If you can't name exactly what you're missing from plain text search, you probably don't need o…
8 min lukuaika
- API Design
- Type Safety
Typed contracts between backend and client: why shared schemas cut integration bugs
"It worked in Postman" is a sentence nearly every team has said at least once, right before discovering the real problem was somewhere else. A typed contract between backend and c…
9 min lukuaika
- Trust
- News Systems
Scoring news trustworthiness: designing a system that doesn't claim to be neutral
A trust label is an editorial decision encoded in software, not a measurement. Every design choice downstream of that fact — from data model to what you show the reader — depends…
9 min lukuaika
Onko sinulla jotain rakennettavaa?
Kerro mitä työstät. Sanomme rehellisesti, olemmeko siihen oikea tiimi.
tai lähetä sähköpostia osoitteeseen hello@larsima.com
