- Vector Search
- Retrieval
- Infrastructure
Vector retrieval in production: when a vector database earns its cost
A vector database is an architectural decision, not an automatic upgrade to search. If you can't name exactly what you're missing from plain text search, you probably don't need one yet.
8 min lexim

Adding a vector database to a system is one of those decisions that's easy to make for the wrong reason: because everyone else is talking about it, because it sounds smarter, or because a library wires it up in two lines of code. But vector retrieval is a tool for a specific problem — finding things that are semantically close even when the words aren't the same — and like any specialized tool, it carries real operational cost: another index to maintain, another latency hop in the request path, and a genuinely new failure mode called embedding drift that plain text search never has to deal with. The right question isn't "is vector retrieval good?" It's "what exactly is my problem, and which tool solves it at the lowest cost?"
Semantic search versus full-text search
Full-text search (what Postgres or Elasticsearch give you natively) works on word matching, weighted by frequency and rarity, with rules like stemming and limited synonym expansion. That's excellent when the user knows roughly which word they're looking for. Semantic search works on closeness in vector space: two sentences can share no words at all and still be very close in meaning, and vector retrieval finds exactly that closeness.
The difference shows up concretely: a user searching for "quarterly revenue growth" against a document that says "three-month increase in sales" — plain full-text search struggles to connect the two without a lot of synonym work. Semantic search finds it naturally. But semantic search is weaker exactly where literal word precision matters — proper names, product codes, exact legal phrasing — because it's measuring meaning, not character strings.
Hybrid retrieval: not a choice, a combination
The practical consequence of that difference is that most real systems shouldn't pick one over the other; they should run both and merge the results. The common pattern: both searches run in parallel, each returns its own list of candidate results, and a fusion layer (often RRF-based, or a lightweight re-ranking model) merges them. That's more complexity than a single engine, but in practice it's a trade worth making, because the two approaches fail in different places — where word matching fails, meaning usually holds, and the reverse is often true too.
Before building that combination, it's worth honestly asking whether you need both at all. If your users' queries are mostly short and precise (product name searches, order IDs), full-text search alone is probably sufficient. Hybrid retrieval earns its architectural cost when queries are natural and open-ended — exactly what you'd see in a news reader or a content recommendation engine.
The fusion logic is usually about this simple — each result is scored by its rank in each list, not by that engine's raw score, because the two engines' scores aren't directly comparable:
score(doc) = 1 / (k + rank_fulltext(doc)) + 1 / (k + rank_semantic(doc))
This formula (a version of RRF) is deliberately simple: it doesn't ask you to bring two different engines' scores onto one shared scale, it just counts each result's rank within its own list. For most systems it's a better starting point than a trained re-ranking model, precisely because it needs no labeled data at all.
Embedding drift and re-indexing
One problem that doesn't exist in plain text search at all is embedding drift. When the model you use to generate vectors changes — through a version upgrade or a provider switch — the new vectors don't mean the same thing in the same space as the old ones. Comparing a vector built with model A against one built with model B returns a meaningless result, even though both numbers look perfectly valid. This isn't a failure that announces itself with an error; it shows up as a quiet degradation in result quality that's hard to diagnose.
The consequence is that swapping an embedding model isn't a simple upgrade — it's a full migration. You need to plan a complete index rebuild, not just for new data but for the entire collection, and you need a strategy for the transition window — either keeping both indexes live until the rebuild finishes, or accepting a short period where search quality is temporarily degraded. A team that hasn't planned for this gets caught off guard the day it's forced to switch models — and that day arrives, because embedding models become obsolete the same way any other model does.
The latency budget of a retrieval hop
Every retrieval hop — to a vector database or to any other external service — adds a network round trip to the request path, and that hop has to fit inside the overall latency budget. If a feature is supposed to respond in under a second and vector retrieval alone takes a hundred to two hundred milliseconds, that hop leaves very little room to maneuver for the language model call that usually follows it.
The practical rule: measure retrieval latency before adding the language model call, not after. If vector search is already slow, stacking a generation step on top doesn't just double the slowness — it compounds it in the worst case. Caching results for frequent queries, capping candidate list size, and running independent searches in parallel are the first three simple tools for keeping that budget under control.
Monitoring retrieval quality after deployment
A vector index that's up and running, unlike many other services, can quietly get worse without a single error. New content gets added, the topical distribution of the data shifts, and result quality for specific queries degrades without triggering any alarm, because from a technical standpoint the service is perfectly "healthy" — it's just returning less relevant results. Monitoring retrieval quality can't rely on latency and error rate alone; it needs a separate evaluation set: a collection of real queries paired with acceptable results, run periodically against the live production index, reporting metrics like recall@k.
This evaluation needs to be independent of the development team and run against data that resembles real user queries, not the clean, simple queries an engineer types during manual testing. Without this layer, the only way to find out retrieval quality has degraded is a user complaint — and by then it's already too late.
A useful practice is running this evaluation whenever a large batch of new content lands, not only on a fixed calendar schedule. Adding an entirely new category of content to the collection, even without touching a line of code, can shift the result distribution for old queries, because the vector space now has new neighbors that weren't there before.
Operating Qdrant, in general terms
Without going into the specifics of any particular deployment, operating a vector database like Qdrant carries a few recurring operational concerns that differ from a relational database: index parameter choices (like HNSW settings) trade off search speed, memory use and accuracy, and need to be tuned against real data rather than left at defaults; metadata filters (by date or category, for example) need to be designed into the collection from the start, because adding a filter after the index has grown large carries a rebuild cost; and collection growth needs monitoring, because unlike many relational databases, vector retrieval doesn't degrade gradually as data volume grows — it tends to fall off a cliff at specific memory thresholds instead.
When not to use a vector database
The most important part of this decision might be knowing when to skip it entirely. If your dataset is small (a few thousand documents), in-memory vector search or even a simple vector-extension filter on Postgres gives you the same accuracy without a separate service to run. If your users' queries are fundamentally precise (codes, IDs, exact legal phrasing), semantic search adds nothing but complexity. And if your team isn't ready to maintain a separate service with its own lifecycle — indexing, rebuilding, memory monitoring — the operational cost usually outweighs whatever benefit it looked like it would bring in the first week.
A system like a news reader that keeps Qdrant alongside Postgres and Redis needs it for exactly this reason: its problem — finding semantically related stories across a large and varied body of content — is precisely the problem full-text search alone doesn't solve. If that isn't your problem, a vector database is just another service to keep alive, without the payoff.
Më shumë për të lexuar

- API Design
- Type Safety
Typed contracts between backend and client: why shared schemas cut integration bugs
"It worked in Postman" is a sentence nearly every team has said at least once, right before discovering the real problem was somewhere else. A typed contract between backend and c…
9 min lexim
- Trust
- News Systems
Scoring news trustworthiness: designing a system that doesn't claim to be neutral
A trust label is an editorial decision encoded in software, not a measurement. Every design choice downstream of that fact — from data model to what you show the reader — depends…
9 min lexim
- AI Infrastructure
- Model Routing
One model is a single point of failure: routing between hosted and local AI models
Wiring a feature to a single model provider looks like the simple choice — until that provider gets slow, changes price, or fails on one kind of input. Routing across models isn't…
8 min lexim
Keni diçka për të ndërtuar?
Na tregoni se çfarë keni në dorë. Ju themi ndershmërisht nëse jemi ekipi i duhur për të.
ose na shkruani në hello@larsima.com
