← All articles
Engineering10 min read

Vector Databases for RAG: Practical Architecture Choices

Vector databases for RAG: practical architecture choices — managed vs self-hosted, hybrid search, tenancy, and ops patterns QuantaloomAI ships in production.

Vector Databases for RAG: Practical Architecture Choices

Vector databases for RAG are not a checkbox — they are an architecture decision that shapes latency, tenancy, cost, and how badly your assistant fails when documents change. Teams that pick a vendor from a Twitter thread often rediscover basic requirements later: filtered search, deletions that actually delete, and auditability.

QuantaloomAI designs RAG stacks around operational constraints first, model fashion second. Below are practical architecture choices we make with mid-market and enterprise product teams.

Vector databases for RAG: what you are really buying

A production RAG path includes chunking, embeddings, indexing, retrieval, reranking, and grounding in the generator. The vector store is one component. Treating it as the whole system creates brittle apps. See LLM application development: RAG and fine-tuning for the broader map.

You need the store to support:

  • Metadata filters (tenant, ACL, document type, freshness)
  • Hybrid search (keyword + vector) for IDs, codes, and rare terms
  • Efficient upserts and deletes when policies change
  • Observability of recall quality over time

Managed vs self-hosted

Managed wins when your team is small, data residency options match requirements, and you want less ops burden. Self-hosted / VPC wins for strict residency, air-gapped environments, or when query patterns need deep customization.

Do not self-host for prestige. Do self-host when compliance or cost at scale demands it — often alongside compliance-ready AI requirements.

Architecture patterns that work

Pattern A: pgvector-in-primary for early products

Good for MVPs with modest corpora and strong desire for transactional consistency. Watch index build times and vacuum behavior as you grow.

Pattern B: dedicated vector service + object/doc store

Scale retrieval independently; keep canonical documents in blob/SQL. Best default for multi-product platforms and SaaS platforms.

Pattern C: hybrid search with rerankers

Combine BM25/keyword with vectors, then rerank top-k. Dramatically improves exact SKU, policy ID, and proper-noun queries common in ERP and healthcare corpora.

QuantaloomAI data engineering pipelines own chunking and embedding jobs so the vector index is a projection of governed source data — not a shadow CMS.

Tenancy and security

Never rely on "the model will ignore other tenants' chunks." Enforce tenant_id (and finer ACLs) in every query filter. Encrypt at rest, rotate embedding keys if you version models, and log retrieval sets for audit. Customer-facing systems that skip filters will eventually leak — a failure mode we design against when reducing hallucinations and building trust UX.

Chunking and metadata beat fancy indexes

Bad chunks poison every database. Prefer structure-aware splitting (headings, sections), attach source URLs/versions, and store last-updated timestamps. Re-embed on model change with a dual-read period. Measure recall@k on a labeled set — wire it into evaluation pipelines.

Latency and cost knobs

  • Cache frequent queries for stable corpora
  • Cap top-k; rerank fewer candidates
  • Separate hot vs cold indexes
  • Track cost per successful grounded answer, not only storage GB

Observability should include empty-result rate and citation miss rate — see building observability into production LLM applications. Also monitor embedding job lag: if documents publish to the CMS but take twelve hours to appear in the index, users will blame "the AI" for stale answers. Publish freshness SLAs next to retrieval SLAs so product and ops share one definition of healthy.

Decision checklist

1. Corpus size and update frequency? 2. Residency and VPC requirements? 3. Need hybrid search on day one? 4. Multi-tenant ACL complexity? 5. Team ops capacity?

Answer those before comparing QPS charts. QuantaloomAI AI product development engagements prototype retrieval quality on your documents — not synthetic Wikipedia dumps — before locking a vendor.

Migration and dual-write periods

When you change embedding models or chunking strategies, run dual indexes and compare recall on the golden set before cutover. Abrupt reindexes without shadow evaluation are a common cause of sudden hallucination spikes. Budget time for backfills; large corpora are not overnight jobs.

Document a rollback: keep the previous index queryable until the new path clears eval gates for a full week of online sampling. Vector databases for RAG are operational systems — treat releases like database migrations, not config toggles.

When not to over-invest

If your corpus is a handful of stable PDFs and traffic is low, a simple managed index may be enough. Complexity should follow risk and scale. Mid-market teams often waste quarters evaluating exotic ANN algorithms while chunk quality remains poor. Fix content and metadata first; then tune the store. Revisit architecture when corpus size, tenancy complexity, or query QPS crosses thresholds you wrote down at kickoff — not when a conference talk makes you anxious.

Related builds such as NeuralDesk and knowledge-heavy platforms succeed when the vector layer is boring, filtered, and measured. Pair retrieval work with workflow automation so answers can become actions without inventing a second integration story. That is the practical standard for vector databases for RAG in production products.


*Written by Sharjeel Ahmed, QuantaloomAI. Book a briefing to review your RAG architecture choices.*

Building something worth shipping?

We take on a small number of AI product engagements. Tell us what you are building — we reply within 48 hours.