How I Approach Production RAG Architecture

A production architecture guide to ingestion, hybrid retrieval, authorization, reranking, citations, evaluation, observability, cost and failure handling in enterprise RAG systems.

Romharshan Singh
Romharshan SinghSenior Solution Architect • AI & Cloud Mentor
15 August 20264 min read5 viewsUpdated 27 Sept 2026
How I Approach Production RAG Architecture
Key takeaways
  • A production RAG system is a distributed application; the LLM is only one component.
  • Document quality, metadata, authorization and retrieval quality matter more than an impressive demo prompt.
  • Filter by user permissions before context reaches the model.
  • Treat prompts, embeddings, chunking and model configuration as versioned production assets.
  • Measure retrieval quality, groundedness, latency and cost together.

Building a RAG demo is easy. Building a production RAG platform that users can trust is much harder. A proof of concept often looks like PDF → embeddings → vector database → LLM → answer. That is useful for learning, but it is not enough for an enterprise application.

A production platform must handle document updates, permissions, bad extraction, duplicate content, chunk quality, retrieval quality, hallucination risk, cost, latency, observability and auditability.

Production RAG architecture with ingestion and query pipelines
Separate ingestion from query-time orchestration, and make authorization, retrieval, reranking, citations and observability explicit components.

Start with the business risk, not the vector database

Before choosing an embedding model or vector store, ask what decisions users will make based on the answer. Searching a public FAQ has a different risk profile from answering questions about financial policy, contracts or operational procedures.

The higher the impact of a wrong answer, the stronger your evidence, authorization and verification controls must be.

Separate the ingestion and query pipelines

I treat ingestion and query-time orchestration as independently operated capabilities. Ingestion owns extraction, cleaning, chunking, metadata, embedding generation and index lifecycle. Query-time logic owns identity, retrieval, filtering, reranking, context assembly, generation and citations.

Document ingestion is more important than the model

Poor source data creates poor answers. Preserve metadata such as document ID, title, section, page, version, department, classification, effective date and source. This becomes essential for authorization, filtering and citations.

Chunking needs experimentation

A contract, FAQ, table, policy document and technical manual have different structures. Prefer semantic boundaries—headings, paragraphs and sections—over blindly cutting every N characters. Track the chunking strategy as a versioned configuration so quality changes can be explained later.

Retrieval should not depend only on vectors

Enterprise systems frequently benefit from hybrid retrieval: semantic search + keyword search + metadata filters, followed by reranking. Exact product codes, policy IDs and domain terminology may not rank reliably with vector similarity alone.

python
from dataclasses import dataclass

@dataclass
class RetrievedChunk:
    text: str
    document_id: str
    page: int
    score: float

async def retrieve(query: str, user_scope: list[str]):
    semantic = await vector_search(query, filters={"acl": user_scope})
    lexical = await keyword_search(query, filters={"acl": user_scope})

    candidates = merge_and_dedupe(semantic, lexical)
    return await reranker.rank(query, candidates, top_k=8)

Authorization must happen before generation

If a user is not authorized to access a document, that document should not be retrieved into the model context. Instructing the model not to reveal confidential data after sending that confidential content to the model is not a strong security boundary.

Citations should be part of the response contract

I prefer structured output that keeps the answer and evidence separate, so the UI can render verifiable sources.

json
{
  "answer": "International travel requires manager approval before booking.",
  "confidence": "high",
  "sources": [
    {
      "document": "Travel Policy v4.2",
      "page": 14,
      "section": "International Travel",
      "chunkId": "travel-4.2-14-03"
    }
  ]
}

The model must be allowed to say “I don't know”

A dangerous RAG system is one that always produces a confident answer. If retrieval quality is weak, a safer response is: “I could not find enough approved evidence to answer this reliably.” Refusal is a feature when evidence is insufficient.

Prompt design should enforce grounding

Your system prompt should require source-grounded claims, citations and explicit refusal when evidence is missing. This does not eliminate hallucination, but it creates predictable behavior that can be tested.

Use a clear orchestration boundary

python
async def answer_question(user, question: str):
    scope = await authorization_service.document_scope(user)
    chunks = await retrieve(question, scope)

    if not quality_gate.has_sufficient_evidence(chunks):
        return {
            "answer": "I could not find enough approved evidence to answer reliably.",
            "sources": [],
        }

    context = context_builder.build(chunks, token_budget=8_000)
    result = await llm.generate(
        prompt_version="rag-system-v12",
        question=question,
        context=context,
    )

    return citation_service.attach_sources(result, chunks)

Treat prompts and models as versioned software

Store prompt version, model version, embedding version, chunking strategy, retrieval parameters and reranker version with evaluation results. If answer quality changes after a deployment, you need to know what changed.

Cache carefully

Caching can reduce cost and latency, but the key must account for authorization scope and document version. A cached answer that ignores permissions can become a data leak; a cached answer that ignores document versions can become stale policy.

Deploy with explicit resource and health controls

yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: rag-orchestrator
spec:
  replicas: 3
  selector:
    matchLabels:
      app: rag-orchestrator
  template:
    metadata:
      labels:
        app: rag-orchestrator
    spec:
      containers:
        - name: api
          image: registry.example.com/rag-orchestrator:2.3.0
          ports:
            - containerPort: 8000
          readinessProbe:
            httpGet:
              path: /health/ready
              port: 8000
          resources:
            requests:
              cpu: "250m"
              memory: "512Mi"
            limits:
              cpu: "2"
              memory: "2Gi"

What I measure in production

LayerMetrics
Retrievalrecall@k, reranker scores, empty retrievals, filter exclusions
Generationgroundedness, refusal rate, citation coverage, model errors
Performanceretrieval latency, model latency, end-to-end p95/p99
Costinput/output tokens, embedding volume, cache hit rate
User valuehelpful votes, follow-up rate, source clicks, escalation rate

Production failure scenarios I test

  • vector database unavailable
  • LLM provider timeout or rate limit
  • document re-index in progress
  • retrieval returns unauthorized chunks
  • OCR/parser produces corrupt text
  • embedding model version changes
  • prompt change reduces citation quality
  • token budget exceeded

Final architecture principle

Evaluate a RAG system by whether it retrieves the right evidence, respects permissions, explains its sources, refuses safely and can be operated—not by how impressive one demo answer sounds.
Continue exploring
Was this article useful?

Your feedback helps prioritize deeper technical content.

Romharshan Singh
ABOUT THE AUTHOR

Romharshan Singh

Senior Solution Architect and Full Stack Technology Leader with 20+ years of enterprise engineering experience across AI, cloud, distributed systems, Java, Node.js, React, Angular, Kafka and Kubernetes.