- A production RAG system is a distributed application; the LLM is only one component.
- Document quality, metadata, authorization and retrieval quality matter more than an impressive demo prompt.
- Filter by user permissions before context reaches the model.
- Treat prompts, embeddings, chunking and model configuration as versioned production assets.
- Measure retrieval quality, groundedness, latency and cost together.
Building a RAG demo is easy. Building a production RAG platform that users can trust is much harder. A proof of concept often looks like PDF → embeddings → vector database → LLM → answer. That is useful for learning, but it is not enough for an enterprise application.
A production platform must handle document updates, permissions, bad extraction, duplicate content, chunk quality, retrieval quality, hallucination risk, cost, latency, observability and auditability.
Start with the business risk, not the vector database
Before choosing an embedding model or vector store, ask what decisions users will make based on the answer. Searching a public FAQ has a different risk profile from answering questions about financial policy, contracts or operational procedures.
The higher the impact of a wrong answer, the stronger your evidence, authorization and verification controls must be.
Separate the ingestion and query pipelines
I treat ingestion and query-time orchestration as independently operated capabilities. Ingestion owns extraction, cleaning, chunking, metadata, embedding generation and index lifecycle. Query-time logic owns identity, retrieval, filtering, reranking, context assembly, generation and citations.
Document ingestion is more important than the model
Poor source data creates poor answers. Preserve metadata such as document ID, title, section, page, version, department, classification, effective date and source. This becomes essential for authorization, filtering and citations.
Chunking needs experimentation
A contract, FAQ, table, policy document and technical manual have different structures. Prefer semantic boundaries—headings, paragraphs and sections—over blindly cutting every N characters. Track the chunking strategy as a versioned configuration so quality changes can be explained later.
Retrieval should not depend only on vectors
Enterprise systems frequently benefit from hybrid retrieval: semantic search + keyword search + metadata filters, followed by reranking. Exact product codes, policy IDs and domain terminology may not rank reliably with vector similarity alone.
from dataclasses import dataclass
@dataclass
class RetrievedChunk:
text: str
document_id: str
page: int
score: float
async def retrieve(query: str, user_scope: list[str]):
semantic = await vector_search(query, filters={"acl": user_scope})
lexical = await keyword_search(query, filters={"acl": user_scope})
candidates = merge_and_dedupe(semantic, lexical)
return await reranker.rank(query, candidates, top_k=8)Authorization must happen before generation
If a user is not authorized to access a document, that document should not be retrieved into the model context. Instructing the model not to reveal confidential data after sending that confidential content to the model is not a strong security boundary.
Citations should be part of the response contract
I prefer structured output that keeps the answer and evidence separate, so the UI can render verifiable sources.
{
"answer": "International travel requires manager approval before booking.",
"confidence": "high",
"sources": [
{
"document": "Travel Policy v4.2",
"page": 14,
"section": "International Travel",
"chunkId": "travel-4.2-14-03"
}
]
}The model must be allowed to say “I don't know”
A dangerous RAG system is one that always produces a confident answer. If retrieval quality is weak, a safer response is: “I could not find enough approved evidence to answer this reliably.” Refusal is a feature when evidence is insufficient.
Prompt design should enforce grounding
Your system prompt should require source-grounded claims, citations and explicit refusal when evidence is missing. This does not eliminate hallucination, but it creates predictable behavior that can be tested.
Use a clear orchestration boundary
async def answer_question(user, question: str):
scope = await authorization_service.document_scope(user)
chunks = await retrieve(question, scope)
if not quality_gate.has_sufficient_evidence(chunks):
return {
"answer": "I could not find enough approved evidence to answer reliably.",
"sources": [],
}
context = context_builder.build(chunks, token_budget=8_000)
result = await llm.generate(
prompt_version="rag-system-v12",
question=question,
context=context,
)
return citation_service.attach_sources(result, chunks)Treat prompts and models as versioned software
Store prompt version, model version, embedding version, chunking strategy, retrieval parameters and reranker version with evaluation results. If answer quality changes after a deployment, you need to know what changed.
Cache carefully
Caching can reduce cost and latency, but the key must account for authorization scope and document version. A cached answer that ignores permissions can become a data leak; a cached answer that ignores document versions can become stale policy.
Deploy with explicit resource and health controls
apiVersion: apps/v1
kind: Deployment
metadata:
name: rag-orchestrator
spec:
replicas: 3
selector:
matchLabels:
app: rag-orchestrator
template:
metadata:
labels:
app: rag-orchestrator
spec:
containers:
- name: api
image: registry.example.com/rag-orchestrator:2.3.0
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /health/ready
port: 8000
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "2"
memory: "2Gi"What I measure in production
| Layer | Metrics |
|---|---|
| Retrieval | recall@k, reranker scores, empty retrievals, filter exclusions |
| Generation | groundedness, refusal rate, citation coverage, model errors |
| Performance | retrieval latency, model latency, end-to-end p95/p99 |
| Cost | input/output tokens, embedding volume, cache hit rate |
| User value | helpful votes, follow-up rate, source clicks, escalation rate |
Production failure scenarios I test
- vector database unavailable
- LLM provider timeout or rate limit
- document re-index in progress
- retrieval returns unauthorized chunks
- OCR/parser produces corrupt text
- embedding model version changes
- prompt change reduces citation quality
- token budget exceeded
Final architecture principle
Evaluate a RAG system by whether it retrieves the right evidence, respects permissions, explains its sources, refuses safely and can be operated—not by how impressive one demo answer sounds.
Your feedback helps prioritize deeper technical content.






