Generative AI

Generative AI & RAG Architecture

I approach enterprise AI as a governed software system around a probabilistic model: retrieval quality, security, evaluation, cost, observability and fallback matter as much as the prompt.

My approach

A RAG prototype can be built quickly. A production RAG system is a different problem. It has to ingest changing documents, preserve permissions, retrieve relevant evidence, control context size, handle model failures, expose citations, protect sensitive data and prove that changes improve quality rather than simply changing behavior.

I treat the LLM as one component inside a larger architecture. The surrounding retrieval pipeline, data governance, evaluation harness, caching, observability and application workflows determine whether the system is trustworthy enough for real enterprise use.

The architecture should also make cost visible. Model size, context length, embedding volume, reranking, repeated retrieval and unnecessary real-time inference can turn a technically successful AI feature into an expensive platform.

Capability

What I focus on

I prefer to describe expertise through architecture decisions and production responsibilities rather than a list of tools.

01

RAG ingestion and knowledge preparation

Design parsing, normalization, chunking, metadata, document lineage and incremental ingestion so retrieval quality starts with clean, governable knowledge.

02

Embeddings and vector retrieval

Choose embedding models, vector stores and indexing strategies based on corpus size, latency, filtering, tenancy and operational requirements.

03

Retrieval quality and reranking

Use hybrid search, metadata filters, query rewriting and reranking when vector similarity alone is not enough to return trustworthy evidence.

04

LLM orchestration and tool use

Control prompts, structured outputs, function calling and workflows so the model performs bounded tasks instead of becoming an ungoverned application layer.

05

Security and permission-aware retrieval

Preserve source-system authorization, prevent cross-tenant leakage and treat retrieved context as untrusted input that can contain prompt-injection attempts.

06

Evaluation, guardrails and cost

Build repeatable offline and online evaluation, token/cost metrics, quality thresholds and safe fallbacks before scaling model usage.

Architecture

A production RAG flow I can reason about

The model comes near the end of the pipeline. Retrieval, permissions and evidence quality determine what information the model receives; evaluation determines whether the complete system is improving.

Source Systems
→
Ingestion + Chunking
→
Embeddings + Vector DB
→
Retrieve + Rerank
→
LLM / Tools
→
Citations + Evaluation

Evidence first

Optimize retrieval and provenance before attempting to solve poor answers with longer prompts.

Permission aware

Filter by the user’s authorized corpus before content reaches the model, not after an answer is generated.

Measurable quality

Maintain evaluation datasets and score retrieval, groundedness, correctness, latency and cost independently.

Production practice

How I approach production design

01

Design ingestion as a data pipeline

Documents arrive in different formats, change over time and often carry access-control metadata. I separate ingestion from query-time retrieval so parsing, chunking, embedding and indexing can be retried, audited and evolved independently.

Chunking is not a cosmetic parameter. It changes retrieval recall, context coherence and token consumption. I normally test chunk strategies against representative questions instead of selecting one global size by convention.

  • Document identity and version
  • Chunk lineage
  • Access metadata
  • Incremental updates
  • Embedding version
  • Delete / re-index workflow
02

Treat retrieval as an information-retrieval problem

Semantic similarity is useful but not sufficient for every corpus. Product codes, policy numbers, names and exact terminology may benefit from keyword or hybrid search. Metadata filters can narrow the candidate set before ranking.

For high-value workflows, reranking a smaller candidate set can materially improve relevance. I measure retrieval independently from answer generation so the team can see whether a miss came from search or from the model.

03

Bound the model with contracts

LLMs are probabilistic. Enterprise integrations should therefore request structured outputs, validate them and keep irreversible actions behind deterministic authorization and business rules.

Tool calling is most useful when the model decides among explicitly allowed capabilities. The tool layer should enforce schema, authorization, timeout and audit behavior regardless of what the model requests.

  • Structured output validation
  • Tool allow-list
  • Timeouts
  • Business-rule enforcement
  • Audit trail
  • Human approval for high-risk actions
04

Build evaluation before optimization

Without a fixed evaluation set, teams can change prompts, models and retrieval settings without knowing whether quality improved. I want a representative set of questions, expected evidence and failure categories before tuning.

Cost optimization then becomes safer. We can compare smaller models, caching, context trimming or selective reranking while verifying that quality remains within an acceptable range.

Decision framework

Questions I want answered before approving the design

Do we need RAG at all?

Use RAG when answers require changing or private knowledge that should remain outside model weights. For static transformations or extraction, a simpler prompt or deterministic pipeline may be better.

Which vector database should we use?

Choose based on scale, filtering, tenancy, latency, operational model and existing platform skills. FAISS can be excellent for embedded/local use; managed or distributed stores help when operational scale and filtering become primary concerns.

Should every query use the largest model?

No. Route by task complexity. Smaller or cheaper models can handle classification, rewriting or straightforward answers while larger models are reserved for cases where they add measurable value.

How do we prevent hallucination?

You cannot guarantee zero hallucination. Improve retrieval, require evidence, constrain output, detect unsupported answers, provide citations and fall back when confidence or evidence is insufficient.

How will we know a new prompt is better?

Run it against a versioned evaluation set and compare correctness, groundedness, retrieval quality, latency and cost before promoting the change.

Reliability

Failure, scale and operational reality

Risk

Relevant document is not retrieved

Architecture response

Measure recall, inspect chunking and metadata, consider hybrid retrieval/query rewriting, and separate retrieval evaluation from generation evaluation.

Risk

Cross-user or cross-tenant data leakage

Architecture response

Apply authorization filters before retrieval, maintain tenant-aware indexes or namespaces where appropriate, and test negative permission cases continuously.

Risk

Prompt injection through documents

Architecture response

Treat retrieved content as untrusted data, separate system instructions, restrict tool capabilities and never let retrieved text override authorization or policy controls.

Risk

Model or provider outage

Architecture response

Use bounded timeouts, provider abstraction where justified, cached or deterministic fallbacks, graceful degradation and clear user messaging.

Risk

Cost grows faster than usage

Architecture response

Track tokens and cost per feature/tenant, trim context, cache stable results, batch embeddings and route tasks to the smallest model that meets the quality target.

Security

Security by architecture

  • Preserve document-level and tenant-level authorization through the retrieval path.
  • Classify prompts, retrieved context and model outputs for sensitive data exposure.
  • Validate tool calls server-side and keep the LLM outside the authorization boundary.
  • Protect against prompt injection, data exfiltration and malicious uploaded content.
  • Audit model, prompt, retrieval and tool versions for high-value workflows.

Observability

Operate what we design

  • Trace the complete request: query rewrite, retrieval, reranking, prompt, model and tools.
  • Measure retrieval recall/precision on evaluation sets in addition to end-answer scores.
  • Track model latency, token usage, cache hit rate and cost by workflow.
  • Capture groundedness/citation failures and user feedback as quality signals.
  • Monitor ingestion lag, failed documents, stale embeddings and index health.

Continue reading

Related architecture guides

FAQ

Frequently asked questions

What is RAG architecture?

Retrieval-Augmented Generation combines information retrieval with an LLM. The system retrieves relevant authorized evidence from an external knowledge source and provides that evidence to the model so answers can use current or private information without retraining the model.

Do I always need a vector database for RAG?

No. Small corpora can use local vector indexes, and some use cases work well with keyword or hybrid search. The database choice should follow scale, filtering, tenancy, latency and operational requirements.

What is the biggest difference between a RAG demo and production RAG?

Production systems need data ingestion, permissions, evaluation, observability, cost controls, failure handling, provenance, prompt/version management and security. The model call is only one part of the platform.

How should RAG quality be measured?

Measure retrieval relevance/recall and answer groundedness/correctness separately. Also track citation quality, latency, cost and failure rates using a versioned evaluation dataset that represents real user questions.

How do you secure enterprise RAG?

Filter retrieval using the caller’s permissions, isolate tenants where appropriate, validate tool calls, protect secrets, treat retrieved text as untrusted, test for prompt injection and retain audit evidence for sensitive workflows.

About the author

Romharshan Singh

Senior Solution Architect • AI & Cloud Mentor

I write about architecture from a production perspective: how systems fail, how design decisions affect cost and operability, and how teams can turn technology choices into maintainable enterprise platforms.