My approach
Frameworks make endpoints easy to create. Production backend engineering is about everything that happens around the endpoint: validation, authorization, transaction boundaries, idempotency, database behavior, caching, messaging, dependency timeouts and telemetry.
I prefer domain and application logic that can be understood independently from transport or persistence details. This keeps the codebase easier to test and makes future framework, database or integration changes less disruptive.
Performance work also starts with measurement. Before adding caches, queues or replicas, I want to identify whether the real constraint is CPU, I/O, database locking, connection pools, downstream latency or inefficient application behavior.
Capability
What I focus on
I prefer to describe expertise through architecture decisions and production responsibilities rather than a list of tools.
Service and domain design
Keep business rules explicit, define module/service boundaries and avoid controllers or ORM entities becoming the architecture.
API design
Use stable resource/command contracts, validation, versioning, consistent errors and idempotency where repeated requests are possible.
Database and transaction design
Choose relational, document and cache technologies based on access patterns, consistency, query requirements and operational ownership.
Caching
Use Redis or application caching where measured value exists, with explicit invalidation, TTL, consistency and failure behavior.
Messaging and asynchronous work
Move long-running or decoupled work to queues/events when that improves user latency and reliability, while keeping processing repeatable.
Resilience and observability
Bound external calls with timeouts, retries and circuit breakers; expose metrics and traces that show where requests actually spend time.
Architecture
A backend structure that separates responsibilities
I prefer request flow where transport concerns, application use cases, domain rules and infrastructure adapters are visible boundaries. This keeps the core behavior testable and reduces framework coupling.
Dependency direction
Business rules should not need to know HTTP, ORM or vendor SDK details.
Transaction clarity
Keep transaction boundaries visible around business operations; do not let implicit persistence behavior define them.
Measured optimization
Profile queries, pools, caches and downstream latency before introducing additional infrastructure.
Production practice
How I approach production design
Keep controllers thin and use cases explicit
Controllers should translate transport into application inputs and outputs. When business rules, persistence and external calls accumulate in controllers, testing becomes difficult and the same logic is duplicated across HTTP, jobs or consumers.
I organize core use cases around business actions so they can be invoked from an API, a message consumer or a scheduled workflow without rewriting the business behavior.
Design database access from query and consistency needs
Relational databases are excellent when transactions, constraints and rich querying matter. Document databases can be valuable when aggregate-shaped data and flexible schema are primary. Redis is useful for low-latency ephemeral or cached state, not as a universal replacement for durable storage.
The data model should follow access patterns but still preserve business integrity. Denormalization is an optimization with update consequences, not a free performance trick.
- Indexes from real queries
- Connection pool sizing
- Lock analysis
- Isolation level
- Read replicas
- Migration strategy
Treat external calls as unreliable
Every network dependency needs a timeout. Retries should be limited to operations that are safe and errors that are likely to be transient. Idempotency matters when clients, gateways or workers can repeat requests.
For slow workflows, asynchronous processing can protect user-facing latency, but it also introduces state transitions, retries and operational queues that must be designed explicitly.
Make performance observable
An API response time is the sum of work across application code, databases, caches and downstream services. Distributed tracing and dependency metrics help identify where the budget is spent.
I also track pool saturation, query latency, queue depth and event-loop/GC behavior depending on the runtime. Scaling the API tier does not solve a saturated database.
Decision framework
Questions I want answered before approving the design
Spring Boot or NestJS?
Choose based on team skills, ecosystem, workload, platform standards and operational requirements. Both can support strong enterprise services; architecture quality matters more than language preference.
SQL or NoSQL?
Start from transaction, query, consistency, schema and scale requirements. Many systems need more than one data technology, but each additional store adds operational and consistency cost.
Should this endpoint be synchronous?
Keep it synchronous when the caller genuinely needs the completed outcome. Move long-running or decoupled work to asynchronous processing when it improves responsiveness and resilience.
Should we add Redis?
Add a cache only for a measured latency or load problem and define invalidation, TTL, stampede behavior and what happens when Redis is unavailable.
Where should validation happen?
Validate transport shape at the boundary and enforce business invariants in the domain/application layer. Database constraints remain valuable as the final integrity boundary.
Reliability
Failure, scale and operational reality
Risk
Database connection pool exhaustion
Architecture response
Measure active/waiting connections, query duration and transaction length. Bound concurrency and fix slow queries before blindly increasing pool size.
Risk
Cache stampede or stale data
Architecture response
Use TTL jitter, request coalescing/locking where appropriate and clear consistency expectations. The source of truth must remain recoverable when the cache fails.
Risk
Downstream API becomes slow
Architecture response
Apply short bounded timeouts, circuit breakers and selective retries. Protect worker/thread/event-loop capacity so one dependency does not consume the service.
Risk
Duplicate command/request
Architecture response
Use idempotency keys, natural business keys or conditional state transitions for operations with side effects such as payment, provisioning or publishing.
Security
Security by architecture
- Validate and normalize input at trust boundaries; do not rely on UI validation.
- Separate authentication from fine-grained business authorization.
- Use parameterized queries/ORM safely and protect dynamic query construction.
- Keep secrets outside source code and rotate privileged credentials.
- Audit security-sensitive state changes and administrative actions.
Observability
Operate what we design
- Track endpoint latency by route and status, including percentiles.
- Measure database query latency, pool utilization and error classes.
- Trace downstream HTTP, messaging and cache calls with correlation IDs.
- Monitor runtime-specific signals such as event-loop delay or JVM GC where relevant.
- Publish business counters for critical workflows, not only infrastructure metrics.
Continue reading
Related architecture guides
FAQ
Frequently asked questions
Is Spring Boot better than NestJS for enterprise backend systems?
Neither is universally better. Spring Boot offers a mature Java ecosystem and is common in large enterprises; NestJS provides strong TypeScript structure and fits Node.js teams well. Team capability, platform standards, workload and operational requirements should drive the decision.
When should a backend use Redis?
Use Redis when low-latency cache, ephemeral state, distributed coordination or specific data structures provide measurable value. Define TTL, invalidation, durability expectations and failure behavior before adding it.
How should backend APIs handle retries?
Only retry transient failures and safe/idempotent operations, with bounded attempts, exponential backoff and jitter. Avoid retrying at multiple layers because attempts multiply quickly during outages.
What makes an API production-ready?
Clear contracts, validation, authentication/authorization, timeouts, consistent error behavior, observability, capacity limits, dependency resilience, versioning and safe deployment all matter.
How do you choose between SQL and MongoDB?
Choose from the access and consistency model. SQL is strong for transactions, constraints and relational querying; document databases are useful for aggregate-shaped flexible data. Operational skills and migration/backup requirements also matter.