My approach
For me, solution architecture starts with the problem, not the framework. Before choosing microservices, Kafka, Kubernetes, a database or an AI platform, I want to understand the business capability, the expected scale, the critical user journeys, the regulatory constraints, the recovery objectives and the way the system will actually be operated.
A good architecture should make delivery easier rather than adding ceremony. I use HLDs, LLDs, architecture decision records and non-functional requirements to make important choices visible: boundaries, data ownership, failure behavior, security controls, integration contracts, deployment topology, observability and cost.
The goal is not to produce the most complex architecture. The goal is to design the simplest system that can safely meet the business and operational requirements, while leaving teams enough room to evolve it without repeated rewrites.
Capability
What I focus on
I prefer to describe expertise through architecture decisions and production responsibilities rather than a list of tools.
Business-to-technology mapping
Translate business capabilities and critical journeys into clear service boundaries, ownership, data flows and measurable quality attributes.
HLD, LLD and architecture decisions
Use diagrams and ADRs to capture important choices, rejected alternatives, assumptions and consequences rather than creating documentation that becomes stale.
Non-functional requirements
Treat availability, latency, throughput, security, RTO/RPO, scalability, maintainability and cost as first-class architecture inputs.
Integration architecture
Choose synchronous APIs, events, queues, batch or streaming based on coupling, consistency and operational behavior—not fashion.
Modernization
Break large transformations into controlled migration stages using strangler patterns, anti-corruption layers, data migration plans and measurable cutovers.
Architecture governance
Keep standards lightweight: reusable reference patterns, decision records, security baselines and production-readiness checks that help teams ship.
Architecture
From business objective to operable platform
I prefer an architecture chain where each technical layer can be traced back to a requirement and each requirement can be verified in production. That avoids “architecture by technology list.”
Traceability
Every major component should exist for a reason that can be explained in business, risk or operational terms.
Explicit trade-offs
Consistency, availability, latency, cost and delivery speed rarely maximize together. Record which dimension wins and why.
Evolution
Design interfaces and ownership boundaries so teams can change parts of the system without coordinating every release.
Production practice
How I approach production design
Start with quality attributes, not a reference diagram
Two systems with the same features can require very different architectures. A payment path, an internal content portal and a real-time operational platform may all expose APIs, but their tolerance for downtime, stale data, duplicate processing and recovery time can be completely different.
I therefore convert vague expectations such as “highly available” or “fast” into measurable targets. The architecture can then be tested against those targets instead of being defended by opinion.
- Peak and sustained throughput
- Latency budgets
- Availability target
- RTO and RPO
- Data retention
- Security classification
Define ownership before distributing the system
Microservices without clear ownership create a distributed monolith. I look for business-aligned boundaries, one source of truth for important data and contracts that minimize cross-service transactions.
The same principle applies to teams. If every change requires three teams and two shared databases, the logical diagram may look modular while delivery remains tightly coupled.
- Bounded contexts
- Service ownership
- Data ownership
- API contracts
- Event ownership
- Versioning strategy
Choose communication patterns by failure behavior
A synchronous API is useful when the caller needs an immediate answer. An event is useful when work can be decoupled, replayed or processed independently. Queues help absorb bursts. Streaming helps when ordered event flows and continuous processing matter.
I evaluate what happens when the dependency is slow, unavailable or partially successful. The failure path often tells us more about the right integration pattern than the happy path.
Make migration architecture part of the target architecture
Enterprise systems rarely move from old to new in one release. I include coexistence, routing, data synchronization, rollback and cutover in the design. This is especially important when databases or integration contracts change.
A technically elegant target architecture is not useful if there is no safe route from the current state to that target.
Decision framework
Questions I want answered before approving the design
Do we actually need microservices?
Use independently deployable services when domain boundaries, scaling, release independence or team ownership justify the distributed-systems cost. A modular monolith can be the better architecture when those pressures do not exist.
Where does the source of truth live?
Define authoritative ownership for important entities. Avoid shared-write databases across services because they create hidden coupling and unclear recovery behavior.
What happens when a dependency fails?
Document timeout, retry, circuit-breaker, fallback, queueing and recovery behavior for critical calls. “The service returns an error” is not a complete resilience design.
How will this be operated at 2 AM?
Metrics, logs, traces, dashboards, runbooks, ownership and safe rollback need to exist in the architecture—not be left as post-development tasks.
How will the design evolve?
Prefer stable contracts, versioning rules and replaceable components over framework-specific coupling. Architecture should reduce the blast radius of future change.
Reliability
Failure, scale and operational reality
Risk
Distributed dependency failure
Architecture response
Use bounded timeouts, selective retries, circuit breakers, bulkheads and degraded modes. Avoid retries at every layer because they can multiply load during incidents.
Risk
Cross-service data inconsistency
Architecture response
Use explicit consistency models, idempotency, outbox/event patterns and reconciliation. Do not hide distributed transactions inside service code.
Risk
Release coupling
Architecture response
Version contracts, apply consumer-driven testing where useful and design backward-compatible changes so deployments remain independently schedulable.
Risk
Architecture drift
Architecture response
Use ADRs, automated checks, reference implementations and production-readiness reviews. Governance should surface risk without becoming a delivery bottleneck.
Security
Security by architecture
- Classify data and trust boundaries before choosing controls.
- Keep authentication centralized enough for consistency while authorization remains close to business rules.
- Use least privilege for service identities, databases, queues and cloud resources.
- Design secrets, key rotation, encryption and auditability as deployment concerns.
- Threat-model external integrations, administrative paths and sensitive data flows.
Observability
Operate what we design
- Define service-level indicators around business journeys, not only CPU and memory.
- Correlate logs, traces and metrics with request or business identifiers.
- Measure dependency latency, error rates, saturation and queue/consumer lag.
- Create actionable alerts linked to ownership and runbooks.
- Track architectural assumptions in production and revisit them when scale changes.
Continue reading
Related architecture guides
FAQ
Frequently asked questions
What does a Solution Architect actually own?
The role typically connects business requirements with application, integration, data, security, cloud and operational design. The architect should make important trade-offs explicit, help teams choose workable patterns and ensure critical non-functional requirements are addressed.
What is the difference between HLD and LLD?
HLD explains the major components, boundaries, integrations, deployment and data flows. LLD goes deeper into service contracts, internal components, schemas, sequences and implementation decisions. The exact boundary varies by organization, but both should support delivery rather than duplicate it.
Do enterprise systems always need microservices?
No. Microservices are valuable when independent scaling, deployment, team ownership or domain boundaries justify their operational complexity. A modular monolith is often the better choice when a system does not need those properties.
What makes architecture production-ready?
A production-ready architecture defines failure behavior, security, observability, capacity, deployment, rollback, data recovery and ownership—not only the happy-path component diagram.
How should architecture decisions be documented?
I prefer concise architecture decision records that capture context, options, the chosen decision and consequences. They are easier to maintain than large documents and create a useful history of why the system evolved.