AWS Solutions Architect Handbook SAA-C03, from zero Bipin Singh
Designing for the four domains

Domain 2: Designing resilient architectures

3 min readChapter 41 of 48By Bipin Singh

Domain 2 (26%) asks whether your designs keep working when parts fail and when load changes. Its two task statements are scalability through loose coupling (2.1) and high availability / fault tolerance (2.2).

Key vocabulary

Term Meaning
High availability The system stays up with minimal downtime (e.g. a brief failover)
Fault tolerance The system keeps running with no interruption when components fail (full redundancy)
RPO (recovery point objective) How much data loss is acceptable, measured in time ("we can lose up to 15 minutes of data")
RTO (recovery time objective) How much downtime is acceptable ("we must be back within 1 hour")
Single point of failure Any component whose failure brings the whole system down

2.1 Scalable and loosely coupled architectures

Requirement Pattern
Absorb spikes, protect backends SQS between producers and consumers
Notify many systems of one event SNS fan-out to SQS queues; EventBridge rules
Multi-step workflows Step Functions
Stateless web/app tier Sessions in ElastiCache/DynamoDB; files in S3/EFS
Scale horizontally Auto Scaling groups / ECS service scaling behind ALB
Scale reads Read replicas, Aurora replicas, caching (ElastiCache, DAX, CloudFront)
Serverless scaling Lambda, API Gateway, DynamoDB, Fargate, Aurora Serverless v2
Containerise apps ECS/EKS (on Fargate for least overhead)
Choose storage type Object (S3), file (EFS/FSx), block (EBS) by access pattern
Purpose-built services Transfer Family for SFTP, managed AI services, managed databases

2.2 Highly available and fault-tolerant architectures

Eliminate single points of failure

Layer Highly available design
DNS Route 53 (100% availability SLA) with health checks
Edge CloudFront / Global Accelerator
Load balancing ALB/NLB across ≥2 AZs
Compute Auto Scaling group across ≥2 AZs; ECS services across AZs; Lambda (multi-AZ by design)
NAT One NAT gateway per AZ
Database RDS Multi-AZ, Aurora (replicas in other AZs), DynamoDB (multi-AZ by design)
Storage S3 (multi-AZ), EFS Regional, EBS snapshots
Connectivity Redundant VPN tunnels / two Direct Connect locations
Connections RDS Proxy to survive DB failovers gracefully

Other resilience practices

The four disaster recovery strategies

Strategy Setup in the DR Region RPO / RTO Cost
Backup and restore Only backups (snapshots, AMIs, S3 replication); rebuild with IaC when needed Hours Lowest
Pilot light Core data replicated and live (e.g. database replica); app servers off/not created; scale up on disaster Tens of minutes Low
Warm standby A scaled-down but fully working copy running; scale up on disaster Minutes Medium
Multi-site active/active Full capacity in multiple Regions serving traffic Near zero Highest
Cost & speed of recovery →
Backup & restore  →  Pilot light  →  Warm standby  →  Active/active
   (hours)            (10s of min)     (minutes)       (near zero)

Supporting services: AWS Backup cross-Region copies, S3 CRR, RDS cross-Region read replicas, Aurora Global Database, DynamoDB global tables, Route 53 failover/latency routing, Global Accelerator, CloudFormation to rebuild, AWS Elastic Disaster Recovery for continuous replication of servers.

Key idea

Match the DR strategy to the stated RPO/RTO at the lowest cost. RTO of a few hours and lowest cost → backup and restore. RTO in minutes with minimal running cost → warm standby. "No downtime" → active/active.

Exam patterns

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I design and run production systems on AWS — serverless, data and AI.

Work with me