Designing for the four domains
Domain 2: Designing resilient architectures
Domain 2 (26%) asks whether your designs keep working when parts fail and when load changes. Its two task statements are scalability through loose coupling (2.1) and high availability / fault tolerance (2.2).
Key vocabulary
| Term | Meaning |
|---|---|
| High availability | The system stays up with minimal downtime (e.g. a brief failover) |
| Fault tolerance | The system keeps running with no interruption when components fail (full redundancy) |
| RPO (recovery point objective) | How much data loss is acceptable, measured in time ("we can lose up to 15 minutes of data") |
| RTO (recovery time objective) | How much downtime is acceptable ("we must be back within 1 hour") |
| Single point of failure | Any component whose failure brings the whole system down |
2.1 Scalable and loosely coupled architectures
| Requirement | Pattern |
|---|---|
| Absorb spikes, protect backends | SQS between producers and consumers |
| Notify many systems of one event | SNS fan-out to SQS queues; EventBridge rules |
| Multi-step workflows | Step Functions |
| Stateless web/app tier | Sessions in ElastiCache/DynamoDB; files in S3/EFS |
| Scale horizontally | Auto Scaling groups / ECS service scaling behind ALB |
| Scale reads | Read replicas, Aurora replicas, caching (ElastiCache, DAX, CloudFront) |
| Serverless scaling | Lambda, API Gateway, DynamoDB, Fargate, Aurora Serverless v2 |
| Containerise apps | ECS/EKS (on Fargate for least overhead) |
| Choose storage type | Object (S3), file (EFS/FSx), block (EBS) by access pattern |
| Purpose-built services | Transfer Family for SFTP, managed AI services, managed databases |
2.2 Highly available and fault-tolerant architectures
Eliminate single points of failure
| Layer | Highly available design |
|---|---|
| DNS | Route 53 (100% availability SLA) with health checks |
| Edge | CloudFront / Global Accelerator |
| Load balancing | ALB/NLB across ≥2 AZs |
| Compute | Auto Scaling group across ≥2 AZs; ECS services across AZs; Lambda (multi-AZ by design) |
| NAT | One NAT gateway per AZ |
| Database | RDS Multi-AZ, Aurora (replicas in other AZs), DynamoDB (multi-AZ by design) |
| Storage | S3 (multi-AZ), EFS Regional, EBS snapshots |
| Connectivity | Redundant VPN tunnels / two Direct Connect locations |
| Connections | RDS Proxy to survive DB failovers gracefully |
Other resilience practices
- Health checks at every layer, with automatic replacement.
- Retries with exponential backoff and jitter; timeouts; idempotent operations.
- Immutable infrastructure and infrastructure as code for consistent rebuilds.
- Service quotas raised in DR Regions ahead of time.
- Workload visibility with CloudWatch and X-Ray, and metrics tied to business needs (e.g. orders per minute).
The four disaster recovery strategies
| Strategy | Setup in the DR Region | RPO / RTO | Cost |
|---|---|---|---|
| Backup and restore | Only backups (snapshots, AMIs, S3 replication); rebuild with IaC when needed | Hours | Lowest |
| Pilot light | Core data replicated and live (e.g. database replica); app servers off/not created; scale up on disaster | Tens of minutes | Low |
| Warm standby | A scaled-down but fully working copy running; scale up on disaster | Minutes | Medium |
| Multi-site active/active | Full capacity in multiple Regions serving traffic | Near zero | Highest |
Cost & speed of recovery →
Backup & restore → Pilot light → Warm standby → Active/active
(hours) (10s of min) (minutes) (near zero)
Supporting services: AWS Backup cross-Region copies, S3 CRR, RDS cross-Region read replicas, Aurora Global Database, DynamoDB global tables, Route 53 failover/latency routing, Global Accelerator, CloudFormation to rebuild, AWS Elastic Disaster Recovery for continuous replication of servers.
Key idea
Match the DR strategy to the stated RPO/RTO at the lowest cost. RTO of a few hours and lowest cost → backup and restore. RTO in minutes with minimal running cost → warm standby. "No downtime" → active/active.
Exam patterns
- "Must remain available if an AZ fails" → Multi-AZ everything; no multi-Region needed.
- "RPO 1 hour, RTO 24 hours, minimise cost" → backup and restore to another Region.
- "RTO 15 minutes; minimise always-on cost" → pilot light or warm standby (warm standby if the app takes long to scale).
- "Orders must never be lost even if processing fails" → SQS (with DLQ) between tiers.