Operations, cost & migration
Cost management and optimisation
Domain 4 is 20% of the exam. Every task statement in it lists the same cost-management tools, so learn them once, then learn the levers for storage, compute, databases and networks.
The tools
| Tool | Purpose | Cue |
|---|---|---|
| AWS Cost Explorer | Visualise and analyse costs and usage; forecasts; Savings Plans/RI recommendations and coverage | "Analyse spending trends", "which service costs most" |
| AWS Budgets | Set cost, usage, RI/Savings Plans budgets; alerts when actual or forecast exceeds thresholds; budget actions (e.g. apply an SCP or stop instances) | "Alert when spend exceeds $X", "prevent overspend" |
| AWS Cost and Usage Report (CUR) / Data Exports | The most detailed line-item billing data, delivered to S3; query with Athena or QuickSight | "Detailed, hourly, resource-level billing analysis" |
| Cost allocation tags | Tag resources (e.g. project, team, environment) and activate the tags to see costs by them |
"Charge back costs to departments" |
| Consolidated billing (Organizations) | One bill, aggregated volume discounts, shared RI/Savings Plans | "Multiple accounts, single bill, maximise discounts" |
| AWS Pricing Calculator | Estimate cost before building | "Estimate cost of a proposed architecture" |
| Trusted Advisor / Compute Optimizer | Idle resources, right-sizing | "Find underutilised resources" |
| Cost Anomaly Detection | ML-based alerts on unusual spend | "Detect unexpected cost spikes" |
Cost levers by area
Compute
- Right-size (Compute Optimizer); Graviton; turn off non-production at night (Instance Scheduler / scheduled scaling).
- Savings Plans / Reserved Instances for steady usage; Spot for interruptible work.
- Serverless (Lambda, Fargate) for spiky or low utilisation.
- Auto Scaling to avoid idle capacity; hibernation for stop/start workloads with long warm-up.
Storage
- Right S3 storage class; lifecycle rules; Intelligent-Tiering for unknown access.
- Delete incomplete multipart uploads, old versions and unattached EBS volumes/old snapshots.
- gp3 instead of gp2; HDD (st1/sc1) for sequential/cold data.
- EFS lifecycle to IA/Archive; EFS One Zone where acceptable.
- Batch uploads of small files (fewer requests) — the exam guide's "batch uploads compared with individual uploads".
- Requester Pays buckets when others download your data.
Databases
- Right-size; reserved instances for steady DBs; Aurora Serverless v2 / DynamoDB on-demand for variable load; DynamoDB provisioned + auto scaling for predictable load.
- Caching to reduce expensive reads; TTL and retention policies; snapshot frequency matched to RPO.
Networking
- Gateway VPC endpoints instead of NAT for S3/DynamoDB.
- Keep traffic in one AZ where possible; avoid unnecessary cross-Region data.
- CloudFront for content delivery.
- Shared NAT gateway for non-production; per-AZ for production.
- Direct Connect for large, steady data transfer volumes.
Key idea
The exam's "most cost-effective" answer still has to meet every requirement. Cheaper-but-not-available is wrong. Find the cheapest option among those that satisfy all constraints.
Workload classes
The exam guide asks you to determine "the required availability for different classes of workloads". Production may need Multi-AZ, reserved capacity and per-AZ NAT; development may run in one AZ, use Spot, shut down at night and share a NAT gateway.
Exam patterns
- "Notify finance when monthly spend is forecast to exceed budget" → AWS Budgets.
- "Show costs per project across 20 accounts" → cost allocation tags + Cost Explorer (consolidated billing).
- "Detailed billing data for custom analysis in Athena" → Cost and Usage Report to S3.
- "Dev environment runs only in office hours" → scheduled stop/start.