Operations, cost & migration
Monitoring, logging and auditing
You can't run what you can't see. These services provide metrics, logs, traces, audit trails and configuration history — and the exam expects you to know exactly which one answers which question.
The one-line distinction
| Question | Service |
|---|---|
| "How is my resource performing?" (metrics, logs, alarms) | CloudWatch |
| "Who did what, and when?" (API calls) | CloudTrail |
| "What did my resource look like, and is it compliant?" | AWS Config |
| "Where is the latency in my distributed request?" | AWS X-Ray |
Amazon CloudWatch
- Metrics — built-in metrics for AWS services; custom metrics from your apps; high-resolution metrics (down to 1 second).
- Alarms — trigger on thresholds or anomaly detection; actions include SNS notifications, Auto Scaling, EC2 actions (stop, terminate, reboot, recover) and Systems Manager actions.
- CloudWatch agent — collects OS-level metrics (memory, disk space) and log files from EC2 and on-premises servers.
- CloudWatch Logs — central log storage; metric filters turn log patterns into metrics; Logs Insights queries logs; subscription filters stream logs to Lambda, Kinesis or OpenSearch; export to S3.
- Dashboards across accounts and Regions.
- Synthetics (canaries that test endpoints) and RUM (real user monitoring).
- Container Insights / Lambda Insights / Application Signals for deeper visibility.
AWS X-Ray
Traces requests as they travel through API Gateway, Lambda, ECS, EC2 and downstream services, building a service map and showing where time and errors occur. The exam guide lists it under workload visibility.
AWS CloudTrail
- Records API calls and account activity (console, CLI, SDK, services).
- Event history — last 90 days of management events, free, per Region.
- Trails — deliver logs to S3 (and optionally CloudWatch Logs) long-term; create an organization trail for all accounts.
- Management events (control plane, logged by default) vs data events (e.g. S3 object reads/writes, Lambda invocations — must be enabled, extra cost).
- Log file integrity validation proves logs weren't tampered with.
- CloudTrail Lake — query events with SQL.
Tip
"Find out who deleted the S3 bucket / terminated the instance" → CloudTrail.
AWS Config
- Records configuration changes of resources over time (configuration history and snapshots).
- Config rules (managed or custom) evaluate compliance — e.g. "S3 buckets must block public access", "EBS volumes must be encrypted".
- Remediation — automatically fix non-compliant resources with Systems Manager Automation.
- Conformance packs and aggregators across accounts and Regions.
Reacting automatically
EventBridge rules can react to events from CloudTrail, Config, GuardDuty, Health and services — e.g. "when someone opens port 22 to the world, invoke Lambda to revert the change and notify security".
Advisory and visibility services
| Service | Purpose |
|---|---|
| AWS Trusted Advisor | Best-practice checks across cost, performance, security, fault tolerance, service limits (full checks with Business/Enterprise support) |
| AWS Compute Optimizer | ML-based right-sizing recommendations for EC2, ASGs, EBS, Lambda, ECS on Fargate |
| AWS Health Dashboard | AWS service events and scheduled maintenance affecting your resources; integrates with EventBridge |
| Amazon Managed Grafana | Managed Grafana dashboards over many data sources |
| Amazon Managed Service for Prometheus | Managed Prometheus-compatible metrics for containers |
| AWS Well-Architected Tool | Review workloads against the framework |
Exam patterns
- "Alert when free memory on EC2 drops below 10%" → CloudWatch agent + alarm.
- "Count application errors in logs and alarm" → CloudWatch Logs metric filter + alarm.
- "Prove which IAM user changed a security group" → CloudTrail.
- "Continuously check all buckets are encrypted and auto-remediate" → Config rules + remediation.
- "Find which microservice adds latency" → X-Ray.
- "Automatically recover an instance if the underlying host fails" → CloudWatch alarm with the EC2 recover action.