Auto Scaling
Auto Scaling adjusts capacity automatically to match demand and replaces unhealthy instances. Combined with a load balancer across multiple AZs, it's the standard answer for scalable, highly available EC2 workloads.
Auto Scaling groups (ASGs)
An ASG manages a fleet of EC2 instances:
- Launch template — what to launch (AMI, instance type(s), security groups, IAM role, user data).
- Minimum, desired and maximum capacity.
- Subnets in multiple AZs — the ASG balances instances across them and rebalances after an AZ recovers.
- Load balancer target groups — new instances register automatically.
Health checks
- EC2 status checks (default) — replace instances with failed hardware/OS checks.
- ELB health checks — replace instances that fail the application health check. Enable this when behind a load balancer so a hung app is replaced, not just a dead server.
- Grace period — time for a new instance to start before health checks count.
Scaling policies
| Policy | How it works | Use when |
|---|---|---|
| Target tracking | Keep a metric at a target (e.g. average CPU 50%, ALB requests per target 1,000) | Most cases — simplest and recommended |
| Step scaling | Add/remove amounts based on alarm breach size | Need different responses to small vs large spikes |
| Simple scaling | One adjustment per alarm, then a cooldown | Legacy; prefer target tracking |
| Scheduled scaling | Change capacity at set times | Predictable patterns (business hours, weekly reports) |
| Predictive scaling | Forecasts load from history and scales ahead of time | Recurring daily/weekly cycles with slow-starting instances |
Scale on the metric that represents load for your app — CPU, ALB request count per target, or SQS queue backlog per instance for worker fleets.
Speeding up scaling
- Pre-baked AMIs — less work at boot.
- Warm pools — keep pre-initialised instances stopped (or hibernated) and ready.
- Instance refresh — roll out a new launch template version gradually.
Lifecycle hooks
Pause instances during launch or termination to run custom actions — install software, register with a system, or copy logs off an instance before it's terminated.
Termination behaviour
On scale-in, the default policy first balances AZs, then prefers instances with the oldest launch template/configuration, then those closest to the next billing hour. Scale-in protection can protect specific instances.
Mixed instances and Spot
An ASG can mix On-Demand and Spot across multiple instance types, e.g. a guaranteed On-Demand base plus Spot for the rest — a strong cost-optimized pattern for stateless fleets.
AWS Auto Scaling (beyond EC2)
Application Auto Scaling and the AWS Auto Scaling console scale other resources too: ECS services, DynamoDB capacity, Aurora replicas, Spot Fleets and more.
Exam patterns
- "Handle unpredictable traffic spikes and replace failed instances automatically" → ASG across AZs behind an ALB with target tracking.
- "Traffic jumps every weekday at 9 a.m." → scheduled (or predictive) scaling.
- "Workers process SQS messages; scale with the backlog" → target tracking on backlog per instance.
- "Instances are marked healthy but the app has crashed" → enable ELB health checks on the ASG.
- "Save logs before instances terminate" → lifecycle hook.