Prev Next

Cloud / AWS Auto Scaling Interview questions

Last updated

1. What is AWS Auto Scaling? 2. What is an Auto Scaling group? 3. What are the benefits of using EC2 Auto Scaling? 4. What is a launch template in Auto Scaling? 5. What are the types of scaling in AWS Auto Scaling? 6. What are minimum, maximum, and desired capacity in an Auto Scaling group? 7. What is dynamic scaling in Auto Scaling? 8. What is scheduled scaling in Auto Scaling? 9. What is predictive scaling in Auto Scaling? 10. What is a target tracking scaling policy? 11. What is a cooldown period in Auto Scaling? 12. What are health checks in an Auto Scaling group? 13. What is the health check grace period in Auto Scaling? 14. What are lifecycle hooks in Auto Scaling? 15. What is the difference between horizontal and vertical scaling? 16. What is the difference between EC2 Auto Scaling and Application Auto Scaling? 17. What is the difference between a launch configuration and a launch template? 18. What are the termination policies in Auto Scaling? 19. What are Auto Scaling group metrics? 20. What is instance warmup in Auto Scaling? 21. How does step scaling differ from simple scaling? 22. How does Auto Scaling balance instances across Availability Zones? 23. Which is better, target tracking or step scaling, and why? 24. How does an Auto Scaling group work with an Application Load Balancer? 25. What is scale-in protection in Auto Scaling? 26. What happens when an instance is put into Standby in an Auto Scaling group? 27. What happens when you suspend an Auto Scaling process? 28. How do you scale an Auto Scaling group based on SQS queue depth? 29. Explain the lifecycle of an instance in an Auto Scaling group? 30. When should you use a warm pool in Auto Scaling? 31. How does instance refresh work in Auto Scaling? 32. Why use a mixed instances policy in an Auto Scaling group? 33. How do you use Spot Instances safely in an Auto Scaling group? 34. What is attribute-based instance type selection? 35. How do you handle a graceful shutdown during Auto Scaling scale-in? 36. How do you troubleshoot an Auto Scaling group that fails to launch instances? 37. Why are instances terminated right after launch in an Auto Scaling group? 38. When should you use scheduled scaling together with dynamic scaling? 39. What happens when multiple scaling policies trigger at the same time? 40. How do you test predictive scaling safely before enabling it? 41. How do you perform a blue/green deployment with Auto Scaling groups? 42. How can you optimize Auto Scaling costs? 43. What is an instance maintenance policy in Auto Scaling? 44. How do you scale Amazon ECS services and capacity with Auto Scaling? 45. How do you scale DynamoDB capacity automatically? 46. How do you get notified about Auto Scaling events? 47. How do you secure instances launched by an Auto Scaling group? 48. How do you design stateful applications to run in Auto Scaling groups? 49. Explain the execution flow of a scale-out event with target tracking? 50. What happens when an Availability Zone becomes impaired for an Auto Scaling group?

1. What is AWS Auto Scaling?

AWS Auto Scaling is a set of capabilities that automatically adjusts the capacity of your AWS resources to match demand. For EC2 workloads this is done by Amazon EC2 Auto Scaling, which adds or removes instances in a group as load rises and falls.

You define a minimum, maximum, and desired number of instances, plus rules that say when to change that number. AWS then launches or terminates instances for you and replaces any that fail health checks.

The same idea also applies beyond EC2 through Application Auto Scaling, which handles resources like ECS services, DynamoDB tables, and Aurora replicas.

The result is better availability during spikes and lower cost during quiet periods, without anyone watching dashboards.

Take quiz
Which service adds or removes EC2 instances automatically based on demand?
AWS Config
Amazon Route 53
Amazon EC2 Auto Scaling
AWS CloudTrail
Besides adding capacity at peak, what else does Auto Scaling do for a group?
Rewrites application code on failure
Encrypts instance traffic by default
Converts instances to bigger sizes in place
Replaces instances that fail health checks

2. What is an Auto Scaling group?

An Auto Scaling group (ASG) is the logical collection of EC2 instances that Auto Scaling manages as a single unit. Every instance in it is launched from the same template, and the group keeps the instance count within the limits you set.

An ASG is defined by a few core settings:

  • Launch template that describes what to launch (AMI, instance type, security groups, and so on)
  • Subnets that decide which Availability Zones are used
  • Min, max, and desired capacity
  • Optional load balancer target groups and scaling policies

If an instance dies, the group notices and launches a replacement to get back to the desired capacity.

Take quiz
What does an Auto Scaling group use to know what kind of instance to launch?
A CloudWatch dashboard
A Route 53 hosted zone
An IAM permission boundary
A launch template
If one instance in an ASG crashes, the group will typically:
Launch a replacement to restore desired capacity
Shrink the max capacity permanently
Wait for an administrator to approve a new instance
Stop all other instances in the group

3. What are the benefits of using EC2 Auto Scaling?

The main benefit is that capacity follows demand automatically, so you stop guessing how many servers you need.

  • Availability: unhealthy instances are replaced and instances are spread across multiple Availability Zones.
  • Cost control: you scale in when traffic drops, so you pay for fewer instance hours.
  • Performance: new instances come up before users feel the slowdown.
  • Less manual work: no one has to launch or terminate instances by hand.
  • No extra charge: EC2 Auto Scaling itself is free; you pay for the EC2 and CloudWatch resources it uses.

It also pairs naturally with Elastic Load Balancing, so new instances start receiving traffic as soon as they are healthy.

Take quiz
What do you pay for when using EC2 Auto Scaling itself?
Nothing extra for the service; you pay for the EC2 and CloudWatch resources used
A flat monthly fee per Auto Scaling group
A per-scaling-activity charge
A per-vCPU licensing fee
Which benefit comes from spreading an ASG over several Availability Zones?
Lower EBS storage prices
Higher availability if one zone has a problem
Automatic database backups
Faster DNS propagation

4. What is a launch template in Auto Scaling?

A launch template stores the settings an Auto Scaling group needs to launch an EC2 instance, so every new instance comes up configured the same way.

Typical contents include the AMI ID, instance type, key pair, security groups, IAM instance profile, EBS volume settings, user data, and metadata options such as IMDSv2.

Launch templates are versioned. You can create version 2 with a new AMI, then point the ASG at it (or at $Latest or $Default) without rebuilding everything. They also support newer features such as mixed instance types, Spot options, and attribute-based instance selection, which is why AWS recommends them over launch configurations.

Take quiz
What makes launch templates easier to maintain than older approaches?
They can only hold one fixed instance type
They support multiple versions you can reference
They cannot be edited after creation
They run inside the instance as a daemon
Which setting is typically defined inside a launch template?
Route 53 failover records
CloudTrail log retention
AMI ID and security groups
Scaling cooldown alarms for other accounts

5. What are the types of scaling in AWS Auto Scaling?

EC2 Auto Scaling supports four ways of changing capacity:

Type How it works Typical use
Manual You change desired capacity yourself One-off changes, testing
Scheduled Capacity changes at set times Known patterns like office hours
Dynamic Reacts to CloudWatch metrics in real time Unpredictable traffic
Predictive Forecasts load from history and scales ahead Recurring daily or weekly cycles

Dynamic scaling itself has three policy styles: target tracking, step, and simple scaling. Most teams start with target tracking and add scheduled or predictive scaling when traffic has a clear rhythm.

Take quiz
Which scaling type uses machine learning on historical load to add capacity before demand arrives?
Manual scaling
Simple scaling
Predictive scaling
Scheduled scaling
Which scaling type is the best fit when load jumps at 9 AM every weekday and you know it in advance?
Manual scaling on demand
Lifecycle hook scaling
Health check scaling
Scheduled scaling

6. What are minimum, maximum, and desired capacity in an Auto Scaling group?

These three numbers control how many instances the group runs.

  • Minimum: the group never goes below this, even if load is zero.
  • Maximum: the group never goes above this, even under heavy load. It acts as a cost safety net.
  • Desired: the number the group tries to maintain right now.

Scaling policies change the desired value, always staying inside the min and max range. If you set desired to 4 with min 2 and max 6, the group launches or terminates instances until it has exactly 4. Setting desired outside the min-max range is rejected.

Take quiz
An ASG has min 2, max 6, and desired 4. A scaling policy asks for 9. What will the group run?
9 instances, because policies override max
4 instances, because desired is locked
2 instances, because min wins
6 instances, because max caps it
Which value acts as the cost safety net during a runaway traffic spike?
Maximum capacity
Minimum capacity
Health check grace period
Cooldown

7. What is dynamic scaling in Auto Scaling?

Dynamic scaling changes group capacity automatically in response to live CloudWatch metrics such as CPU utilization, request count per target, or a custom metric.

You attach a scaling policy to the group. When the metric crosses a threshold (or drifts from a target value), Auto Scaling raises or lowers desired capacity. The policy types are:

  • Target tracking: keep a metric near a target, like 50% average CPU.
  • Step scaling: add or remove different amounts depending on how far the alarm is breached.
  • Simple scaling: one adjustment per alarm, followed by a cooldown.

It handles surprises well, but it reacts after load changes, so there is always a short lag while instances boot.

Take quiz
What triggers a dynamic scaling action?
A CloudWatch metric moving past a threshold or away from a target
A calendar entry only
A manual AWS Support request
A change in IAM policy
What is the main limitation of purely dynamic scaling?
It cannot scale in
It reacts after load changes, so there is boot lag
It needs a fixed number of instances
It only works with Spot Instances

8. What is scheduled scaling in Auto Scaling?

Scheduled scaling changes the min, max, or desired capacity at specific times, so capacity is ready before a known load change.

You create a scheduled action with either a one-time date or a recurring cron expression, and you can set a time zone. For example, raise the minimum to 10 at 8:30 AM on weekdays and drop it back to 2 at 7 PM.

aws autoscaling put-scheduled-update-group-action \
  --auto-scaling-group-name web-asg \
  --scheduled-action-name weekday-morning \
  --recurrence "30 8 * * 1-5" \
  --time-zone "America/New_York" \
  --min-size 10 --desired-capacity 10

It works well for marketing launches, batch windows, and office-hours apps, and it can run alongside dynamic scaling policies.

Take quiz
Which expression format does a recurring scheduled action use?
A CloudFormation intrinsic function
A cron expression
An IAM condition key
A Lambda alias name
Which workload fits scheduled scaling best?
A game with random viral spikes
A service with zero traffic pattern knowledge
A payroll app with a heavy load every Friday afternoon
A one-time database migration script

9. What is predictive scaling in Auto Scaling?

Predictive scaling uses machine learning to study your group's past load and forecast future demand. It then schedules capacity ahead of the forecasted spike instead of waiting for a metric to breach.

It needs at least 24 hours of metric history and works best for traffic that repeats daily or weekly, such as business-hours web apps. Forecasts look ahead 48 hours and are refreshed regularly.

It has two modes:

  • Forecast only: generates forecasts so you can review them, but changes nothing.
  • Forecast and scale: actually launches instances based on the forecast.

Predictive scaling is usually combined with dynamic scaling, which still handles unexpected spikes.

Take quiz
What does predictive scaling use to decide when to add capacity?
A fixed cron expression only
Manual alarms set by the user
A forecast built from historical metrics
The AMI creation date
Which mode lets you inspect forecasts without launching any instances?
Forecast and scale
Simple scaling
Standby mode
Forecast only

10. What is a target tracking scaling policy?

A target tracking policy keeps a chosen metric close to a target value, the way a thermostat holds a room temperature. You pick the metric and the target, and Auto Scaling does the math.

Predefined metrics include ASGAverageCPUUtilization, ASGAverageNetworkIn, ASGAverageNetworkOut, and ALBRequestCountPerTarget. You can also use a custom metric.

Behind the scenes it creates and manages the CloudWatch alarms for you, so you should not edit them. It scales out quickly when the metric is above target and scales in more cautiously when it is below, to avoid flapping.

aws autoscaling put-scaling-policy \
  --auto-scaling-group-name web-asg \
  --policy-name cpu50 \
  --policy-type TargetTrackingScaling \
  --target-tracking-configuration '{"PredefinedMetricSpecification":{"PredefinedMetricType":"ASGAverageCPUUtilization"},"TargetValue":50.0}'

Take quiz
With a target tracking policy, who manages the CloudWatch alarms?
You must create both alarms manually
AWS Config manages them
They are not used at all
Auto Scaling creates and manages them for you
Which is a predefined metric for target tracking on an ASG behind an ALB?
ALBRequestCountPerTarget
VolumeQueueLength
ConsumedWriteCapacityUnits
BucketSizeBytes

11. What is a cooldown period in Auto Scaling?

A cooldown period is a pause after a scaling activity during which a simple scaling policy cannot trigger another one. It gives the new instances time to start and take load before the group decides whether more capacity is needed.

The default is 300 seconds, and you can set it for the whole group or override it for a single policy. Without it, a metric that is still high while new instances are booting could trigger repeated scale-outs and leave you with far more capacity than needed.

Target tracking and step scaling do not rely on cooldown. They use instance warmup instead, which is a more precise way to handle the same problem.

Take quiz
What is the default cooldown value for an Auto Scaling group?
300 seconds
30 seconds
60 minutes
24 hours
Which policy type depends mainly on cooldown rather than instance warmup?
Target tracking scaling
Simple scaling
Step scaling
Predictive scaling

12. What are health checks in an Auto Scaling group?

Health checks let the group decide whether an instance is fit to serve. When an instance is marked unhealthy, Auto Scaling terminates it and launches a replacement.

Several health check sources exist:

  • EC2 status checks: the default. They catch hardware or hypervisor level failures, not a crashed web app.
  • ELB health checks: optional. They use the load balancer's HTTP/HTTPS checks, so they detect application-level failures.
  • Custom health checks: your own tooling reports status through the set-instance-health API.

If your instances sit behind a load balancer, turn on ELB health checks. Otherwise an instance with a dead application can look perfectly healthy to EC2.

Take quiz
Which health check type catches a web application that has crashed while the VM is still running?
EC2 status checks only
ELB health checks
CloudTrail data events
Billing alarms
Which API lets you report an instance as unhealthy from your own monitoring tool?
put-scaling-policy
attach-instances
set-instance-health
describe-launch-templates

13. What is the health check grace period in Auto Scaling?

The health check grace period is the time Auto Scaling waits after an instance enters InService before it starts checking that instance's health. It exists because a new instance needs time to boot, run user data, and start the application.

The commonly used default is 300 seconds. If the grace period is shorter than your real startup time, the load balancer will report the instance unhealthy, Auto Scaling will terminate it, and the group can get stuck in a launch-and-terminate loop.

Measure how long your application takes to become ready, then set the grace period slightly above that. Setting it far too high delays replacement of instances that really are broken.

Take quiz
What problem does a too-short grace period typically cause?
Instances never get public IPs
Scaling policies stop evaluating
Instances are terminated before the app finishes starting
The launch template gets deleted
When does the grace period timer start?
When the AMI is created
When the scaling policy is attached
When the max capacity is reached
When the instance enters InService

14. What are lifecycle hooks in Auto Scaling?

A lifecycle hook pauses an instance at a transition point so you can run custom actions before it moves on. There are two hook types:

  • autoscaling:EC2_INSTANCE_LAUNCHING puts the instance in Pending:Wait, useful for installing agents or pulling config before it receives traffic.
  • autoscaling:EC2_INSTANCE_TERMINATING puts it in Terminating:Wait, useful for draining connections, uploading logs, or deregistering from a service.

Auto Scaling notifies you through EventBridge, SNS, or SQS. Your automation does its work, then calls complete-lifecycle-action with CONTINUE or ABANDON. If nothing responds, the hook times out (one hour by default) and the default result applies.

You can extend the wait by sending a heartbeat with record-lifecycle-action-heartbeat.

Take quiz
In which state does an instance wait when a launch lifecycle hook is active?
Standby
Detached
Terminated
Pending:Wait
What do you call to tell Auto Scaling your custom action is finished?
complete-lifecycle-action
put-notification-configuration
enter-standby
suspend-processes

15. What is the difference between horizontal and vertical scaling?

Horizontal scaling (scale out/in) changes the number of instances. Vertical scaling (scale up/down) changes the size of a single instance, such as moving from m5.large to m5.2xlarge.

Aspect Horizontal Vertical
What changes Instance count Instance size
Downtime None, new nodes join the pool Usually a stop and start
Upper limit Very high Largest instance type
Auto Scaling support Yes, this is what an ASG does No, manual or scripted
Needs stateless app Generally yes Not necessarily

EC2 Auto Scaling is built for horizontal scaling. Vertical scaling still has a place for databases or legacy apps that cannot run on several nodes.

Take quiz
Which action is horizontal scaling?
Adding more instances behind a load balancer
Changing m5.large to m5.4xlarge
Adding more RAM to one VM
Upgrading an EBS volume type
Which statement about vertical scaling on EC2 is accurate?
It is performed by target tracking policies
It usually requires stopping and resizing the instance
It has no upper limit
It never causes downtime

16. What is the difference between EC2 Auto Scaling and Application Auto Scaling?

EC2 Auto Scaling manages groups of EC2 instances. Application Auto Scaling is a separate API that scales other AWS resources through a common model of scalable targets and scaling policies.

Aspect EC2 Auto Scaling Application Auto Scaling
Scales EC2 instances in an ASG ECS services, DynamoDB, Aurora replicas, Lambda provisioned concurrency, SageMaker endpoints and more
Core object Auto Scaling group Scalable target
Instance replacement and health checks Yes No, it only adjusts capacity
Predictive scaling Yes No
CLI namespace aws autoscaling aws application-autoscaling

A containerized app on ECS often uses both: Application Auto Scaling for the service task count and EC2 Auto Scaling for the cluster instances.

Take quiz
Which service scales a DynamoDB table's provisioned throughput?
EC2 Auto Scaling
Application Auto Scaling
AWS Backup
Amazon Inspector
Which feature is available in EC2 Auto Scaling but not in Application Auto Scaling?
Target tracking policies
Scheduled actions
Instance health replacement and lifecycle hooks
CloudWatch metric support

17. What is the difference between a launch configuration and a launch template?

Both describe how to launch instances for an ASG, but launch templates are the modern replacement. Launch configurations are legacy: they are immutable and have not received new features, and AWS no longer lets new accounts create them.

Feature Launch configuration Launch template
Versioning No, create a new one for every change Yes, multiple versions
Edit after creation No Create a new version
Mixed instance types and Spot/On-Demand mix No Yes
Newer instance features (for example T3 unlimited, Capacity Reservations) Limited or no Yes
Also usable for plain EC2 launches No Yes

If you still have an ASG on a launch configuration, AWS provides a way to migrate it to a launch template from the console or CLI.

Take quiz
Which statement about launch configurations is correct?
They support versioning like launch templates
They are required for mixed instances policies
They are immutable and considered legacy
They are the recommended default today
Which capability requires a launch template?
Setting a min capacity
Attaching an ALB target group
Using CloudWatch alarms
Mixing On-Demand and Spot with several instance types

18. What are the termination policies in Auto Scaling?

A termination policy decides which instance is removed first during a scale-in. The default policy first picks the Availability Zone with the most instances so the group stays balanced, then prefers instances on the oldest launch template or configuration, then those closest to the next billing hour, and finally picks randomly.

You can replace the default with one or more of these:

  • OldestInstance and NewestInstance
  • OldestLaunchTemplate and OldestLaunchConfiguration
  • ClosestToNextInstanceHour
  • AllocationStrategy to align with your Spot and On-Demand mix
  • A custom Lambda function for your own selection logic

Instances with scale-in protection are skipped. A common pattern is OldestLaunchTemplate so outdated instances disappear first after a rollout.

Take quiz
What does the default termination policy look at first?
Which instance has the biggest EBS volume
Which instance has the lowest CPU
Which instance has the longest hostname
Which Availability Zone has the most instances
Which termination policy helps retire instances still on an old template after a rollout?
OldestLaunchTemplate
NewestInstance
ClosestToNextInstanceHour
AllocationStrategy only

19. What are Auto Scaling group metrics?

Group metrics are CloudWatch metrics that describe the state of the ASG itself, not the instances inside it. They report counts such as:

  • GroupMinSize, GroupMaxSize, GroupDesiredCapacity
  • GroupInServiceInstances and GroupPendingInstances
  • GroupStandbyInstances and GroupTerminatingInstances
  • GroupTotalInstances

They are sent at one-minute granularity and cost nothing, but you have to switch them on with enable-metrics-collection. They are handy for alarms like "desired capacity has been at max for 30 minutes", which usually means you are being capped.

Do not confuse them with EC2 instance metrics like CPUUtilization, which scaling policies actually act on. Turning on detailed monitoring gives those one-minute data points too.

Take quiz
Which metric tells you how many instances the group currently wants?
GroupDesiredCapacity
CPUCreditBalance
StatusCheckFailed
NetworkPacketsIn
What must you do before group metrics appear in CloudWatch?
Install the CloudWatch agent on every instance
Enable metrics collection on the group
Create a Lambda forwarder
Attach a lifecycle hook

20. What is instance warmup in Auto Scaling?

Instance warmup is the time a newly launched instance needs before it can handle a full share of load. During that time, its metrics are left out of the group's aggregate, so Auto Scaling does not act on numbers that are still ramping up.

Without warmup, a freshly booted instance reports low CPU, drags the average down, and may trigger an early scale-in. Or the opposite: the group keeps adding instances while earlier ones are still starting.

Set a default instance warmup on the group so target tracking, step scaling, and instance refresh all use the same value. Pick a number close to the real time from launch until the app is serving normally.

Take quiz
What happens to a new instance's metrics during warmup?
They are doubled to compensate
They are excluded from the group's aggregated metrics
They are sent only to CloudTrail
They trigger an immediate scale-in
Which policy types use instance warmup instead of cooldown?
Simple scaling only
Scheduled actions only
Target tracking and step scaling
Lifecycle hooks only

21. How does step scaling differ from simple scaling?

Both are triggered by a CloudWatch alarm, but step scaling can pick different adjustments based on how far the metric has gone past the threshold, while simple scaling always applies a single adjustment.

With step scaling you define ranges. For example, with an alarm at 60% CPU: 60-70% adds 1 instance, 70-85% adds 2, and anything above 85% adds 4. The bigger the breach, the bigger the response.

The other difference is timing. Simple scaling must finish its activity and sit out the cooldown before it reacts again. Step scaling keeps evaluating the alarm while a scaling activity is running and uses instance warmup to avoid double-counting capacity that is still starting. This makes it far more responsive during fast spikes.

AWS now recommends target tracking or step scaling over simple scaling for most cases.

Take quiz
What lets step scaling react proportionally to the size of a spike?
A fixed single adjustment
A manual approval step
Step adjustments tied to alarm breach ranges
A larger cooldown value
Which policy type must wait for cooldown before responding to the next alarm?
Step scaling
Target tracking
Predictive scaling
Simple scaling

22. How does Auto Scaling balance instances across Availability Zones?

When you give an ASG subnets in several Availability Zones, it tries to keep the instance count even across zones. On scale-out, it launches into the zone with the fewest instances. On scale-in, the default termination policy removes from the zone with the most.

If the spread gets uneven, for instance after an AZ recovers from an outage or a launch fails in one zone, the AZRebalance process steps in. It launches a new instance first in the under-populated zone and only then terminates one from the crowded zone, so capacity never dips.

During rebalancing, the group can briefly go above its maximum (by up to about 10%, or one instance, whichever is larger). If you want strict numbers, you can suspend AZRebalance, but you give up automatic balancing.

Take quiz
In which order does AZRebalance work?
Terminate first, then launch
Stop all instances, then start new ones
Only launch, never terminate
Launch a new instance first, then terminate the old one
On scale-out, where does the group place the next instance?
In the AZ with the fewest instances
In the AZ with the cheapest storage
Always in the first listed AZ
In a randomly chosen Region

23. Which is better, target tracking or step scaling, and why?

For most workloads, target tracking is the better starting point because it is simpler and self-tuning. You state a goal like "keep average CPU at 50%" and Auto Scaling works out how many instances are needed, creating and adjusting alarms for you.

Aspect Target tracking Step scaling
Setup effort One metric, one target Alarms plus step table
Alarm management Automatic You manage them
Control over amounts Proportional, calculated for you Fully manual per step
Metric requirement Must move inversely with capacity (for example average CPU) Any metric you can alarm on
Best for Steady, proportional workloads Custom or uneven response rules

Choose step scaling when the metric does not fall as capacity rises, when you want an aggressive jump at a certain threshold, or when scale-in and scale-out need very different shapes. You can attach both types to the same group if you need to.

Take quiz
Which metric is a poor fit for target tracking?
A metric that does not decrease when you add instances
Average CPU across the group
ALB requests per target
Average network in per instance
When is step scaling the better choice?
When you want AWS to manage alarms
When you need custom jump sizes at specific thresholds
When you have no CloudWatch access
When you only need a fixed schedule

24. How does an Auto Scaling group work with an Application Load Balancer?

You attach one or more target groups to the ASG. From then on, the group registers each new instance with the target group once it is ready, and deregisters it before it is terminated.

flowchart LR
  A["Scale-out event"] --> B["Instance launches"]
  B --> C["Launch hook completes"]
  C --> D["Registered to target group"]
  D --> E{ALB health check passes?}
  E -- Yes --> F["Receives traffic"]
  E -- No --> G["Marked unhealthy and replaced"]

A few details matter in practice:

  • Enable ELB health checks on the ASG so a failing app triggers replacement, not just a failing VM.
  • On scale-in, the instance is deregistered and the target group's deregistration delay lets in-flight requests finish.
  • Use the ALBRequestCountPerTarget metric for target tracking so scaling follows actual request load.
  • The ASG's subnets should cover the same Availability Zones the ALB is enabled in.
Take quiz
What lets in-flight requests complete when an instance is removed from a target group?
Health check grace period
Deregistration delay
Cooldown period
Instance warmup
Which ASG setting makes an application-level failure trigger instance replacement?
Disabling AZRebalance
Raising max capacity
Turning on ELB health checks
Switching to a larger instance type

25. What is scale-in protection in Auto Scaling?

Scale-in protection stops Auto Scaling from selecting an instance for termination during a scale-in event. It is useful for instances doing long jobs, like video encoding or a batch worker holding a lease.

It can be set at two levels:

  • Group level: every newly launched instance starts protected.
  • Instance level: you protect or unprotect individual instances, usually from the application itself when it starts and finishes work.

Be aware of what it does not cover. A protected instance can still be replaced after a failed health check, terminated manually, or lost to a Spot interruption. And if every instance is protected, a scale-in simply cannot remove anything, so cost may stay high.

Take quiz
Which event can still remove a scale-in protected instance?
A normal target tracking scale-in
An AZ rebalance termination
A Spot interruption
A scheduled scale-in
Which workload is a good candidate for instance-level scale-in protection?
A stateless web server idle at night
A freshly launched empty instance
An instance in the warm pool
A worker in the middle of a long encoding job

26. What happens when an instance is put into Standby in an Auto Scaling group?

An instance in Standby is still part of the group but is taken out of service. It is deregistered from the load balancer, no longer receives traffic, and Auto Scaling stops running health checks on it, so it will not be replaced while you work.

This is the clean way to patch, debug, or inspect a live instance. When you call enter-standby, you choose whether to decrement desired capacity. If you do, the group will not launch a substitute. If you do not, it launches a new instance to keep the count.

Calling exit-standby puts the instance back in service and re-registers it. You keep paying for a standby instance, so do not leave it there for weeks.

Take quiz
What does Auto Scaling do with health checks for a Standby instance?
It runs them twice as often
It terminates the instance immediately
It moves the instance to another group
It does not run them, so the instance is not replaced
Why would you decrement desired capacity when entering Standby?
To avoid launching a replacement instance
To delete the launch template
To increase the max capacity
To reset the cooldown timer

27. What happens when you suspend an Auto Scaling process?

Auto Scaling does its job through named internal processes. Suspending one stops that behavior until you resume it, while the others keep running.

Process Effect when suspended
Launch No new instances are launched
Terminate No instances are terminated
HealthCheck Health checks stop being evaluated
ReplaceUnhealthy Unhealthy instances are not replaced
AZRebalance Zone balancing stops
AlarmNotification Alarm-driven scaling is ignored
ScheduledActions Scheduled scaling does not run
AddToLoadBalancer New instances are not registered with the LB

Teams suspend AZRebalance or ReplaceUnhealthy during troubleshooting or deployments. The risk is forgetting to resume. A group with Launch suspended will not heal itself, which can quietly cause an outage.

Take quiz
What is the effect of suspending the ReplaceUnhealthy process?
Unhealthy instances are left in place instead of being replaced
All instances are terminated
Scheduled actions are deleted
The launch template is locked
What is the biggest operational risk of leaving the Launch process suspended?
The group deletes itself
The group cannot heal or scale out
Instance metadata becomes public
Billing doubles automatically

28. How do you scale an Auto Scaling group based on SQS queue depth?

Queue length alone is a poor scaling signal, because the right number of workers depends on how many instances you already have. The recommended approach is a backlog per instance metric used with a target tracking policy.

The formula is ApproximateNumberOfMessagesVisible / number of InService instances. You can build it with CloudWatch metric math directly in the policy, so no Lambda is needed.

Then work out the target. If a worker handles one message in 10 seconds and you accept a 10-minute wait, the target backlog per instance is 600 / 10 = 60 messages. Auto Scaling then adds or removes instances to hold that ratio.

Pair it with a sensible minimum so the first message is picked up quickly, and use scale-in protection or lifecycle hooks so a worker is not killed halfway through a message.

Take quiz
Workers take 5 seconds per message and the accepted wait is 300 seconds. What is the target backlog per instance?
15 messages
60 messages
300 messages
5 messages
Why is raw queue length a weak scaling metric by itself?
It is not available in CloudWatch
It only updates once a day
It ignores how many workers are already running
It cannot be graphed

29. Explain the lifecycle of an instance in an Auto Scaling group?

Every instance moves through a defined set of states from launch to termination, and lifecycle hooks can pause it at two of the transitions.

flowchart LR
  A[Pending] --> B["Pending:Wait (launch hook)"]
  B --> C["Pending:Proceed"]
  C --> D[InService]
  D --> E[Terminating]
  E --> F["Terminating:Wait (terminate hook)"]
  F --> G["Terminating:Proceed"]
  G --> H[Terminated]
  D --> I[EnteringStandby]
  I --> J[Standby]
  J --> D
  D --> K[Detaching]
  K --> L[Detached]
  • Pending: the instance is being launched. If a launch hook exists, it waits in Pending:Wait until the hook completes.
  • InService: running, registered with the load balancer, and counted as capacity.
  • Standby / Detached: manually removed from service or from the group.
  • Terminating: being shut down, optionally held in Terminating:Wait for cleanup, then Terminated.

Only InService instances count toward the group's healthy capacity.

Take quiz
In which state does an instance count as healthy serving capacity?
Pending:Wait
Terminating:Wait
InService
Detached
Where does an instance sit while a termination lifecycle hook runs?
Pending:Proceed
EnteringStandby
InService
Terminating:Wait

30. When should you use a warm pool in Auto Scaling?

Use a warm pool when new instances take a long time to become useful, for example a large AMI, heavy user data, or a slow application bootstrap. Scale-out then pulls an already-initialized instance from the pool instead of starting from scratch.

Pool instances can be kept in one of three states:

State Compute cost Notes
Stopped None (you still pay for EBS) Cheapest, resume takes a short while
Running Full instance price Fastest, but you pay for idle capacity
Hibernated None (RAM saved to EBS) Resumes with in-memory state intact

A launch lifecycle hook is commonly used to run the long initialization once, while the instance is in the pool. You can also let instances return to the pool on scale-in so they can be reused. If your app starts in under a minute, a warm pool adds complexity without much benefit.

Take quiz
Which warm pool state avoids paying instance compute charges while still pre-initializing?
Running
Pending
Detached
Stopped
A warm pool is most useful when:
New instances need a long initialization time
The app starts in two seconds
The group never scales
You use only one Availability Zone

31. How does instance refresh work in Auto Scaling?

Instance refresh replaces the instances in an ASG in rolling batches so that every instance picks up a new launch template version, a new AMI, or changed settings. You trigger it once, and Auto Scaling handles the replacement order.

aws autoscaling start-instance-refresh \
  --auto-scaling-group-name web-asg \
  --preferences '{"MinHealthyPercentage":90,"InstanceWarmup":120,"AutoRollback":true}'

The main preferences to know:

  • MinHealthyPercentage: how much of the group must stay healthy during the roll (90% by default).
  • InstanceWarmup: how long to wait before a replacement counts as healthy.
  • Checkpoints: pause at chosen percentages, such as 20% and 50%, so you can verify before continuing.
  • AutoRollback: revert to the previous configuration if the refresh fails or an alarm fires.
  • SkipMatching: skip instances that already match the desired configuration.

Instance refresh replaces instances in place within the same group, so it suits routine AMI updates. For a safer cutover with instant rollback, a blue/green approach with two groups is more suitable.

Take quiz
Which instance refresh option reverts the group if the rollout goes wrong?
AutoRollback
SkipMatching
ScaleInProtectedInstances
DefaultCooldown
What does MinHealthyPercentage control during an instance refresh?
How many instances are added permanently
How much capacity must stay healthy while instances are replaced
The size of the new instances
How often alarms are evaluated

32. Why use a mixed instances policy in an Auto Scaling group?

A mixed instances policy lets one ASG use several instance types and combine On-Demand with Spot capacity. The main reasons are lower cost and better resilience to capacity shortages.

You control the blend with a few settings:

  • On-Demand base capacity: a fixed number of On-Demand instances that always run.
  • On-Demand percentage above base: the share of extra capacity that stays On-Demand, with the rest on Spot.
  • Instance type overrides: a list such as m5.large, m5a.large, m6i.large, c5.xlarge.
  • Allocation strategy: for Spot, price-capacity-optimized is the usual recommendation.

If one instance type runs out in a zone, the group can launch another, so you are less likely to hit InsufficientInstanceCapacity. It requires a launch template.

Take quiz
What does On-Demand base capacity guarantee?
A hard cap on Spot price
A fixed number of On-Demand instances always running
A minimum of one Spot instance
A reserved IP for each instance
Which benefit is specific to listing several instance types in the policy?
Automatic AMI patching
Free data transfer
Lower risk of insufficient capacity errors
Elimination of health checks

33. How do you use Spot Instances safely in an Auto Scaling group?

Spot Instances can be reclaimed with a two-minute warning, so the goal is to make an interruption boring rather than rare.

  1. Diversify: use many instance types and several Availability Zones so no single capacity pool dominates.
  2. Choose a good allocation strategy: price-capacity-optimized picks pools that are both cheap and less likely to be interrupted.
  3. Keep an On-Demand base for the minimum capacity your service cannot lose.
  4. Enable Capacity Rebalancing: the group launches a replacement as soon as AWS signals a Spot instance is at elevated risk, and then drains the old one.
  5. React to the notice: use a termination lifecycle hook or the EventBridge interruption event to drain connections and checkpoint work.
  6. Stay stateless: store sessions and results outside the instance.

Spot works best for stateless web tiers, queue workers, CI runners, and batch jobs. It is a poor fit for a single-node database.

Take quiz
How much warning does AWS give before reclaiming a Spot Instance?
Thirty minutes
One hour
Two minutes
No warning at all
What does Capacity Rebalancing do?
Converts Spot to Reserved Instances
Moves instances to another Region
Raises the Spot bid price
Launches a replacement when a Spot instance is at elevated interruption risk

34. What is attribute-based instance type selection?

Attribute-based instance type selection (ABIS) lets you describe the instance requirements instead of listing instance types by name. Auto Scaling then chooses every matching type available in your Region.

You specify limits such as vCPU count, memory size, CPU architecture, and instance generation, and you can exclude families you do not want.

"InstanceRequirements": {
  "VCpuCount": {"Min": 2, "Max": 4},
  "MemoryMiB": {"Min": 8192},
  "CpuManufacturers": ["intel", "amd"],
  "ExcludedInstanceTypes": ["t2.*"]
}

The big win is future proofing: when AWS releases a new instance type that fits your rules, the group can use it without anyone editing the template. It also widens the pool of Spot capacity. It works with mixed instances policies, so you can still define the On-Demand and Spot split.

Take quiz
What do you provide with attribute-based selection instead of a list of instance types?
A list of IAM roles
A VPC peering ID
A CloudFront distribution
Requirements such as vCPU and memory ranges
Why does attribute-based selection help over time?
New matching instance types can be used without editing the template
It freezes the instance type forever
It disables Spot interruptions
It removes the need for AMIs

35. How do you handle a graceful shutdown during Auto Scaling scale-in?

Scale-in terminates instances, so the application needs a window to finish what it is doing. Several mechanisms combine to give it one.

  1. Deregistration delay: the target group stops sending new requests and gives in-flight ones time to complete (default 300 seconds).
  2. Termination lifecycle hook: the instance waits in Terminating:Wait while a script or Lambda drains work, flushes logs, or hands off a lease.
  3. Heartbeat: if cleanup takes longer, call record-lifecycle-action-heartbeat to extend the timeout.
  4. Complete the action: call complete-lifecycle-action with CONTINUE so termination proceeds.

For workers handling long jobs, add scale-in protection while a job is running and clear it afterward. Also make sure your process handles SIGTERM properly, since the OS signal is what actually reaches the application.

Take quiz
Which feature holds an instance in Terminating:Wait for cleanup?
A termination lifecycle hook
A scheduled action
A launch template version
A cooldown
What extends a lifecycle hook when cleanup takes longer than expected?
Raising max capacity
Recording a lifecycle action heartbeat
Detaching the instance
Disabling health checks

36. How do you troubleshoot an Auto Scaling group that fails to launch instances?

Start with the Activity history of the group (console or describe-scaling-activities). Each failed launch has a status message that usually names the cause.

Typical message or symptom Likely cause Fix
InsufficientInstanceCapacity No capacity for that type in that AZ Add instance types or AZs, use mixed instances
VcpuLimitExceeded Account vCPU quota reached Request a quota increase
Invalid AMI or snapshot not found AMI deregistered or not shared Update the launch template
Security group does not exist Group belongs to a different VPC Use a group from the ASG's VPC
Client.InternalError with encrypted EBS Auto Scaling role lacks KMS key access Grant the service-linked role key permissions
No free IPs in subnet Subnet CIDR exhausted Add larger or more subnets

Also check that desired capacity is not already at the maximum, that the instance profile can be passed, and that the key pair still exists. If launches keep failing for a long time, Auto Scaling backs off and may stop retrying, so fix the cause and verify that the group has resumed.

Take quiz
Where should you look first when instances will not launch?
The billing dashboard
The group's Activity history
The S3 access logs
The Route 53 query log
Which message suggests picking more instance types or Availability Zones?
VcpuLimitExceeded
InvalidKeyPair.NotFound
InsufficientInstanceCapacity
SecurityGroup.NotFound

37. Why are instances terminated right after launch in an Auto Scaling group?

If instances launch fine and then disappear within minutes, Auto Scaling is almost always replacing them as unhealthy. Check the activity history for a cause such as "Health check failed" or "ELB health check failed".

The usual suspects:

  • Grace period too short: the load balancer checks before the app is ready.
  • Wrong health check path or port: the target group probes a URL that returns 404 or 500.
  • Security group blocks the ALB: the instance cannot receive health check traffic.
  • User data failed: the application never started because a script errored out.
  • Encrypted volume problem: the instance cannot start because of missing KMS permissions.
  • Lifecycle hook abandoned: a hook timed out with ABANDON as the default result.

To investigate, temporarily put an instance in Standby or suspend ReplaceUnhealthy, then connect and read /var/log/cloud-init-output.log and the application logs.

Take quiz
Which setting is most likely wrong if a healthy app is killed during boot?
Max capacity is too high
The cooldown is too long
Health check grace period is too short
Predictive scaling is on
How can you keep a failing instance alive long enough to debug it?
Delete the launch template
Raise the desired capacity
Turn off CloudWatch
Suspend ReplaceUnhealthy or move it to Standby

38. When should you use scheduled scaling together with dynamic scaling?

Use them together when you know a surge is coming but cannot predict its exact size. Dynamic scaling alone reacts after traffic arrives, and a fresh instance needs minutes to boot. Scheduled scaling closes that gap.

A practical pattern: at 8:00 AM, a scheduled action raises the minimum capacity from 4 to 12. Dynamic scaling is still active and can go above 12 if load demands it, and can come back down to 12 but not below. At 7:00 PM, another scheduled action lowers the minimum to 4.

Raising the minimum, not the desired value, is the key trick. Dynamic policies keep working on top of that floor, whereas a scheduled change to desired alone can be overwritten by the next policy evaluation.

For launches, sales events, or school-term traffic, schedule extra headroom 15 to 30 minutes early. If the pattern recurs reliably, consider predictive scaling instead of maintaining many schedules.

Take quiz
Which value should a scheduled action raise so dynamic scaling can still go higher?
The health check grace period
The cooldown
The instance warmup
Minimum capacity
Why does scheduled scaling help before a known traffic surge?
Instances are already running when the load arrives
It removes the need for load balancers
It makes instances boot faster
It disables CloudWatch alarms

39. What happens when multiple scaling policies trigger at the same time?

An ASG can have several policies, for example one tracking CPU and another tracking request count. If more than one wants to change capacity at once, Auto Scaling resolves it by choosing the option that gives the largest capacity.

  • Scale-out: Policy A wants +1 and policy B wants +3, so the group adds 3.
  • Scale-in: Policy A wants -1 and policy B wants -2, so the group removes only 1, because that leaves more capacity.
  • Scale-out vs scale-in: scale-out wins over scale-in.

For target tracking, there is an extra rule: the group scales in only when all target tracking policies are ready to scale in. This prevents one metric from removing capacity that another metric still needs.

The practical effect is that the group favors availability over savings. Overlapping policies are safe, but each additional policy can slow scale-in.

Take quiz
Two policies request +1 and +3 instances. What does the group do?
Adds 3 instances
Adds 1 instance
Adds 4 instances
Ignores both requests
When does a group with several target tracking policies scale in?
When any one of them agrees
Only when all of them agree to scale in
Only at midnight
Never, it can only scale out

40. How do you test predictive scaling safely before enabling it?

Start in forecast only mode. Auto Scaling builds the forecast and shows it next to actual load, but does not launch anything. You get a risk-free way to judge accuracy.

  1. Create the predictive policy with mode ForecastOnly after the group has at least a day of data (a week or more gives better results).
  2. Open the policy's graphs and compare forecasted load and forecasted capacity with what actually happened over several days.
  3. Check that the chosen scaling metric (for example CPU) really tracks the load metric (for example request count). A weak correlation gives poor forecasts.
  4. Review how the forecast treats your maximum capacity, and decide whether it may exceed it using MaxCapacityBreachBehavior.
  5. Switch to ForecastAndScale and tune SchedulingBufferTime so instances launch early enough to finish booting.

Keep your dynamic policies attached. Predictive scaling covers the pattern, and dynamic scaling still covers the surprise.

Take quiz
Which mode generates forecasts without changing capacity?
ForecastAndScale
ForecastOnly
SuspendOnly
StandbyMode
What does SchedulingBufferTime help with?
Delaying scale-in by a day
Setting the maximum instance price
Launching instances early enough to finish booting before the forecast load
Selecting a different Region

41. How do you perform a blue/green deployment with Auto Scaling groups?

The idea is to run the new version in a second ASG next to the current one, shift traffic gradually, and keep the old group around until you are confident.

flowchart LR
  U[Users] --> L["ALB listener with weighted rules"]
  L -- 90 percent --> TB["Target group blue"]
  L -- 10 percent --> TG["Target group green"]
  TB --> AB["ASG blue - v1"]
  TG --> AG["ASG green - v2"]
  1. Create a launch template version with the new AMI and a new ASG (green) attached to its own target group.
  2. Wait until green instances pass health checks.
  3. Change the ALB listener to weighted forwarding, for example 90/10, then 50/50, then 0/100 while watching error rate and latency.
  4. If metrics degrade, set the weights back to blue. Rollback takes seconds because blue is still running.
  5. After a bake period, scale blue down to zero or delete it.

The trade-off is paying for roughly double capacity during the switch. If that is too costly and a brief rolling change is acceptable, instance refresh with checkpoints and auto rollback is simpler. AWS CodeDeploy can also automate the blue/green flow for ASGs.

Take quiz
How is traffic shifted between blue and green behind an ALB?
By editing the launch template cooldown
By renaming the instances
By changing weights on the listener rule's target groups
By changing the IAM role
What is the main advantage of blue/green over instance refresh?
It uses half the capacity
It needs no load balancer
It avoids creating a new AMI
Near-instant rollback because the old group is still running

42. How can you optimize Auto Scaling costs?

Cost in an ASG comes from how many instances run, how long, and at what price. Work on each of those three.

  • Run fewer instance hours: use target tracking with a realistic target (60 to 70% CPU is often safe, 30% wastes money) and scheduled scale-in for nights and weekends.
  • Pay a lower price: use a mixed instances policy with Spot for the flexible share and an On-Demand base for the floor.
  • Cover the baseline: buy Savings Plans or Reserved Instances for the minimum capacity that always runs.
  • Right-size the instance: review utilization, and consider Graviton-based types if your software supports ARM.
  • Scale in faster where safe: shorten instance warmup and cooldown values to your real startup time, so excess capacity is released sooner.
  • Use warm pools carefully: prefer the Stopped state over Running.

Set a sensible maximum capacity and a CloudWatch billing alarm too, so a bug or traffic attack cannot scale you into a surprise bill.

Take quiz
Which approach reduces cost for the capacity that always runs?
Raising the max capacity
Disabling scale-in
Using only the largest instance type
Savings Plans or Reserved Instances for the baseline
Why set a realistic target value in target tracking?
A very low target keeps too many instances running
It disables health checks
It increases the Spot price
It changes the AMI

43. What is an instance maintenance policy in Auto Scaling?

An instance maintenance policy controls how much capacity an ASG keeps during replacements that Auto Scaling starts, such as replacing an unhealthy instance or one that has reached its maximum lifetime. You set it by choosing a minimum and maximum healthy percentage.

Behavior Min healthy Max healthy Effect
Launch before terminating 100% 110% or more A new instance starts first, so capacity never dips, at a small temporary cost
Terminate and launch 90% or less 100% The old one is removed first, so cost stays flat, but capacity may drop
Custom Your choice Your choice Balance between the two

The max cannot be more than 100 percentage points above the min. Without a policy, the default behavior is to terminate first and then launch.

Choose launch before terminating for latency-sensitive services where losing even one instance hurts. Choose terminate-and-launch for batch groups where a brief dip is fine.

Take quiz
What does a 100% minimum healthy and 110% maximum healthy policy do?
Launches a replacement before terminating the old instance
Terminates first, then launches
Prevents all replacements
Doubles the group size permanently
Which kind of workload benefits most from launch-before-terminate?
A nightly batch job tolerant of dips
A latency-sensitive web tier
A dev environment used once a week
An idle test group with min 0

44. How do you scale Amazon ECS services and capacity with Auto Scaling?

ECS has two layers, and each one scales separately.

Layer What scales Service used
Service Number of tasks Application Auto Scaling (target tracking on ECSServiceAverageCPUUtilization, ECSServiceAverageMemoryUtilization or ALBRequestCountPerTarget)
Cluster capacity EC2 instances hosting the tasks EC2 Auto Scaling through an ECS capacity provider

For EC2-backed clusters, create an ASG capacity provider with managed scaling turned on. ECS publishes a CapacityProviderReservation metric and creates a target tracking policy to keep it at your target (commonly 100%), which means "just enough instances for the tasks that need to run".

Turn on managed termination protection as well. ECS then protects instances that have running tasks from scale-in, which prevents tasks from being killed mid-flight.

With Fargate there are no instances to manage, so only the service-level task scaling applies.

Take quiz
Which ECS feature lets Auto Scaling add EC2 instances when tasks cannot be placed?
A task definition revision
A capacity provider with managed scaling
An ECR lifecycle policy
A CloudMap namespace
What does managed termination protection prevent?
Creation of new tasks
Deregistration from the cluster console
Scale-in of instances that are running tasks
Use of Spot Instances

45. How do you scale DynamoDB capacity automatically?

For tables in provisioned mode, use Application Auto Scaling. It adjusts read and write capacity units for the table and each global secondary index to follow actual usage.

  1. Register the table (or index) as a scalable target with a minimum and maximum capacity.
  2. Attach a target tracking policy for read or write usage, with a target utilization such as 70%.
  3. Application Auto Scaling creates the alarms and raises or lowers provisioned capacity when consumption drifts from the target.
aws application-autoscaling register-scalable-target \
  --service-namespace dynamodb \
  --resource-id table/Orders \
  --scalable-dimension dynamodb:table:WriteCapacityUnits \
  --min-capacity 5 --max-capacity 500

It reacts to consumed capacity over a few minutes, so very sharp spikes can still throttle before scaling catches up. For unpredictable or spiky traffic, on-demand mode is the simpler choice. For known peaks, add scheduled actions to the same scalable target.

Take quiz
Which service handles automatic capacity changes for provisioned DynamoDB tables?
EC2 Auto Scaling
Amazon Macie
Application Auto Scaling
AWS Glue
Which table mode avoids capacity planning for highly spiky traffic?
Provisioned mode with fixed units
Reserved capacity only
Streams-enabled mode
On-demand mode

46. How do you get notified about Auto Scaling events?

There are three main channels, and they suit different needs.

Channel What it gives you Typical use
SNS notifications on the group Messages for launch, terminate, and their failures Email or chat alerts for the team
EventBridge events Structured events like EC2 Instance Launch Successful and EC2 Instance Terminate Unsuccessful Trigger Lambda, update a CMDB, open a ticket
CloudWatch alarms on group metrics Alerts based on counts such as desired versus in-service capacity Detect being stuck at max or having too few healthy hosts

A solid minimum is an alarm on failed launches and one on desired capacity equal to max for a prolonged period. The first catches broken templates or quota problems, the second tells you traffic is being capped. Lifecycle hook events also arrive through EventBridge, which is how automation reacts to instances entering Pending:Wait.

Take quiz
Which channel triggers a Lambda function when an instance launch fails?
CloudTrail Insights only
Route 53 health checks
AWS Artifact reports
EventBridge events from Auto Scaling
What does an alarm on desired capacity staying at maximum usually signal?
Traffic demand is being capped by the max setting
The AMI is corrupted
A key pair has expired
The VPC has no route table

47. How do you secure instances launched by an Auto Scaling group?

Because every instance comes from the same template, security is mostly decided once in the launch template and the surrounding IAM setup.

  • Instance profile: give instances a least-privilege IAM role, not broad admin access.
  • IMDSv2 required: set HttpTokens to required in the template's metadata options to block SSRF-style credential theft.
  • Encryption: enable EBS encryption and make sure the KMS key policy allows the Auto Scaling service-linked role, AWSServiceRoleForAutoScaling.
  • Network: launch into private subnets with no public IP, and let security groups allow only the load balancer.
  • Access: use Systems Manager Session Manager instead of open SSH ports and shared keys.
  • Patching: bake updates into a new AMI and roll it out with instance refresh.

The people creating the group also need iam:PassRole for the instance profile, so limit who holds that permission.

Take quiz
Which launch template setting blocks IMDSv1 and requires token-based metadata access?
HttpTokens set to required
AssociatePublicIpAddress set to true
Monitoring set to detailed
DeleteOnTermination set to false
Why must the KMS key policy include the Auto Scaling service-linked role?
So CloudWatch can graph CPU
So instances with encrypted EBS volumes can launch
So the ALB can resolve DNS
So Spot prices are visible

48. How do you design stateful applications to run in Auto Scaling groups?

Auto Scaling treats instances as disposable, so the safest design is to keep state off the instance. Anything that must survive a replacement should live somewhere else.

  • Sessions: store them in ElastiCache or DynamoDB instead of local memory.
  • Files and uploads: use S3 for objects, or EFS when many instances need a shared file system.
  • Databases: use RDS, Aurora, or DynamoDB, not a database on the ASG instance.
  • Configuration: pull from Parameter Store or Secrets Manager at boot.

Sometimes you truly need per-node data, for example a clustered service with its own EBS volume. A workable pattern is one ASG per Availability Zone with min and max both set to 1, plus a launch lifecycle hook that attaches the persistent EBS volume to the new instance before it goes InService. You get automatic healing without losing the data.

Sticky sessions on the load balancer can help short term, but they make scale-in disruptive and should not replace external session storage.

Take quiz
Where should user session data live in an Auto Scaling design?
In the instance's memory only
In ElastiCache or DynamoDB
In the AMI
In the launch template
What does an ASG with min 1 and max 1 plus a lifecycle hook that attaches an EBS volume provide?
Unlimited horizontal scaling
Free Spot capacity
Self-healing for a single node that keeps its data
Automatic multi-Region failover

49. Explain the execution flow of a scale-out event with target tracking?

Target tracking turns a metric target into instance launches through a short chain of automated steps.

sequenceDiagram
  participant I as Instances
  participant CW as CloudWatch
  participant AS as Auto Scaling
  participant EC2 as EC2
  participant LB as Load Balancer
  I->>CW: Publish metric (CPU 75%)
  CW->>CW: AlarmHigh breaches target 50%
  CW->>AS: Alarm triggers scaling policy
  AS->>AS: Compute new desired capacity
  AS->>EC2: Launch new instances
  EC2->>AS: Instances reach InService
  AS->>LB: Register with target group
  Note over AS: Warmup excludes new instances from metrics
  1. Instances publish their metric to CloudWatch (one-minute data with detailed monitoring).
  2. The alarm that the policy created goes into ALARM after a few consecutive breaching data points, while the scale-in alarm needs far more, so the group scales out fast and in slowly.
  3. Auto Scaling calculates the capacity that would bring the metric back to the target. With 4 instances at 75% CPU and a 50% target, it needs about 4 x 75 / 50 = 6 instances, so it adds 2.
  4. EC2 launches instances from the launch template, any launch hook runs, and the instances are registered with the load balancer.
  5. During instance warmup, the new instances are not counted in the aggregate, which prevents another round of over-scaling.

The new capacity is capped by the group's maximum, and the metric settles back near the target once the new instances share the load.

Take quiz
Four instances run at 75% average CPU with a 50% target. About how many instances does target tracking aim for?
5 instances
8 instances
6 instances
4 instances
Why does the scale-in alarm need more data points than the scale-out alarm?
Because CloudWatch cannot evaluate lower values
Because scale-in launches instances
Because it uses a different Region
To scale out quickly but scale in conservatively

50. What happens when an Availability Zone becomes impaired for an Auto Scaling group?

Instances in the impaired zone start failing health checks. Auto Scaling marks them unhealthy, terminates them, and launches replacements in the remaining healthy zones of the group, so the instance count recovers without anyone acting.

flowchart TD
  A["AZ-a becomes impaired"] --> B["Instances fail ELB or EC2 health checks"]
  B --> C["Marked unhealthy and terminated"]
  C --> D["Replacements launched in AZ-b and AZ-c"]
  D --> E["AZ-a recovers"]
  E --> F["AZRebalance evens out the zones"]

A few things decide how smooth this is:

  • Headroom: if you run 3 zones, size the group so that losing one still leaves enough capacity. Otherwise the survivors are overloaded until replacements are ready.
  • Subnets in several zones: an ASG with a single AZ cannot fail over.
  • ELB health checks: without them, the load balancer may keep sending traffic to dead targets.
  • AZRebalance: after recovery, it gradually moves capacity back, launching before terminating.

For a prolonged, partial impairment you can also remove the affected subnet from the group. Auto Scaling then replaces its instances in the other zones.

Take quiz
Where does Auto Scaling launch replacements for instances lost in an impaired zone?
In a different AWS Region
Only in the same impaired zone
It does not launch replacements
In the group's remaining healthy Availability Zones
Which design choice makes zone failure survivable?
Sizing capacity so surviving zones can carry the load
Using a single large subnet
Disabling health checks
Setting min capacity to zero
«
»

Comments & Discussions