Prev Next

Cloud / AWS Fargate Interview questions

Last updated

1. What is AWS Fargate? 2. What container orchestrators does Fargate work with? 3. What is a task definition in Fargate? 4. What CPU and memory combinations does Fargate support? 5. What is the awsvpc network mode in Fargate? 6. What is a Fargate platform version? 7. What is the difference between task role and task execution role? 8. What are the ephemeral storage limits for Fargate tasks? 9. What is Fargate Spot? 10. How is AWS Fargate priced? 11. What is an ECS service on Fargate? 12. How do you run a standalone task on Fargate? 13. What CPU architectures does Fargate support? 14. What operating systems does Fargate support? 15. How do you view container logs for Fargate tasks? 16. What is ECS Exec and how do you use it? 17. What is a capacity provider strategy? 18. How do you pass secrets to Fargate containers? 19. How do you define a Fargate service using Terraform? 20. What happens when an essential container in a Fargate task exits? 21. What is the difference between Fargate and EC2 launch types? 22. When should you choose Fargate over EC2? 23. When would you choose EC2 over Fargate? 24. Explain the lifecycle of a Fargate task? 25. Explain the execution flow of launching a Fargate task? 26. How does Fargate isolate tasks from each other? 27. How do you troubleshoot CannotPullContainerError on Fargate? 28. How do you troubleshoot a Fargate task stuck in PENDING? 29. How do you troubleshoot exit code 137 on Fargate? 30. How does ECS service auto scaling work with Fargate? 31. How can you reduce AWS Fargate costs? 32. How do Savings Plans apply to Fargate? 33. How do you handle Fargate Spot interruptions? 34. How do you attach persistent storage to Fargate tasks? 35. How do you attach an EBS volume to a Fargate task? 36. How do you load balance traffic to Fargate tasks? 37. What is the difference between Service Connect and Cloud Map discovery? 38. How do rolling deployments work for Fargate services? 39. How does the ECS deployment circuit breaker work? 40. How do you configure health checks for Fargate tasks? 41. How do you secure Fargate workloads? 42. Why can't Fargate run privileged containers? 43. How does Fargate encrypt ephemeral storage? 44. How can you speed up Fargate container startup? 45. How do you monitor Fargate tasks? 46. How do you use FireLens for logging on Fargate? 47. How do Fargate profiles work in EKS? 48. What are the limitations of Fargate on EKS? 49. How do you run scheduled jobs on Fargate? 50. How would you architect a highly available Fargate service?

1. What is AWS Fargate?

AWS Fargate is a serverless compute engine for containers. You declare the CPU and memory a workload needs, and AWS finds the capacity, runs the container, and removes it when it stops. There are no EC2 instances to choose, patch, or scale.

It is not an orchestrator on its own. Fargate is the compute layer behind Amazon ECS and Amazon EKS, which decide what runs and when. Each task or pod runs in its own isolated environment instead of sharing a kernel with other workloads.

You pay for the vCPU and memory a task requests, not for idle servers. The trade-off is less control: no SSH access to hosts, no privileged containers, and no custom AMIs.

Take quiz
Which part of container infrastructure does Fargate take off your hands?
Writing the Dockerfile for your application
Choosing which image tag to deploy
Writing your application's IAM policies
Provisioning, patching, and scaling the underlying servers
What does Fargate need in order to decide which containers run and when?
Nothing, it schedules containers by itself
An orchestrator such as ECS or EKS
An EC2 Auto Scaling group
One Lambda function per container

2. What container orchestrators does Fargate work with?

Fargate works with Amazon ECS and Amazon EKS. AWS Batch can also run jobs on Fargate. The way you ask for Fargate differs for each.

Service How you select Fargate Unit of work
Amazon ECS Launch type FARGATE or the FARGATE / FARGATE_SPOT capacity providers Task
Amazon EKS A Fargate profile whose selectors match the pod's namespace and labels Pod
AWS Batch A compute environment of type FARGATE or FARGATE_SPOT Job

Self-managed Kubernetes and Docker Swarm cannot use Fargate as a compute backend.

Take quiz
In EKS, what makes a pod run on Fargate?
A Fargate profile whose selectors match the pod's namespace and labels
A nodeSelector that names an instance type
The container image name in the pod spec
A DaemonSet running on every worker node
What is the unit of work Fargate runs when it is used with ECS?
A pod
A node group
A task
A Lambda invocation

3. What is a task definition in Fargate?

A task definition is the JSON blueprint ECS uses to launch a task. It lists the containers, images, ports, environment variables, log settings, and IAM roles. Revisions are grouped under a family such as orders-api:7.

Fargate adds a few mandatory settings: networkMode must be awsvpc, requiresCompatibilities must include FARGATE, and cpu and memory must be set at the task level.

{
  "family": "orders-api",
  "networkMode": "awsvpc",
  "requiresCompatibilities": ["FARGATE"],
  "cpu": "512",
  "memory": "1024",
  "executionRoleArn": "arn:aws:iam::111122223333:role/ecsTaskExecutionRole",
  "containerDefinitions": [{
    "name": "api",
    "image": "111122223333.dkr.ecr.us-east-1.amazonaws.com/orders-api:1.4.2",
    "essential": true,
    "portMappings": [{"containerPort": 8080}]
  }]
}

Revisions are immutable. Editing a definition registers a new revision, and running tasks keep the revision they started with until you update the service.

Take quiz
On Fargate, where must CPU and memory be set?
Only per container, never per task
At the task level in the task definition
In the ECS cluster settings
On the service's load balancer
What happens when you change a registered task definition?
Running tasks restart automatically with the change
The existing revision is overwritten in place
The change is rejected until the family is deleted
A new revision is created and the old one stays unchanged

4. What CPU and memory combinations does Fargate support?

Fargate only accepts fixed pairings of task-level vCPU and memory. Each CPU size has its own memory range, so you cannot mix them freely.

Task vCPU Memory range (Linux) Step
0.25 0.5, 1, 2 GB fixed values
0.5 1 - 4 GB 1 GB
1 2 - 8 GB 1 GB
2 4 - 16 GB 1 GB
4 8 - 30 GB 1 GB
8 16 - 60 GB 4 GB
16 32 - 120 GB 8 GB

In the task definition, CPU is written in units (256, 512, 1024 and so on) and memory in MiB. An invalid pair fails at registration. Windows tasks start at 1 vCPU and top out at 4 vCPU.

You are billed for the size you configure, so pick the smallest combination that covers peak usage.

Take quiz
Which task size is valid on Linux Fargate?
0.25 vCPU with 8 GB
2 vCPU with 2 GB
1 vCPU with 4 GB
4 vCPU with 4 GB
How is a Fargate task billed relative to its size?
For the vCPU and memory configured, not the amount used
For the average CPU actually consumed
For the peak memory seen in CloudWatch
For memory only, since CPU is free

5. What is the awsvpc network mode in Fargate?

awsvpc gives every task its own elastic network interface (ENI) with a private IP from your subnet and its own security groups. It is the only network mode Fargate supports.

Containers in the same task share that interface, so they reach each other over localhost. Port mappings also stay simple, because each task has its own IP and there is no host port to collide on.

Two practical consequences follow:

  • Every running task consumes one IP address, so subnets must be sized for peak task count including rolling deployments.
  • In a public subnet the task needs assignPublicIp=ENABLED to reach the internet. In a private subnet it needs a NAT gateway or VPC endpoints.

Load balancers register tasks by IP, so target groups must use target type ip.

Take quiz
In awsvpc mode, how do containers in the same Fargate task talk to each other?
Through the Docker bridge on the host
Only through the load balancer
By using each other's public IPs
Over localhost, because they share the task's network interface
Why does subnet sizing matter for Fargate services?
All tasks in a subnet share one IP address
Every task takes a private IP from the subnet for its ENI
Subnets only matter for the EC2 launch type
Each task needs its own Elastic IP

6. What is a Fargate platform version?

A platform version identifies the runtime environment a Fargate task runs on: the kernel, container runtime, and ECS agent. It decides which features the task can use.

Linux and Windows have independent numbering. The LATEST Linux version is 1.4.0 and the LATEST Windows version is 1.0.0. Version 1.4.0 is the one that brought per-task ENIs, EFS support, and configurable ephemeral storage.

AWS patches the platform by publishing new revisions of a version. A running task is never patched in place. New tasks start on the latest revision, and old ones are retired and replaced. Services handle this automatically, but standalone tasks you launched yourself have to be started again.

Take quiz
What is the LATEST Linux platform version on Fargate?
1.4.0
1.0.0
2.0.0
1.3.0
How does a running Fargate task receive a platform security patch?
AWS patches the running task in place
You SSH in and run the package manager
It is retired and replaced by a new task on the patched revision
The task restarts itself nightly

7. What is the difference between task role and task execution role?

Both are IAM roles on a task, but they serve different callers. The execution role is used by ECS and Fargate to set the task up. The task role is used by your application code at runtime.

Task execution role Task role
Used by ECS agent / Fargate infrastructure Your application inside the container
Typical permissions Pull from ECR, write to CloudWatch Logs, read Secrets Manager / SSM values for the secrets block Read an S3 bucket, write to DynamoDB, publish to SQS
When it is used Before and during container start While the code is running
Task definition field executionRoleArn taskRoleArn

A frequent mistake is putting S3 or DynamoDB permissions on the execution role. The application never uses it, so the calls fail with AccessDenied.

Take quiz
Which role does the application code inside the container use to call DynamoDB?
The task execution role
The task role
The ECS service-linked role
The EC2 instance profile of the host
Which role needs permission to pull the image from a private ECR repository?
The task role
The role attached to the load balancer
The IAM user who ran the deployment
The task execution role

8. What are the ephemeral storage limits for Fargate tasks?

Every Fargate task gets 20 GiB of ephemeral storage by default. On Linux platform version 1.4.0 or later you can raise it to a maximum of 200 GiB with the ephemeralStorage setting.

"ephemeralStorage": {
  "sizeInGiB": 100
}

The space is shared by the pulled image layers, each container's writable layer, and any bind-mount volumes the containers share. It exists only for the life of the task, so everything is lost when the task stops.

Storage above the default 20 GiB is billed per GB-hour. If data must outlive the task, use EFS or an EBS volume instead of growing ephemeral storage.

Take quiz
What is the maximum ephemeral storage you can configure for a Linux Fargate task?
20 GiB
100 GiB
200 GiB
1 TiB
What happens to data in a task's ephemeral storage when the task stops?
It is lost
It is copied to S3 automatically
It moves to the next task in the service
It stays on the host for 24 hours

9. What is Fargate Spot?

Fargate Spot runs ECS tasks on spare Fargate capacity at a discount of up to about 70% compared with On-Demand. The catch is that AWS can reclaim the capacity, so a task can be interrupted.

When that happens, the task gets a two-minute warning: a SIGTERM is sent to the containers, and after the stopTimeout (up to 120 seconds on Fargate) they are killed. Your code should shut down cleanly inside that window.

Spot suits workloads that tolerate interruption: batch jobs, queue workers, CI runners, and stateless services running at the same time as On-Demand tasks. It is a poor fit for a single-task service or anything holding state in memory.

You enable it through the FARGATE_SPOT capacity provider.

Take quiz
How much warning does a Fargate Spot task get before it is stopped?
Thirty minutes by email
None, it is killed instantly
Ten seconds, delivered as SIGKILL
Two minutes, delivered as SIGTERM
Which workload is the best fit for Fargate Spot?
A single-instance primary database
A stateless queue worker that can retry a message
A task holding unsaved user sessions in memory
A service that must never lose a task

10. How is AWS Fargate priced?

Fargate bills for the vCPU and memory configured on a task, charged per second with a one-minute minimum for Linux. Billing runs from the moment the image pull starts until the task terminates, so startup time costs money too.

Several levers change the rate:

Factor Effect on price
ARM64 (Graviton) Roughly 20% lower than x86 for vCPU and memory
Fargate Spot Up to about 70% off, with interruption risk
Compute Savings Plans Discount for a 1 or 3 year hourly spend commitment
Ephemeral storage above 20 GiB Extra charge per GB-hour
Windows containers Higher rate that includes the Windows license

Remember the costs around the task as well: NAT gateway data processing, load balancers, CloudWatch Logs ingestion, and cross-AZ traffic often add up faster than people expect.

Take quiz
From which point does Fargate start billing a task?
When the container image pull begins
When the first request is served
When the service is created
When the task passes its load balancer health check
Which option lowers the Fargate rate through a time-based spend commitment?
Reserved Instances for EC2
Fargate Spot
Compute Savings Plans
A larger ephemeral storage volume

11. What is an ECS service on Fargate?

An ECS service keeps a specified number of copies of a task definition running and replaces any that fail. It is how you run long-lived workloads such as web APIs on Fargate.

The scheduler compares desiredCount with the running count. If a task crashes, fails its health check, or is retired, the scheduler starts a replacement, possibly in another Availability Zone.

A service also handles:

  • registering and deregistering tasks with an Application or Network Load Balancer
  • rolling out new task definition revisions with deployment settings
  • scaling the task count through Application Auto Scaling

Fargate only supports the REPLICA scheduling strategy. The DAEMON strategy, which places one task per instance, is not available because there are no instances to place on.

Take quiz
What does an ECS service do when one of its Fargate tasks crashes?
Leaves the count lower until you intervene
Starts a replacement to restore the desired count
Restarts the whole cluster
Emails the task owner and waits
Which service scheduling strategy does Fargate support?
DAEMON
Both REPLICA and DAEMON
Neither, services are not allowed on Fargate
REPLICA

12. How do you run a standalone task on Fargate?

Use aws ecs run-task (or the equivalent API call). It starts one or more tasks that run to completion and are not restarted, which suits batch jobs and one-off scripts.

aws ecs run-task \
  --cluster prod \
  --launch-type FARGATE \
  --task-definition report-job:3 \
  --count 1 \
  --network-configuration "awsvpcConfiguration={subnets=[subnet-0a1b2c],securityGroups=[sg-0d4e5f],assignPublicIp=DISABLED}" \
  --overrides '{"containerOverrides":[{"name":"job","command":["python","run.py","--date","2026-10-01"]}]}'

The network configuration is required because of awsvpc. The overrides block lets you change the command or environment for this run without registering a new revision.

Once the essential container exits, the task moves to STOPPED and billing ends. Check stoppedReason and the container exit code to see how it went.

Take quiz
Why must run-task include a network configuration for Fargate?
Fargate tasks cannot run without a public IP
The command needs a VPN tunnel
Each task needs subnets and security groups for its own ENI
It selects the Availability Zone's instance type
What does the overrides parameter let you do?
Change the command or environment for one run without a new revision
Raise the task above the maximum vCPU limit
Switch the task to the EC2 launch type
Skip the task execution role

13. What CPU architectures does Fargate support?

Fargate runs Linux tasks on X86_64 and ARM64 (AWS Graviton). Windows tasks run on X86_64 only.

You choose the architecture in the task definition:

"runtimePlatform": {
  "cpuArchitecture": "ARM64",
  "operatingSystemFamily": "LINUX"
}

ARM64 tasks are typically about 20% cheaper at the same size, and many workloads run as fast or faster. The image must match, though. If an x86 image is launched on an ARM64 task, the container fails with an exec format error.

Build multi-architecture images with docker buildx build --platform linux/amd64,linux/arm64 so one tag works for both, and check that native dependencies have ARM builds.

Take quiz
What error typically appears when an x86 image runs on an ARM64 Fargate task?
CannotPullContainerError due to a missing role
InsufficientFreeAddressesInSubnet
Task failed ELB health checks
exec format error
Where do you select ARM64 for a Fargate task?
In the cluster's instance type setting
In runtimePlatform.cpuArchitecture of the task definition
In the ALB listener rules
In the VPC route table

14. What operating systems does Fargate support?

Fargate supports Linux and Windows Server containers. Linux is the common case and has the most features.

For Windows you pick the Windows Server version in runtimePlatform.operatingSystemFamily, such as WINDOWS_SERVER_2022_CORE or WINDOWS_SERVER_2019_FULL. It must match the Windows Server version the container image was built on.

Windows tasks differ in a few ways:

  • They use platform version 1.0.0, numbered separately from Linux.
  • Tasks start at 1 vCPU, so there are no 0.25 or 0.5 vCPU sizes, and the maximum is 4 vCPU.
  • Only X86_64 is available, with no ARM64 and no Fargate Spot.
  • The hourly rate is higher because it includes the Windows license.
Take quiz
What must the Windows Server setting in runtimePlatform match?
The Windows Server version the container image was built on
The Linux kernel version of the host
The ECS cluster's creation date
The version of the AWS CLI in use
Which statement about Windows tasks on Fargate is true?
They can use 0.25 vCPU sizes
They can run on ARM64 Graviton
They start at 1 vCPU and run on X86_64 only
They share platform version numbers with Linux

15. How do you view container logs for Fargate tasks?

Fargate has no host to SSH into, so logs have to be shipped. The simplest option is the awslogs log driver, which sends container stdout and stderr to CloudWatch Logs.

"logConfiguration": {
  "logDriver": "awslogs",
  "options": {
    "awslogs-group": "/ecs/orders-api",
    "awslogs-region": "us-east-1",
    "awslogs-stream-prefix": "api"
  }
}

The task execution role needs logs:CreateLogStream and logs:PutLogEvents. Add logs:CreateLogGroup if you set awslogs-create-group to true.

Streams are named prefix/container-name/task-id. You can follow them live with aws logs tail /ecs/orders-api --follow. For routing to other destinations such as S3, Firehose, or third-party tools, use FireLens.

Take quiz
Which IAM role needs logs:PutLogEvents for the awslogs driver?
The task role
The task execution role
The ECS service-linked role only
The user running the AWS CLI
What does the awslogs driver capture from the container?
Only files under /var/log
Only CPU and memory metrics
Network packet captures
Standard output and standard error

16. What is ECS Exec and how do you use it?

ECS Exec opens an interactive shell, or runs a single command, inside a running container without SSH or a bastion. It rides on AWS Systems Manager Session Manager.

To use it on Fargate:

  1. Set enableExecuteCommand on the service or in run-task. Only tasks started afterwards get it.
  2. Add the ssmmessages:CreateControlChannel, CreateDataChannel, OpenControlChannel and OpenDataChannel permissions to the task role.
  3. Make sure the task can reach the SSM endpoints, through NAT or an ssmmessages VPC endpoint.
  4. Install the Session Manager plugin for the AWS CLI and run the command below.
aws ecs execute-command --cluster prod --task <task-id> \
  --container api --interactive --command "/bin/sh"

It is valuable for debugging, but in production you should log the sessions to CloudWatch or S3 through the cluster's execute-command configuration.

Take quiz
Which role needs the ssmmessages permissions for ECS Exec?
The task execution role
The load balancer's role
The task role
The AWS Config service role
What must you do to existing tasks after enabling ECS Exec on a service?
Start new tasks, since only new tasks get the feature
Nothing, running tasks pick it up immediately
Reboot the underlying EC2 instance
Recreate the VPC

17. What is a capacity provider strategy?

A capacity provider strategy tells ECS where to place tasks and in what proportion. On Fargate the two providers are FARGATE (On-Demand) and FARGATE_SPOT.

Each entry has two knobs:

  • base: the minimum number of tasks to run on that provider. Only one provider in a strategy can have a base.
  • weight: the relative share of the remaining tasks.
capacityProviderStrategy=[
  {"capacityProvider":"FARGATE",      "base":2, "weight":1},
  {"capacityProvider":"FARGATE_SPOT", "weight":3}
]

With ten tasks, two land on FARGATE as the base. The other eight split 1:3, so about two more go On-Demand and six go to Spot. If Spot capacity dries up, the guaranteed base keeps the service available.

Take quiz
In a capacity provider strategy, what does base mean?
The maximum price per task
The share of tasks as a percentage
The number of Availability Zones used
The minimum number of tasks placed on that provider
With base 2 on FARGATE and weights 1:3 for FARGATE and FARGATE_SPOT, where do the first two tasks go?
To FARGATE_SPOT, since it is cheaper
To FARGATE, because of the base
One to each provider
They wait until Spot capacity is confirmed

18. How do you pass secrets to Fargate containers?

Reference the secret in the container's secrets block instead of putting the value in environment. ECS fetches it from AWS Secrets Manager or SSM Parameter Store when the task starts and injects it as an environment variable.

"secrets": [
  {
    "name": "DB_PASSWORD",
    "valueFrom": "arn:aws:secretsmanager:us-east-1:111122223333:secret:prod/db-AbCdEf:password::"
  }
]

The task execution role must be allowed secretsmanager:GetSecretValue (or ssm:GetParameters) and kms:Decrypt if a customer-managed key is used.

Values are resolved only at launch. Rotating a secret does not update running tasks, so you must start new ones, for example with update-service --force-new-deployment. A task that cannot fetch the secret fails with a ResourceInitializationError.

Take quiz
When does ECS resolve a value from the secrets block?
When the task starts
Every time the application reads the variable
Once a day on a schedule
Only when the secret is rotated
Which role needs secretsmanager:GetSecretValue to inject a secret?
The task role
The load balancer role
The task execution role
The Secrets Manager service role

19. How do you define a Fargate service using Terraform?

You need a task definition with Fargate settings and an aws_ecs_service that references it with launch_type = "FARGATE" and a network configuration.

resource "aws_ecs_task_definition" "api" {
  family                   = "orders-api"
  requires_compatibilities = ["FARGATE"]
  network_mode             = "awsvpc"
  cpu                      = 512
  memory                   = 1024
  execution_role_arn       = aws_iam_role.exec.arn
  task_role_arn            = aws_iam_role.app.arn
  container_definitions    = file("api.json")
}

resource "aws_ecs_service" "api" {
  name            = "orders-api"
  cluster         = aws_ecs_cluster.main.id
  task_definition = aws_ecs_task_definition.api.arn
  desired_count   = 2
  launch_type     = "FARGATE"

  network_configuration {
    subnets          = var.private_subnet_ids
    security_groups  = [aws_security_group.api.id]
    assign_public_ip = false
  }
}

Three details trip people up. network_mode must be awsvpc, and cpu and memory must form a valid Fargate pair. The execution role pulls the image and writes logs, while the task role carries the application's own permissions. The service needs subnets and security groups because each task gets its own ENI.

Add a load_balancer block if the service sits behind an ALB, and consider lifecycle { ignore_changes = [desired_count] } when auto scaling manages the task count.

Take quiz
Which Terraform argument marks a task definition as Fargate-compatible?
instance_type = "fargate"
requires_compatibilities = ["FARGATE"]
capacity = "serverless"
compute_mode = "awsvpc"
Why add ignore_changes on desired_count when auto scaling is enabled?
Because desired_count is not supported on Fargate
To stop the service from using the load balancer
To force every apply to restart tasks
So Terraform doesn't reset the count that auto scaling changed

20. What happens when an essential container in a Fargate task exits?

If a container marked essential: true stops, for any reason, ECS stops every other container in the task and the task moves to STOPPED. A task definition needs at least one essential container.

In a service, the scheduler then launches a replacement task. For a standalone task, it simply ends.

Containers marked essential: false can exit without affecting the rest. That is the right setting for init steps and helpers whose failure should not take the app down. Use dependsOn to control start order:

"containerDefinitions": [
  {"name": "migrate", "essential": false, ...},
  {"name": "api", "essential": true,
   "dependsOn": [{"containerName": "migrate", "condition": "SUCCESS"}]}
]

Here api starts only after migrate exits with code 0. Make log-router and monitoring sidecars essential only if you really want the task to die without them.

Take quiz
What does ECS do when an essential container stops?
Restarts only that container forever
Keeps the task running without it
Stops all other containers in the task
Moves the container to another task
What does dependsOn with condition SUCCESS wait for?
The dependency container to exit with code 0
The dependency's first log line
The load balancer health check
The next platform revision

21. What is the difference between Fargate and EC2 launch types?

Both run the same ECS tasks. The difference is who manages the servers. With Fargate, AWS does. With the EC2 launch type, you run and patch a fleet of instances that ECS places tasks on.

Aspect Fargate EC2 launch type
Server management None You manage AMIs, patching, scaling
Billing Per task vCPU and memory, per second Per instance, whether full or not
Task sizes Fixed vCPU/memory pairs Anything the instance can hold
GPUs Not supported Supported
Privileged containers, host access Not allowed Allowed
Daemon services Not supported Supported (DAEMON)
Network mode awsvpc only awsvpc, bridge, host
Isolation Own microVM per task Tasks share the instance kernel

Fargate trades flexibility for less operational work. EC2 gives you control and can be cheaper at sustained high utilization.

Take quiz
What is the main difference between Fargate and the EC2 launch type?
Fargate cannot use ECS services
EC2 launch type cannot use load balancers
Fargate runs only Windows containers
Who manages the servers that run the tasks
Which capability is available on the EC2 launch type but not on Fargate?
Task IAM roles
GPU-backed tasks
awsvpc networking
Auto scaling of services

22. When should you choose Fargate over EC2?

Choose Fargate when you would rather spend engineering time on the application than on the cluster. It is the better default for most new container workloads.

It fits especially well when:

  • Traffic is spiky or unpredictable. You pay per task and there is no spare instance capacity to buy up front.
  • The team is small. No AMI pipelines, patching windows, or capacity planning.
  • Workloads are batch or event-driven and run briefly, so an always-on fleet would sit idle.
  • Isolation matters. Each task has its own kernel boundary, which helps with multi-tenant or compliance-sensitive services.
  • Tasks are small. Packing many 0.25 to 2 vCPU tasks onto instances rarely saves enough to offset the operations effort.

Start on Fargate and move a service to EC2 only when you hit a concrete limit, such as a GPU requirement or a measured cost gap at steady load.

Take quiz
Which situation favors Fargate over EC2?
A small team running spiky workloads without a platform engineer
A job that needs GPUs
A daemon agent required on every host
A container that needs privileged host access
Why does Fargate suit short batch jobs well?
Batch jobs are free on Fargate
Tasks start in under a millisecond
You pay only while each task runs, with no idle fleet
It reserves instances for a year automatically

23. When would you choose EC2 over Fargate?

Choose the EC2 launch type when a requirement is something Fargate cannot offer, or when sustained utilization makes managing instances worth it.

Situation Why EC2 wins
GPU or specialized hardware Fargate has no GPU support
Privileged containers, host mounts, custom kernel modules Fargate blocks host-level access
Daemon or agent per host (security, monitoring) DAEMON scheduling is EC2 only
Very large or odd task sizes Fargate caps at 16 vCPU / 120 GB with fixed pairs
Steady, high-utilization fleet Reserved or Savings Plan instances, tightly packed, can cost less
Need for local NVMe or instance storage Not available on Fargate

A mixed setup is common: Fargate for web services and jobs, EC2 for the one GPU or host-level workload.

Take quiz
Which requirement forces you to use EC2 instead of Fargate?
Using an Application Load Balancer
Running GPU inference containers
Using task IAM roles
Using CloudWatch Logs
Why can a well-packed EC2 fleet cost less than Fargate?
Fargate charges for every request
EC2 has no networking charges
Instances are billed only when tasks run
Tightly packed tasks at high steady utilization waste little capacity

24. Explain the lifecycle of a Fargate task?

An ECS task on Fargate moves through a fixed set of states. Knowing them tells you where to look when something stalls.

stateDiagram-v2
  [*] --> PROVISIONING
  PROVISIONING --> PENDING
  PENDING --> ACTIVATING
  ACTIVATING --> RUNNING
  RUNNING --> DEACTIVATING
  DEACTIVATING --> STOPPING
  STOPPING --> DEPROVISIONING
  DEPROVISIONING --> STOPPED
  STOPPED --> [*]
  1. PROVISIONING: Fargate reserves capacity and creates the task's ENI in your subnet.
  2. PENDING: the agent is starting; images are pulled and secrets fetched.
  3. ACTIVATING: containers are started and the task is being registered with load balancers or service discovery.
  4. RUNNING: the task is serving. Health checks and metrics apply.
  5. DEACTIVATING: the task is being removed from load balancers and deregistered.
  6. STOPPING: containers get SIGTERM, then SIGKILL after stopTimeout.
  7. DEPROVISIONING: the ENI is detached and resources are released.
  8. STOPPED: the task is finished. Its stoppedReason stays visible for a short time.

A task stuck in PROVISIONING or PENDING usually points to subnet IPs, image pull, or secret problems, while repeated STOPPED tasks point to application or health check failures.

Take quiz
In which state does a Fargate task get its ENI created?
RUNNING
DEACTIVATING
PROVISIONING
STOPPED
What happens during the STOPPING state?
Containers receive SIGTERM and then SIGKILL after stopTimeout
Images are pulled again for the next task
The ENI is created
The task is registered with the load balancer

25. Explain the execution flow of launching a Fargate task?

From the moment you call RunTask or a service needs a new task, ECS and Fargate coordinate several steps before your code runs.

sequenceDiagram
  participant U as You or Service Scheduler
  participant E as ECS Control Plane
  participant F as Fargate
  participant V as Your VPC
  participant R as ECR / Secrets
  U->>E: RunTask (task definition, subnets, SGs)
  E->>F: Request capacity (cpu, memory, AZ)
  F->>V: Create ENI in chosen subnet
  F->>R: Pull image, fetch secrets (execution role)
  F->>F: Start containers in microVM
  F->>E: Report RUNNING
  E->>V: Register task IP with load balancer
  1. ECS validates the task definition and picks a subnet and Availability Zone.
  2. Fargate allocates an isolated microVM with the requested vCPU and memory.
  3. An ENI is attached to the task in your subnet using your security groups.
  4. The agent assumes the task execution role, pulls the image, and resolves secrets.
  5. Containers start in dependsOn order and any container health checks begin.
  6. The task reports RUNNING and, for a service, the target group registers its IP.

Any failure in steps 3 or 4, such as no IPs, no route to ECR, or missing permissions, shows up as a stopped task with a stoppedReason.

Take quiz
Which role does Fargate use to pull the image and fetch secrets during launch?
The task role
The service-linked role of the load balancer
The IAM user that called RunTask
The task execution role
Where does the task's ENI get created?
In an AWS-owned VPC outside your account
In a subnet you specify in the network configuration
On an EC2 instance you manage
In the ECR registry

26. How does Fargate isolate tasks from each other?

Every Fargate task runs in its own isolated compute environment built on Firecracker microVM technology. Tasks do not share a kernel, CPU, memory, or network interface with tasks from other customers, or with your other tasks.

That is a real difference from the EC2 launch type, where all tasks on an instance share the host kernel and rely on container namespaces and cgroups for separation.

Within a single task the boundary is softer. Containers in the same task share the task's kernel, ENI, and ephemeral storage, which is why localhost networking works between them. Put workloads with different trust levels in separate tasks, not as sidecars of each other.

Isolation also explains some Fargate limits. With no shared host there is nothing to mount or inspect, so privileged mode, host networking, and host path mounts are not offered.

Take quiz
What technology does Fargate use to isolate each task?
Firecracker microVMs
Shared Docker bridge networks
One EC2 instance per cluster
A single shared kernel for all customers
What do containers in the same Fargate task share?
Nothing at all
Only the Docker socket of the host
The kernel, the ENI, and ephemeral storage
The kernel of other customers' tasks

27. How do you troubleshoot CannotPullContainerError on Fargate?

The task failed to download its image before any container started. Start with aws ecs describe-tasks and read the stoppedReason. It usually names the cause.

Check these in order:

  1. Network path to the registry. A task in a private subnet needs a NAT gateway, or VPC endpoints for ecr.api, ecr.dkr, an S3 gateway endpoint (layers live in S3), and logs if you use awslogs. In a public subnet, assignPublicIp must be ENABLED.
  2. Security group egress. Outbound HTTPS (443) must be allowed.
  3. Execution role. It needs ecr:GetAuthorizationToken, ecr:BatchGetImage, and ecr:GetDownloadUrlForLayer.
  4. Image URI and tag. A typo, a deleted tag, or the wrong region or account causes a not-found error.
  5. External registries. Docker Hub rate limits or missing credentials in repositoryCredentials will fail pulls.

If it only fails sometimes, suspect rate limits or a flaky NAT path rather than permissions.

Take quiz
A Fargate task in a private subnet cannot pull from ECR. Which fix is needed?
Assign an Elastic IP to the container
A NAT gateway, or ECR and S3 VPC endpoints
Raise the task CPU to 4 vCPU
Enable ECS Exec on the service
Why does an ECR setup with VPC endpoints also need an S3 gateway endpoint?
ECR stores task definitions in S3
S3 provides the load balancer for ECR
The S3 endpoint holds the IAM role
Image layers are served from S3

28. How do you troubleshoot a Fargate task stuck in PENDING?

A task sitting in PROVISIONING or PENDING has not yet started its containers. Work out which phase it is stuck in: networking, image pull, or secrets.

Symptom Likely cause Check
Task fails right away, no ENI Subnet has no free IPs Free IPs in the subnet; add subnets or a larger CIDR
Stays PENDING, then CannotPullContainerError No route to ECR or registry NAT / VPC endpoints, security group egress
ResourceInitializationError Cannot fetch secrets or reach SSM / Secrets Manager Execution role permissions, endpoints, KMS access
Stays PENDING for minutes Very large image or slow pull Image size, region, SOCI, endpoint throughput
Service shows events but no tasks Capacity or service-quota limit Service events, Fargate vCPU quotas

Run aws ecs describe-services and read the events list, then describe-tasks for the stopped task's stoppedReason. The service event log is often the fastest clue. If the reason says the ENI could not be created, look at the subnet and security group first.

Take quiz
What is a likely cause when a Fargate task fails before its ENI is created?
The container image uses ARM64
The log retention period is too short
The subnet has run out of free IP addresses
The task role is too permissive
Which error points to a failure fetching secrets or reaching Secrets Manager?
ResourceInitializationError
exec format error
OutOfMemoryError
TaskFailedToStart for ELB health check

29. How do you troubleshoot exit code 137 on Fargate?

Exit code 137 means the container received SIGKILL (128 + 9). On Fargate this is most often the out-of-memory killer, though it can also appear when a container ignores SIGTERM and ECS force-kills it after stopTimeout.

To tell which one it is:

  1. Run describe-tasks and read stoppedReason. OutOfMemoryError: Container killed due to memory usage confirms OOM.
  2. Check MemoryUtilization in CloudWatch Container Insights for the minutes before the kill.
  3. If the memory graph is flat and the stop was a deployment, the app probably does not handle SIGTERM.

Fixes for the OOM case:

  • Raise the task memory, within a valid CPU and memory pairing.
  • Set a hard container memory limit lower than the task size so a leaky sidecar cannot starve the app.
  • Tune the runtime. For Java, use -XX:MaxRAMPercentage=75 so the heap respects the container limit.
  • Profile for leaks and unbounded caches.
Take quiz
What does exit code 137 indicate for a container?
It exited cleanly after finishing its work
It failed to parse its config file
It was stopped by SIGTERM
It was killed with SIGKILL, often by the out-of-memory killer
Which JVM flag helps a Java container stay inside its memory limit?
-XX:+UseFargate
-XX:MaxRAMPercentage
-Xdebug
-XX:AutoScaleHeap

30. How does ECS service auto scaling work with Fargate?

ECS service auto scaling uses Application Auto Scaling to change the service's DesiredCount. Because Fargate has no instances, scaling the task count is the whole story. There is no cluster capacity to scale first.

Policy type How it works Typical use
Target tracking Holds a metric near a target, such as 60% CPU Default choice for web services
Step scaling CloudWatch alarm adds or removes tasks by thresholds Fine control over bursts
Scheduled Changes min/max or count at set times Predictable daily peaks

Built-in metrics are ECSServiceAverageCPUUtilization, ECSServiceAverageMemoryUtilization, and ALBRequestCountPerTarget. For queue consumers, publish a custom metric such as messages per task.

Scale-out is fast, but a new task still takes time to provision and pull its image. Keep a sensible minimum and use shorter scale-out than scale-in cooldowns so traffic spikes are handled without flapping.

Take quiz
What does ECS service auto scaling change on a Fargate service?
The service's DesiredCount
The number of EC2 instances in the cluster
The CPU size of each running task
The Availability Zones in the VPC
Which built-in metric scales tasks based on traffic from an ALB?
ECSTaskDiskUsage
NetworkInterfaceCount
ALBRequestCountPerTarget
FargateSpotPrice

31. How can you reduce AWS Fargate costs?

Cost comes from task size multiplied by running time, so most savings come from smaller tasks, fewer hours, or cheaper rates.

  1. Right-size tasks. Compare actual CPU and memory in Container Insights with the configured size. Compute Optimizer also gives ECS on Fargate recommendations.
  2. Switch to ARM64 (Graviton). About 20% lower rates for compatible images.
  3. Use Fargate Spot for fault-tolerant tasks, mixed with a small On-Demand base.
  4. Buy Compute Savings Plans to cover the steady baseline.
  5. Scale down off-hours with scheduled scaling, especially in dev and test environments.
  6. Trim startup time. Billing starts at image pull, so smaller images save money on every launch.
  7. Cut the extras. Use VPC endpoints instead of NAT for ECR and S3 traffic, and set log retention.

Do the right-sizing first. It is free and usually the largest saving, since many services run at 10-20% of the memory they reserve.

Take quiz
Which change typically lowers the Fargate rate for compatible workloads by about 20%?
Adding more sidecar containers
Moving from x86 to ARM64 Graviton
Using larger ephemeral storage
Switching to the bridge network mode
Why does a smaller container image save money on Fargate?
Smaller images use a cheaper CPU type
Image size is billed separately per request
Small images skip the execution role
Billing starts when the image pull begins

32. How do Savings Plans apply to Fargate?

Compute Savings Plans apply to Fargate usage automatically. You commit to a dollar amount of compute per hour for one or three years, and in return the covered usage is billed at a lower rate.

The same commitment covers EC2, Lambda, and Fargate. It is not tied to a region, instance family, or operating system, so you can move workloads between services and still use the discount.

Things to know:

  • Usage above the commitment is billed at normal On-Demand rates.
  • Fargate Spot is priced on its own and is not discounted further by a Savings Plan.
  • The commitment is in dollars, not vCPUs, so size it to your steady baseline, not your peak.
  • Unused commitment is lost for that hour. Do not cover bursty usage.

A common approach is to use the plan for baseline On-Demand tasks, run extra capacity on Spot, and review coverage and utilization in Cost Explorer.

Take quiz
How is a Compute Savings Plan commitment expressed?
As a fixed number of vCPUs
As a count of running tasks
As a dollar amount of compute spend per hour
As a specific instance family
What happens to usage that exceeds your Savings Plan commitment?
It is billed at normal On-Demand rates
It is stopped automatically
It is billed at the Spot rate
It is free for the remainder of the hour

33. How do you handle Fargate Spot interruptions?

Treat interruption as a normal event, not a failure. When AWS reclaims capacity, ECS sends SIGTERM to the task's containers about two minutes before termination, and SIGKILL follows after stopTimeout (maximum 120 seconds on Fargate).

Make the application cooperate:

  1. Catch SIGTERM. Stop accepting new work, finish or checkpoint what is in flight, and exit with code 0.
  2. Raise stopTimeout in the container definition so the shutdown has time, up to the 120 second limit.
  3. Make work idempotent. A queue message that was not deleted will return after its visibility timeout and be processed again.
  4. Keep an On-Demand base. Use a capacity provider strategy with base on FARGATE so the service never drops to zero.
  5. Spread the risk. Use several subnets in different Availability Zones.

You can also watch the ECS Task State Change events in EventBridge. A stop caused by Spot reclaim is reported in the task's stoppedReason, which helps you track how often it happens.

Take quiz
What should a Fargate Spot task do when it receives SIGTERM?
Ignore it and keep processing
Restart itself immediately
Delete its own task definition
Stop taking new work, checkpoint in-flight work, and exit cleanly
How do you keep a Spot-based service from dropping to zero tasks?
Put every task on FARGATE_SPOT only
Set a base on the FARGATE capacity provider
Disable the load balancer health check
Increase the ephemeral storage

34. How do you attach persistent storage to Fargate tasks?

Ephemeral storage disappears with the task, so durable data needs another option. Fargate supports two: Amazon EFS and Amazon EBS.

Ephemeral Amazon EFS Amazon EBS
Lifetime Task only Independent of tasks Tied to the task by default
Sharing Containers in one task Many tasks at once One task at a time
Type Local disk NFS file system Block volume
Best for Scratch space, caches Shared files, CMS uploads, ML model files Single-writer databases, low-latency block IO
Size 20 - 200 GiB Elastic Configured per volume

For EFS, create mount targets in the task's subnets, allow NFS (port 2049) between the task and mount target security groups, and declare an efsVolumeConfiguration in the task definition. Access points and IAM authorization give per-service control.

Data stores like S3 or DynamoDB are often better than any disk. Prefer them when the application can use an API instead of a file system.

Take quiz
Which Fargate storage option can be mounted by many tasks at the same time?
Amazon EFS
Ephemeral storage
Amazon EBS
The instance store
Which network port must be open between a Fargate task and its EFS mount target?
TCP 22 for SSH
TCP 3306 for MySQL
TCP 2049 for NFS
UDP 53 for DNS only

35. How do you attach an EBS volume to a Fargate task?

ECS can create and attach an encrypted EBS volume to each Fargate task when it launches. You mark the volume as configuredAtLaunch in the task definition, then supply the details when you run the task or create the service.

  1. In the task definition, declare a volume with "configuredAtLaunch": true and mount it in the container with a mountPoints entry.
  2. In run-task or create-service, pass volumeConfigurations with a managedEBSVolume: size, volume type, encryption, optional KMS key or snapshot.
  3. Provide an ECS infrastructure IAM role that lets ECS manage the volume on your behalf.
"volumeConfigurations": [{
  "name": "data",
  "managedEBSVolume": {
    "roleArn": "arn:aws:iam::111122223333:role/ecsInfrastructureRole",
    "volumeType": "gp3",
    "sizeInGiB": 50,
    "encrypted": true,
    "filesystemType": "ext4"
  }
}]

Each volume attaches to a single task and cannot be shared. By default the volume is deleted when the task ends, so use a snapshot or terminationPolicy settings if the data must survive.

Take quiz
Which task definition setting lets a volume be configured when the task launches?
mountHost set to true
configuredAtLaunch set to true
persistAfterStop set to true
efsAuthorization set to enabled
Can an EBS volume attached to a Fargate task be shared by several tasks?
Yes, up to ten tasks
Yes, if they share a security group
Yes, but only in the same subnet
No, it attaches to a single task

36. How do you load balance traffic to Fargate tasks?

Put an Application Load Balancer (HTTP/HTTPS) or Network Load Balancer (TCP/UDP/TLS) in front of the service and register tasks through a target group with target type ip. Instance targets do not work because tasks use awsvpc and have no host instance.

flowchart LR
  C[Client] --> L["ALB in public subnets"]
  L --> T1["Task 1 IP, private subnet AZ-a"]
  L --> T2["Task 2 IP, private subnet AZ-b"]
  L --> T3["Task 3 IP, private subnet AZ-c"]

The service definition ties the pieces together:

"loadBalancers": [{
  "targetGroupArn": "arn:aws:elasticloadbalancing:...:targetgroup/api/abc123",
  "containerName": "api",
  "containerPort": 8080
}],
"healthCheckGracePeriodSeconds": 60

ECS registers each new task IP when it becomes ready and deregisters it on shutdown. The task security group must allow inbound traffic from the load balancer's security group on the container port. Set a grace period so slow-starting apps are not killed by early failed health checks.

Also match the target group's deregistration delay to the container's stopTimeout. That way in-flight requests finish before the task receives SIGKILL, and rolling deployments don't drop connections.

Take quiz
Which target type must a target group use for Fargate tasks?
instance
lambda
ip
host
What should the task security group allow inbound?
Traffic from the load balancer's security group on the container port
All traffic from 0.0.0.0/0 on every port
SSH from the internet
Only traffic from other tasks' ENIs

37. What is the difference between Service Connect and Cloud Map discovery?

Both let services find each other by name. Cloud Map service discovery is plain DNS. ECS Service Connect adds a proxy sidecar that handles the connection, and it uses a Cloud Map namespace underneath.

Aspect Cloud Map service discovery ECS Service Connect
Mechanism DNS records for task IPs ECS-managed proxy sidecar next to your container
Load balancing Client-side, from DNS answers Round robin by the proxy
Failure handling Clients may cache stale IPs until TTL expires Proxy retries and removes unhealthy tasks quickly
Metrics None built in Per-connection metrics in CloudWatch
Client changes Resolve names yourself Use short names like http://orders:8080

Service Connect is the better default for service-to-service calls inside a cluster, because it avoids DNS caching surprises and gives you request metrics. Plain Cloud Map still makes sense when you need DNS records for non-ECS clients or want no proxy.

Take quiz
What does Service Connect add on top of DNS-based discovery?
A public IP for every task
A replacement for the VPC
An EC2 instance per service
A proxy sidecar that retries and tracks connections
What is a drawback of pure DNS discovery with Cloud Map?
It cannot resolve private IPs
Clients can keep using stale IPs until the TTL expires
It needs a GPU
It works only with Windows tasks

38. How do rolling deployments work for Fargate services?

In a rolling update (the default ECS deployment), the scheduler starts tasks from the new task definition revision and stops old ones step by step until all tasks are replaced.

Two settings control the pace:

  • minimumHealthyPercent (default 100): the floor of healthy tasks, as a percentage of desiredCount, during the rollout.
  • maximumPercent (default 200): the ceiling of total tasks, including old and new ones.

With four tasks and the defaults, ECS can launch up to four new tasks beside the four old ones, wait for them to pass health checks, then drain the old ones. Capacity never drops below four. On Fargate there is no instance limit to block this, but you pay for the extra tasks for a few minutes and need spare subnet IPs.

Setting minimumHealthyPercent to 50 and maximumPercent to 100 replaces tasks in place with no extra cost, but at half capacity during the rollout. Pair either with the deployment circuit breaker to roll back failed releases. For traffic shifting and instant rollback, ECS also offers blue/green deployments.

Take quiz
With the defaults of 100 and 200 and four desired tasks, how many tasks can run during a rollout?
Up to eight, with at least four healthy
Exactly four, with none healthy
Up to sixteen with no minimum
Only two at a time
What is the trade-off of minimumHealthyPercent 50 and maximumPercent 100?
Rollouts become instant with no risk
Tasks are billed double
No extra tasks are launched, but capacity halves during the rollout
Old tasks never get stopped

39. How does the ECS deployment circuit breaker work?

The deployment circuit breaker watches a rolling deployment and stops it if new tasks keep failing, so a bad release doesn't loop forever. With rollback enabled it also returns the service to the last good revision.

flowchart TD
  A["New deployment starts"] --> B["Launch tasks from new revision"]
  B --> C{Tasks reach steady state?}
  C -- Yes --> D["Deployment completes"]
  C -- No, launches or health checks keep failing --> E{Failure threshold reached?}
  E -- No --> B
  E -- Yes --> F["Deployment marked FAILED"]
  F --> G["Roll back to last completed deployment"]

A failure is counted when a task cannot start or fails health checks. The threshold scales with the service's desired count, with a minimum of a few tasks. When it is crossed, the deployment is marked failed.

deploymentConfiguration={
  "deploymentCircuitBreaker": {"enable": true, "rollback": true}
}

Typical triggers on Fargate are a wrong image tag, a missing secret, an app that crashes at startup, or health check paths that return errors. Use it with CloudWatch alarms for application-level errors it cannot see.

Take quiz
What does the circuit breaker do when failures reach its threshold with rollback enabled?
Deletes the cluster
Marks the deployment failed and returns to the last good revision
Doubles the desired count
Switches the service to the EC2 launch type
Which problem is the circuit breaker designed to catch?
Slow database queries after startup
High NAT gateway bills
Misspelled log group names after a successful deploy
New tasks that repeatedly fail to start or pass health checks

40. How do you configure health checks for Fargate tasks?

Fargate tasks can be checked at two levels, and they do different jobs. A container health check runs a command inside the container. A load balancer health check makes an HTTP or TCP request to the task's IP.

"healthCheck": {
  "command": ["CMD-SHELL", "curl -f http://localhost:8080/health || exit 1"],
  "interval": 30,
  "timeout": 5,
  "retries": 3,
  "startPeriod": 60
}

  • startPeriod gives the app time to boot before failures count.
  • If an essential container turns UNHEALTHY, ECS stops the task and the service replaces it.
  • The target group health check decides whether the ALB sends traffic to the task. A failing target is deregistered and, in a service, replaced.
  • healthCheckGracePeriodSeconds on the service stops ECS from replacing tasks that are still warming up.

The container command must exist in the image. A distroless or scratch image without curl will report unhealthy even if the app is fine.

Take quiz
What does startPeriod do in a container health check?
Sets how long logs are kept
Defines the task's CPU units
Lets the app boot before failed checks are counted
Schedules the next deployment
Why might a distroless image fail a curl-based container health check?
The curl binary does not exist in the image
Distroless images cannot run on Fargate
Health checks need a public IP
The execution role lacks curl permissions

41. How do you secure Fargate workloads?

AWS secures the underlying host and platform. You secure what runs on it: identity, network, image, and secrets. This is the shared responsibility model.

Area What to do
Identity One task role per service with least privilege. Keep the execution role narrow, limited to ECR, logs, and the specific secrets.
Network Run tasks in private subnets. Give each service its own security group and allow inbound only from the load balancer's security group.
Image Scan images in ECR, pin tags or digests, use minimal base images, and run as a non-root user.
Container Set readonlyRootFilesystem to true and write only to explicit volumes.
Secrets Use the secrets block with Secrets Manager or SSM, never plain environment values.
Access Leave ECS Exec off by default. When on, log sessions to CloudWatch or S3.
Data Encrypt EFS and EBS, and optionally use a customer-managed KMS key for ephemeral storage.

Add VPC endpoints so traffic to AWS APIs stays off the internet, and use CloudTrail to audit who changed task definitions and services.

Take quiz
Which setting prevents a container from writing to its own root file system?
privileged set to false
essential set to true
assignPublicIp set to DISABLED
readonlyRootFilesystem set to true
Where should Fargate tasks normally run for better network security?
In public subnets with open security groups
In private subnets behind a load balancer
Directly on the internet gateway
Inside the default VPC only

42. Why can't Fargate run privileged containers?

There is no host for a container to be privileged on. In Fargate you do not own the instance, and AWS keeps the platform locked down so tasks cannot affect it, or each other.

This shows up as concrete limits:

  • privileged: true is rejected.
  • No host network mode and no host path or /var/run/docker.sock mounts.
  • Linux capabilities cannot be freely added. Only SYS_PTRACE is allowed to be added.
  • No custom kernel modules or sysctl changes beyond what the platform exposes.

Tools that need these rights, such as Docker-in-Docker, host monitoring agents, and eBPF tools, do not work directly. Workarounds are to build images in CodeBuild instead, use rootless build tools like Kaniko, or run the workload on the EC2 launch type.

Because the microVM boundary already isolates the task, the lost flexibility is also what gives Fargate its security guarantee.

Take quiz
Which Linux capability can be added to a Fargate container?
SYS_PTRACE
SYS_ADMIN
NET_ADMIN
SYS_MODULE
What is a common alternative to running Docker-in-Docker on Fargate?
Mount the host Docker socket
Enable privileged mode in the task definition
Use a rootless build tool or build in a service like CodeBuild
Switch the network mode to host

43. How does Fargate encrypt ephemeral storage?

On Linux platform version 1.4.0 and later, a task's ephemeral storage is encrypted at rest with AES-256 by default, using an encryption key managed by AWS Fargate. You do nothing to turn it on.

If your compliance rules require your own key, configure a customer-managed KMS key at the cluster level through managedStorageConfiguration, using the fargateEphemeralStorageKmsKeyId setting.

aws ecs create-cluster --cluster-name secure-prod \
  --configuration 'managedStorageConfiguration={fargateEphemeralStorageKmsKeyId=arn:aws:kms:us-east-1:111122223333:key/1234abcd-12ab-34cd-56ef-1234567890ab}'

The key policy must allow the Fargate service to use the key. Only tasks launched after the key is configured use it. Existing tasks keep the old encryption.

This covers storage inside the task. EFS and EBS volumes are encrypted separately with their own key settings, and data in transit still needs TLS.

Take quiz
How is Fargate ephemeral storage protected by default on platform version 1.4.0?
It is not encrypted unless you enable it
Encrypted at rest with AES-256 using an AWS-managed key
Encrypted only when a public IP is assigned
Encrypted with a key you must rotate manually
Where do you set a customer-managed KMS key for Fargate ephemeral storage?
In each container's logConfiguration
On the load balancer listener
In the task's security group
In the cluster's managedStorageConfiguration

44. How can you speed up Fargate container startup?

Startup time is the sum of ENI creation, image pull, container start, and the health check delay before traffic arrives. Image pull and application boot are the parts you control.

  1. Shrink the image. Use multi-stage builds, slim or distroless bases, and remove build tools and caches.
  2. Use SOCI lazy loading. A Seekable OCI index lets Fargate start a container before the whole image is downloaded, which helps large images.
  3. Keep ECR in the same region and use VPC endpoints so pulls take a short, fast path.
  4. Cut application boot time. Lazy-load, trim startup work, and for Java consider class data sharing or AOT options.
  5. Tune health checks. A startPeriod and grace period that match real boot time avoid false failures that restart tasks.
  6. Scale ahead of demand. Scheduled scaling or a higher minimum hides the delay from users.

Do not rely on layer caching between tasks, since each new task may start on fresh infrastructure. Measure with the task's createdAt, pullStartedAt, pullStoppedAt, and startedAt timestamps to see where time goes.

Take quiz
What does a SOCI index let Fargate do?
Skip the ENI creation step
Run containers without an execution role
Start a container before the entire image has downloaded
Cache images permanently on your laptop
Which task timestamps help measure image pull time?
pullStartedAt and pullStoppedAt
connectivityAt and stopCode
healthStatus and lastStatus
platformVersion and capacityProviderName

45. How do you monitor Fargate tasks?

Monitoring uses four layers: metrics, logs, events, and traces. Since you cannot install agents on the host, everything is delivered through ECS features or sidecars.

Layer Tool What you get
Service metrics CloudWatch AWS/ECS namespace CPUUtilization, MemoryUtilization per service
Task and container metrics Container Insights Per-task CPU, memory, network, storage, and running task counts
Logs awslogs driver or FireLens stdout/stderr in CloudWatch Logs or other targets
Events EventBridge ECS Task State Change and Service Action Alerts on stopped tasks, failed deployments, Spot reclaim
Traces and custom metrics AWS Distro for OpenTelemetry sidecar, X-Ray Request traces and app-level metrics

Inside a task, the task metadata endpoint v4 (from the ECS_CONTAINER_METADATA_URI_V4 variable) returns the task ARN, container stats, and network details for the application to use.

Useful alarms: running task count below desired, memory above 80%, target group unhealthy hosts, and deployment failure events.

Take quiz
Which feature provides per-task CPU and memory metrics for Fargate?
VPC Flow Logs
AWS Trusted Advisor
Amazon Inspector findings
CloudWatch Container Insights
Where can a container find its task ARN and stats from inside the task?
The EC2 instance metadata service of the host
The task metadata endpoint v4
The load balancer access logs
The ECR image manifest

46. How do you use FireLens for logging on Fargate?

FireLens lets you route container logs to many destinations by running a Fluent Bit (or Fluentd) sidecar called the log router. Your app container sends logs with the awsfirelens driver, and the sidecar forwards them.

You need two parts in the task definition:

{
  "name": "log_router",
  "image": "public.ecr.aws/aws-observability/aws-for-fluent-bit:stable",
  "essential": true,
  "firelensConfiguration": {"type": "fluentbit"}
},
{
  "name": "api",
  "image": "...",
  "logConfiguration": {
    "logDriver": "awsfirelens",
    "options": {
      "Name": "firehose",
      "region": "us-east-1",
      "delivery_stream": "app-logs"
    }
  }
}

The destination is chosen by the Name option. Fluent Bit supports CloudWatch, Amazon Data Firehose, S3, OpenSearch, and partner tools such as Datadog and Splunk.

You can also supply a custom config file for parsing, filtering, or multiline handling. Give the log router enough CPU and memory, because it shares the task's resources, and decide deliberately whether it is essential.

Take quiz
Which log driver does an app container use to send logs through FireLens?
awsfirelens
awslogs
json-file
syslog-host
What does the FireLens log router container run?
A second copy of the application
The ECS agent
Fluent Bit or Fluentd
A Prometheus server

47. How do Fargate profiles work in EKS?

A Fargate profile tells EKS which pods should run on Fargate. When a pod is created and matches a profile's selectors (a namespace and optional labels), EKS schedules it onto Fargate instead of an EC2 worker node.

flowchart TD
  A["Pod created"] --> B{Matches a Fargate profile selector?}
  B -- Yes --> C["Mutating webhook sets scheduler to fargate-scheduler"]
  C --> D["Fargate starts a microVM sized to the pod"]
  D --> E["Pod runs on its own virtual node"]
  B -- No --> F["Default scheduler places it on EC2 nodes"]

What happens behind the scenes:

  • A mutating admission webhook changes the pod's scheduler to fargate-scheduler.
  • EKS launches a dedicated microVM for that pod, which appears as its own node.
  • The pod is sized from the sum of its container resource requests, plus a small overhead for Kubernetes components, and rounded up to the next Fargate size.
  • The profile supplies the pod execution role and the private subnets to use.

A pod that does not match any profile will not run on Fargate, even if a Fargate node exists. Always set resource requests, since they determine what you are billed.

Take quiz
What makes an EKS pod run on Fargate?
Adding a GPU limit to the pod
Matching a Fargate profile's namespace and label selectors
Running kubectl on a bastion host
Using a node group with spot instances
What determines the size of the Fargate microVM a pod gets?
The size of the cluster's node group
The pod name length
The EKS control plane version
The pod's resource requests plus a small overhead, rounded up

48. What are the limitations of Fargate on EKS?

Fargate on EKS suits stateless pods, but it removes features that rely on controlling the node. Know these before you move workloads.

Limitation Detail
No DaemonSets There is no shared node to put a per-node pod on. Use sidecars instead.
No privileged pods Also no hostNetwork, hostPort, or host path volumes.
No GPUs Not supported on Fargate.
Private subnets only Pods get ENIs in private subnets, so you need NAT or endpoints for outbound access.
Storage EFS is supported for persistence. EBS volumes are not.
Load balancers Use ALB or NLB with target type ip.
Per-pod cost floor One microVM per pod, so many tiny pods cost more than packing them on a node.
System pods CoreDNS needs a Fargate profile for kube-system or must run on EC2 nodes.

DaemonSet-based tools, such as log collectors and security agents, need a replacement. Use sidecars for logging, or the platform's built-in log router for Fargate pods.

Take quiz
Which Kubernetes feature is not available on Fargate in EKS?
Deployments
Services
DaemonSets
ConfigMaps
Which persistent storage option works for EKS pods on Fargate?
Amazon EFS
Amazon EBS volumes
Instance store
hostPath volumes

49. How do you run scheduled jobs on Fargate?

A job on Fargate is just a task started on a timer. The usual setup is Amazon EventBridge Scheduler (or an EventBridge rule) with an ECS RunTask target.

  1. Create a task definition for the job with a command that exits when finished.
  2. Create a schedule, for example cron(0 2 * * ? *) for 02:00 UTC daily.
  3. Choose the ECS cluster, the task definition, a FARGATE launch type or capacity provider, and the network configuration.
  4. Give the scheduler an IAM role allowed to call ecs:RunTask and iam:PassRole for the task roles.

You pay only while the task runs. Add a FARGATE_SPOT strategy for jobs that can retry.

For heavier needs, there are two other options:

  • AWS Batch with a Fargate compute environment, which adds queues, retries, array jobs, and dependencies.
  • Step Functions with the ecs:runTask.sync integration, which waits for the task and lets you branch on success or failure.

Make the job idempotent, set an alarm on STOPPED tasks with a non-zero exit code, and keep a retry policy on the schedule.

Take quiz
Which service commonly triggers an ECS RunTask on a cron schedule?
AWS Config
Amazon Macie
AWS Shield
Amazon EventBridge Scheduler
Which option adds queues, retries, and job dependencies for Fargate batch work?
An ECS service with desired count of one
AWS Batch with a Fargate compute environment
A bigger NAT gateway
A cross-region ECR replication rule

50. How would you architect a highly available Fargate service?

High availability on Fargate means removing single points of failure in the task layer, the network layer, and the deployment process. Fargate already manages the hosts, so the design focuses on the rest.

flowchart TD
  U[Users] --> ALB["ALB across 3 public subnets"]
  ALB --> S["ECS service, desiredCount 6"]
  S --> A["Tasks in private subnet AZ-a"]
  S --> B["Tasks in private subnet AZ-b"]
  S --> C["Tasks in private subnet AZ-c"]
  A --> D[(Database, Multi-AZ)]
  B --> D
  C --> D
  AS["Application Auto Scaling"] --> S
  1. Spread across at least two Availability Zones. List private subnets from each AZ in the service. ECS balances tasks across them.
  2. Run a minimum of two tasks, and more if losing one AZ must not hurt capacity. Size for N-1 AZs.
  3. Use an ALB with health checks and a grace period, so only ready tasks get traffic.
  4. Mix capacity. Put a base on FARGATE and let extra tasks run on FARGATE_SPOT.
  5. Auto scale on CPU or request count, with a floor equal to your availability minimum.
  6. Deploy safely. Rolling updates with minimumHealthyPercent 100, the circuit breaker with rollback, or blue/green.
  7. Handle shutdown. Catch SIGTERM, set stopTimeout, and keep deregistration delay in line.
  8. Keep state outside the task in a Multi-AZ database, EFS, S3, or a managed cache.

For regional resilience, repeat the stack in a second region behind Route 53 health-checked routing, and replicate images with ECR cross-region replication.

Take quiz
How many Availability Zones should a highly available Fargate service span at minimum?
At least two
Exactly one
None, Fargate is global
Only the zone with the cheapest price
What keeps a service from dropping to zero when Spot capacity is reclaimed?
A bigger ephemeral disk
A second ECR repository
A base of tasks on the FARGATE On-Demand provider
An extra NAT gateway
«
»

Comments & Discussions