Cloud / AWS Fargate Interview questions
Last updated
1. What is AWS Fargate?
AWS Fargate is a serverless compute engine for containers. You declare the CPU and memory a workload needs, and AWS finds the capacity, runs the container, and removes it when it stops. There are no EC2 instances to choose, patch, or scale.
It is not an orchestrator on its own. Fargate is the compute layer behind Amazon ECS and Amazon EKS, which decide what runs and when. Each task or pod runs in its own isolated environment instead of sharing a kernel with other workloads.
You pay for the vCPU and memory a task requests, not for idle servers. The trade-off is less control: no SSH access to hosts, no privileged containers, and no custom AMIs.
Take quiz
Writing the Dockerfile for your application
Choosing which image tag to deploy
Writing your application's IAM policies
Provisioning, patching, and scaling the underlying servers
Nothing, it schedules containers by itself
An orchestrator such as ECS or EKS
An EC2 Auto Scaling group
One Lambda function per container
2. What container orchestrators does Fargate work with?
Fargate works with Amazon ECS and Amazon EKS. AWS Batch can also run jobs on Fargate. The way you ask for Fargate differs for each.
| Service | How you select Fargate | Unit of work |
| Amazon ECS | Launch type FARGATE or the FARGATE / FARGATE_SPOT capacity providers |
Task |
| Amazon EKS | A Fargate profile whose selectors match the pod's namespace and labels | Pod |
| AWS Batch | A compute environment of type FARGATE or FARGATE_SPOT |
Job |
Self-managed Kubernetes and Docker Swarm cannot use Fargate as a compute backend.
Take quiz
A Fargate profile whose selectors match the pod's namespace and labels
A nodeSelector that names an instance type
The container image name in the pod spec
A DaemonSet running on every worker node
A pod
A node group
A task
A Lambda invocation
3. What is a task definition in Fargate?
A task definition is the JSON blueprint ECS uses to launch a task. It lists the containers, images, ports, environment variables, log settings, and IAM roles. Revisions are grouped under a family such as orders-api:7.
Fargate adds a few mandatory settings: networkMode must be awsvpc, requiresCompatibilities must include FARGATE, and cpu and memory must be set at the task level.
{ "family": "orders-api", "networkMode": "awsvpc", "requiresCompatibilities": ["FARGATE"], "cpu": "512", "memory": "1024", "executionRoleArn": "arn:aws:iam::111122223333:role/ecsTaskExecutionRole", "containerDefinitions": [{ "name": "api", "image": "111122223333.dkr.ecr.us-east-1.amazonaws.com/orders-api:1.4.2", "essential": true, "portMappings": [{"containerPort": 8080}] }] }
Revisions are immutable. Editing a definition registers a new revision, and running tasks keep the revision they started with until you update the service.
Take quiz
Only per container, never per task
At the task level in the task definition
In the ECS cluster settings
On the service's load balancer
Running tasks restart automatically with the change
The existing revision is overwritten in place
The change is rejected until the family is deleted
A new revision is created and the old one stays unchanged
4. What CPU and memory combinations does Fargate support?
Fargate only accepts fixed pairings of task-level vCPU and memory. Each CPU size has its own memory range, so you cannot mix them freely.
| Task vCPU | Memory range (Linux) | Step |
| 0.25 | 0.5, 1, 2 GB | fixed values |
| 0.5 | 1 - 4 GB | 1 GB |
| 1 | 2 - 8 GB | 1 GB |
| 2 | 4 - 16 GB | 1 GB |
| 4 | 8 - 30 GB | 1 GB |
| 8 | 16 - 60 GB | 4 GB |
| 16 | 32 - 120 GB | 8 GB |
In the task definition, CPU is written in units (256, 512, 1024 and so on) and memory in MiB. An invalid pair fails at registration. Windows tasks start at 1 vCPU and top out at 4 vCPU.
You are billed for the size you configure, so pick the smallest combination that covers peak usage.
Take quiz
0.25 vCPU with 8 GB
2 vCPU with 2 GB
1 vCPU with 4 GB
4 vCPU with 4 GB
For the vCPU and memory configured, not the amount used
For the average CPU actually consumed
For the peak memory seen in CloudWatch
For memory only, since CPU is free
5. What is the awsvpc network mode in Fargate?
awsvpc gives every task its own elastic network interface (ENI) with a private IP from your subnet and its own security groups. It is the only network mode Fargate supports.
Containers in the same task share that interface, so they reach each other over localhost. Port mappings also stay simple, because each task has its own IP and there is no host port to collide on.
Two practical consequences follow:
- Every running task consumes one IP address, so subnets must be sized for peak task count including rolling deployments.
- In a public subnet the task needs
assignPublicIp=ENABLEDto reach the internet. In a private subnet it needs a NAT gateway or VPC endpoints.
Load balancers register tasks by IP, so target groups must use target type ip.
Take quiz
Through the Docker bridge on the host
Only through the load balancer
By using each other's public IPs
Over localhost, because they share the task's network interface
All tasks in a subnet share one IP address
Every task takes a private IP from the subnet for its ENI
Subnets only matter for the EC2 launch type
Each task needs its own Elastic IP
6. What is a Fargate platform version?
A platform version identifies the runtime environment a Fargate task runs on: the kernel, container runtime, and ECS agent. It decides which features the task can use.
Linux and Windows have independent numbering. The LATEST Linux version is 1.4.0 and the LATEST Windows version is 1.0.0. Version 1.4.0 is the one that brought per-task ENIs, EFS support, and configurable ephemeral storage.
AWS patches the platform by publishing new revisions of a version. A running task is never patched in place. New tasks start on the latest revision, and old ones are retired and replaced. Services handle this automatically, but standalone tasks you launched yourself have to be started again.
Take quiz
1.4.0
1.0.0
2.0.0
1.3.0
AWS patches the running task in place
You SSH in and run the package manager
It is retired and replaced by a new task on the patched revision
The task restarts itself nightly
7. What is the difference between task role and task execution role?
Both are IAM roles on a task, but they serve different callers. The execution role is used by ECS and Fargate to set the task up. The task role is used by your application code at runtime.
| Task execution role | Task role | |
| Used by | ECS agent / Fargate infrastructure | Your application inside the container |
| Typical permissions | Pull from ECR, write to CloudWatch Logs, read Secrets Manager / SSM values for the secrets block |
Read an S3 bucket, write to DynamoDB, publish to SQS |
| When it is used | Before and during container start | While the code is running |
| Task definition field | executionRoleArn |
taskRoleArn |
A frequent mistake is putting S3 or DynamoDB permissions on the execution role. The application never uses it, so the calls fail with AccessDenied.
Take quiz
The task execution role
The task role
The ECS service-linked role
The EC2 instance profile of the host
The task role
The role attached to the load balancer
The IAM user who ran the deployment
The task execution role
8. What are the ephemeral storage limits for Fargate tasks?
Every Fargate task gets 20 GiB of ephemeral storage by default. On Linux platform version 1.4.0 or later you can raise it to a maximum of 200 GiB with the ephemeralStorage setting.
"ephemeralStorage": { "sizeInGiB": 100 }
The space is shared by the pulled image layers, each container's writable layer, and any bind-mount volumes the containers share. It exists only for the life of the task, so everything is lost when the task stops.
Storage above the default 20 GiB is billed per GB-hour. If data must outlive the task, use EFS or an EBS volume instead of growing ephemeral storage.
Take quiz
20 GiB
100 GiB
200 GiB
1 TiB
It is lost
It is copied to S3 automatically
It moves to the next task in the service
It stays on the host for 24 hours
9. What is Fargate Spot?
Fargate Spot runs ECS tasks on spare Fargate capacity at a discount of up to about 70% compared with On-Demand. The catch is that AWS can reclaim the capacity, so a task can be interrupted.
When that happens, the task gets a two-minute warning: a SIGTERM is sent to the containers, and after the stopTimeout (up to 120 seconds on Fargate) they are killed. Your code should shut down cleanly inside that window.
Spot suits workloads that tolerate interruption: batch jobs, queue workers, CI runners, and stateless services running at the same time as On-Demand tasks. It is a poor fit for a single-task service or anything holding state in memory.
You enable it through the FARGATE_SPOT capacity provider.
Take quiz
Thirty minutes by email
None, it is killed instantly
Ten seconds, delivered as SIGKILL
Two minutes, delivered as SIGTERM
A single-instance primary database
A stateless queue worker that can retry a message
A task holding unsaved user sessions in memory
A service that must never lose a task
10. How is AWS Fargate priced?
Fargate bills for the vCPU and memory configured on a task, charged per second with a one-minute minimum for Linux. Billing runs from the moment the image pull starts until the task terminates, so startup time costs money too.
Several levers change the rate:
| Factor | Effect on price |
| ARM64 (Graviton) | Roughly 20% lower than x86 for vCPU and memory |
| Fargate Spot | Up to about 70% off, with interruption risk |
| Compute Savings Plans | Discount for a 1 or 3 year hourly spend commitment |
| Ephemeral storage above 20 GiB | Extra charge per GB-hour |
| Windows containers | Higher rate that includes the Windows license |
Remember the costs around the task as well: NAT gateway data processing, load balancers, CloudWatch Logs ingestion, and cross-AZ traffic often add up faster than people expect.
Take quiz
When the container image pull begins
When the first request is served
When the service is created
When the task passes its load balancer health check
Reserved Instances for EC2
Fargate Spot
Compute Savings Plans
A larger ephemeral storage volume
11. What is an ECS service on Fargate?
An ECS service keeps a specified number of copies of a task definition running and replaces any that fail. It is how you run long-lived workloads such as web APIs on Fargate.
The scheduler compares desiredCount with the running count. If a task crashes, fails its health check, or is retired, the scheduler starts a replacement, possibly in another Availability Zone.
A service also handles:
- registering and deregistering tasks with an Application or Network Load Balancer
- rolling out new task definition revisions with deployment settings
- scaling the task count through Application Auto Scaling
Fargate only supports the REPLICA scheduling strategy. The DAEMON strategy, which places one task per instance, is not available because there are no instances to place on.
Take quiz
Leaves the count lower until you intervene
Starts a replacement to restore the desired count
Restarts the whole cluster
Emails the task owner and waits
DAEMON
Both REPLICA and DAEMON
Neither, services are not allowed on Fargate
REPLICA
12. How do you run a standalone task on Fargate?
Use aws ecs run-task (or the equivalent API call). It starts one or more tasks that run to completion and are not restarted, which suits batch jobs and one-off scripts.
aws ecs run-task \ --cluster prod \ --launch-type FARGATE \ --task-definition report-job:3 \ --count 1 \ --network-configuration "awsvpcConfiguration={subnets=[subnet-0a1b2c],securityGroups=[sg-0d4e5f],assignPublicIp=DISABLED}" \ --overrides '{"containerOverrides":[{"name":"job","command":["python","run.py","--date","2026-10-01"]}]}'
The network configuration is required because of awsvpc. The overrides block lets you change the command or environment for this run without registering a new revision.
Once the essential container exits, the task moves to STOPPED and billing ends. Check stoppedReason and the container exit code to see how it went.
Take quiz
Fargate tasks cannot run without a public IP
The command needs a VPN tunnel
Each task needs subnets and security groups for its own ENI
It selects the Availability Zone's instance type
Change the command or environment for one run without a new revision
Raise the task above the maximum vCPU limit
Switch the task to the EC2 launch type
Skip the task execution role
13. What CPU architectures does Fargate support?
Fargate runs Linux tasks on X86_64 and ARM64 (AWS Graviton). Windows tasks run on X86_64 only.
You choose the architecture in the task definition:
"runtimePlatform": { "cpuArchitecture": "ARM64", "operatingSystemFamily": "LINUX" }
ARM64 tasks are typically about 20% cheaper at the same size, and many workloads run as fast or faster. The image must match, though. If an x86 image is launched on an ARM64 task, the container fails with an exec format error.
Build multi-architecture images with docker buildx build --platform linux/amd64,linux/arm64 so one tag works for both, and check that native dependencies have ARM builds.
Take quiz
CannotPullContainerError due to a missing role
InsufficientFreeAddressesInSubnet
Task failed ELB health checks
exec format error
In the cluster's instance type setting
In runtimePlatform.cpuArchitecture of the task definition
In the ALB listener rules
In the VPC route table
14. What operating systems does Fargate support?
Fargate supports Linux and Windows Server containers. Linux is the common case and has the most features.
For Windows you pick the Windows Server version in runtimePlatform.operatingSystemFamily, such as WINDOWS_SERVER_2022_CORE or WINDOWS_SERVER_2019_FULL. It must match the Windows Server version the container image was built on.
Windows tasks differ in a few ways:
- They use platform version 1.0.0, numbered separately from Linux.
- Tasks start at 1 vCPU, so there are no 0.25 or 0.5 vCPU sizes, and the maximum is 4 vCPU.
- Only X86_64 is available, with no ARM64 and no Fargate Spot.
- The hourly rate is higher because it includes the Windows license.
Take quiz
The Windows Server version the container image was built on
The Linux kernel version of the host
The ECS cluster's creation date
The version of the AWS CLI in use
They can use 0.25 vCPU sizes
They can run on ARM64 Graviton
They start at 1 vCPU and run on X86_64 only
They share platform version numbers with Linux
15. How do you view container logs for Fargate tasks?
Fargate has no host to SSH into, so logs have to be shipped. The simplest option is the awslogs log driver, which sends container stdout and stderr to CloudWatch Logs.
"logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/orders-api", "awslogs-region": "us-east-1", "awslogs-stream-prefix": "api" } }
The task execution role needs logs:CreateLogStream and logs:PutLogEvents. Add logs:CreateLogGroup if you set awslogs-create-group to true.
Streams are named prefix/container-name/task-id. You can follow them live with aws logs tail /ecs/orders-api --follow. For routing to other destinations such as S3, Firehose, or third-party tools, use FireLens.
Take quiz
The task role
The task execution role
The ECS service-linked role only
The user running the AWS CLI
Only files under /var/log
Only CPU and memory metrics
Network packet captures
Standard output and standard error
16. What is ECS Exec and how do you use it?
ECS Exec opens an interactive shell, or runs a single command, inside a running container without SSH or a bastion. It rides on AWS Systems Manager Session Manager.
To use it on Fargate:
- Set
enableExecuteCommandon the service or inrun-task. Only tasks started afterwards get it. - Add the
ssmmessages:CreateControlChannel,CreateDataChannel,OpenControlChannelandOpenDataChannelpermissions to the task role. - Make sure the task can reach the SSM endpoints, through NAT or an
ssmmessagesVPC endpoint. - Install the Session Manager plugin for the AWS CLI and run the command below.
aws ecs execute-command --cluster prod --task <task-id> \ --container api --interactive --command "/bin/sh"
It is valuable for debugging, but in production you should log the sessions to CloudWatch or S3 through the cluster's execute-command configuration.
Take quiz
The task execution role
The load balancer's role
The task role
The AWS Config service role
Start new tasks, since only new tasks get the feature
Nothing, running tasks pick it up immediately
Reboot the underlying EC2 instance
Recreate the VPC
17. What is a capacity provider strategy?
A capacity provider strategy tells ECS where to place tasks and in what proportion. On Fargate the two providers are FARGATE (On-Demand) and FARGATE_SPOT.
Each entry has two knobs:
- base: the minimum number of tasks to run on that provider. Only one provider in a strategy can have a base.
- weight: the relative share of the remaining tasks.
capacityProviderStrategy=[ {"capacityProvider":"FARGATE", "base":2, "weight":1}, {"capacityProvider":"FARGATE_SPOT", "weight":3} ]
With ten tasks, two land on FARGATE as the base. The other eight split 1:3, so about two more go On-Demand and six go to Spot. If Spot capacity dries up, the guaranteed base keeps the service available.
Take quiz
The maximum price per task
The share of tasks as a percentage
The number of Availability Zones used
The minimum number of tasks placed on that provider
To FARGATE_SPOT, since it is cheaper
To FARGATE, because of the base
One to each provider
They wait until Spot capacity is confirmed
18. How do you pass secrets to Fargate containers?
Reference the secret in the container's secrets block instead of putting the value in environment. ECS fetches it from AWS Secrets Manager or SSM Parameter Store when the task starts and injects it as an environment variable.
"secrets": [ { "name": "DB_PASSWORD", "valueFrom": "arn:aws:secretsmanager:us-east-1:111122223333:secret:prod/db-AbCdEf:password::" } ]
The task execution role must be allowed secretsmanager:GetSecretValue (or ssm:GetParameters) and kms:Decrypt if a customer-managed key is used.
Values are resolved only at launch. Rotating a secret does not update running tasks, so you must start new ones, for example with update-service --force-new-deployment. A task that cannot fetch the secret fails with a ResourceInitializationError.
Take quiz
When the task starts
Every time the application reads the variable
Once a day on a schedule
Only when the secret is rotated
The task role
The load balancer role
The task execution role
The Secrets Manager service role
19. How do you define a Fargate service using Terraform?
You need a task definition with Fargate settings and an aws_ecs_service that references it with launch_type = "FARGATE" and a network configuration.
resource "aws_ecs_task_definition" "api" { family = "orders-api" requires_compatibilities = ["FARGATE"] network_mode = "awsvpc" cpu = 512 memory = 1024 execution_role_arn = aws_iam_role.exec.arn task_role_arn = aws_iam_role.app.arn container_definitions = file("api.json") } resource "aws_ecs_service" "api" { name = "orders-api" cluster = aws_ecs_cluster.main.id task_definition = aws_ecs_task_definition.api.arn desired_count = 2 launch_type = "FARGATE" network_configuration { subnets = var.private_subnet_ids security_groups = [aws_security_group.api.id] assign_public_ip = false } }
Three details trip people up. network_mode must be awsvpc, and cpu and memory must form a valid Fargate pair. The execution role pulls the image and writes logs, while the task role carries the application's own permissions. The service needs subnets and security groups because each task gets its own ENI.
Add a load_balancer block if the service sits behind an ALB, and consider lifecycle { ignore_changes = [desired_count] } when auto scaling manages the task count.
Take quiz
instance_type = "fargate"
requires_compatibilities = ["FARGATE"]
capacity = "serverless"
compute_mode = "awsvpc"
Because desired_count is not supported on Fargate
To stop the service from using the load balancer
To force every apply to restart tasks
So Terraform doesn't reset the count that auto scaling changed
20. What happens when an essential container in a Fargate task exits?
If a container marked essential: true stops, for any reason, ECS stops every other container in the task and the task moves to STOPPED. A task definition needs at least one essential container.
In a service, the scheduler then launches a replacement task. For a standalone task, it simply ends.
Containers marked essential: false can exit without affecting the rest. That is the right setting for init steps and helpers whose failure should not take the app down. Use dependsOn to control start order:
"containerDefinitions": [ {"name": "migrate", "essential": false, ...}, {"name": "api", "essential": true, "dependsOn": [{"containerName": "migrate", "condition": "SUCCESS"}]} ]
Here api starts only after migrate exits with code 0. Make log-router and monitoring sidecars essential only if you really want the task to die without them.
Take quiz
Restarts only that container forever
Keeps the task running without it
Stops all other containers in the task
Moves the container to another task
The dependency container to exit with code 0
The dependency's first log line
The load balancer health check
The next platform revision
21. What is the difference between Fargate and EC2 launch types?
Both run the same ECS tasks. The difference is who manages the servers. With Fargate, AWS does. With the EC2 launch type, you run and patch a fleet of instances that ECS places tasks on.
| Aspect | Fargate | EC2 launch type |
| Server management | None | You manage AMIs, patching, scaling |
| Billing | Per task vCPU and memory, per second | Per instance, whether full or not |
| Task sizes | Fixed vCPU/memory pairs | Anything the instance can hold |
| GPUs | Not supported | Supported |
| Privileged containers, host access | Not allowed | Allowed |
| Daemon services | Not supported | Supported (DAEMON) |
| Network mode | awsvpc only | awsvpc, bridge, host |
| Isolation | Own microVM per task | Tasks share the instance kernel |
Fargate trades flexibility for less operational work. EC2 gives you control and can be cheaper at sustained high utilization.
Take quiz
Fargate cannot use ECS services
EC2 launch type cannot use load balancers
Fargate runs only Windows containers
Who manages the servers that run the tasks
Task IAM roles
GPU-backed tasks
awsvpc networking
Auto scaling of services
22. When should you choose Fargate over EC2?
Choose Fargate when you would rather spend engineering time on the application than on the cluster. It is the better default for most new container workloads.
It fits especially well when:
- Traffic is spiky or unpredictable. You pay per task and there is no spare instance capacity to buy up front.
- The team is small. No AMI pipelines, patching windows, or capacity planning.
- Workloads are batch or event-driven and run briefly, so an always-on fleet would sit idle.
- Isolation matters. Each task has its own kernel boundary, which helps with multi-tenant or compliance-sensitive services.
- Tasks are small. Packing many 0.25 to 2 vCPU tasks onto instances rarely saves enough to offset the operations effort.
Start on Fargate and move a service to EC2 only when you hit a concrete limit, such as a GPU requirement or a measured cost gap at steady load.
Take quiz
A small team running spiky workloads without a platform engineer
A job that needs GPUs
A daemon agent required on every host
A container that needs privileged host access
Batch jobs are free on Fargate
Tasks start in under a millisecond
You pay only while each task runs, with no idle fleet
It reserves instances for a year automatically
23. When would you choose EC2 over Fargate?
Choose the EC2 launch type when a requirement is something Fargate cannot offer, or when sustained utilization makes managing instances worth it.
| Situation | Why EC2 wins |
| GPU or specialized hardware | Fargate has no GPU support |
| Privileged containers, host mounts, custom kernel modules | Fargate blocks host-level access |
| Daemon or agent per host (security, monitoring) | DAEMON scheduling is EC2 only |
| Very large or odd task sizes | Fargate caps at 16 vCPU / 120 GB with fixed pairs |
| Steady, high-utilization fleet | Reserved or Savings Plan instances, tightly packed, can cost less |
| Need for local NVMe or instance storage | Not available on Fargate |
A mixed setup is common: Fargate for web services and jobs, EC2 for the one GPU or host-level workload.
Take quiz
Using an Application Load Balancer
Running GPU inference containers
Using task IAM roles
Using CloudWatch Logs
Fargate charges for every request
EC2 has no networking charges
Instances are billed only when tasks run
Tightly packed tasks at high steady utilization waste little capacity
24. Explain the lifecycle of a Fargate task?
An ECS task on Fargate moves through a fixed set of states. Knowing them tells you where to look when something stalls.
stateDiagram-v2 [*] --> PROVISIONING PROVISIONING --> PENDING PENDING --> ACTIVATING ACTIVATING --> RUNNING RUNNING --> DEACTIVATING DEACTIVATING --> STOPPING STOPPING --> DEPROVISIONING DEPROVISIONING --> STOPPED STOPPED --> [*]
- PROVISIONING: Fargate reserves capacity and creates the task's ENI in your subnet.
- PENDING: the agent is starting; images are pulled and secrets fetched.
- ACTIVATING: containers are started and the task is being registered with load balancers or service discovery.
- RUNNING: the task is serving. Health checks and metrics apply.
- DEACTIVATING: the task is being removed from load balancers and deregistered.
- STOPPING: containers get
SIGTERM, thenSIGKILLafterstopTimeout. - DEPROVISIONING: the ENI is detached and resources are released.
- STOPPED: the task is finished. Its
stoppedReasonstays visible for a short time.
A task stuck in PROVISIONING or PENDING usually points to subnet IPs, image pull, or secret problems, while repeated STOPPED tasks point to application or health check failures.
Take quiz
RUNNING
DEACTIVATING
PROVISIONING
STOPPED
Containers receive SIGTERM and then SIGKILL after stopTimeout
Images are pulled again for the next task
The ENI is created
The task is registered with the load balancer
25. Explain the execution flow of launching a Fargate task?
From the moment you call RunTask or a service needs a new task, ECS and Fargate coordinate several steps before your code runs.
sequenceDiagram participant U as You or Service Scheduler participant E as ECS Control Plane participant F as Fargate participant V as Your VPC participant R as ECR / Secrets U->>E: RunTask (task definition, subnets, SGs) E->>F: Request capacity (cpu, memory, AZ) F->>V: Create ENI in chosen subnet F->>R: Pull image, fetch secrets (execution role) F->>F: Start containers in microVM F->>E: Report RUNNING E->>V: Register task IP with load balancer
- ECS validates the task definition and picks a subnet and Availability Zone.
- Fargate allocates an isolated microVM with the requested vCPU and memory.
- An ENI is attached to the task in your subnet using your security groups.
- The agent assumes the task execution role, pulls the image, and resolves secrets.
- Containers start in
dependsOnorder and any container health checks begin. - The task reports RUNNING and, for a service, the target group registers its IP.
Any failure in steps 3 or 4, such as no IPs, no route to ECR, or missing permissions, shows up as a stopped task with a stoppedReason.
Take quiz
The task role
The service-linked role of the load balancer
The IAM user that called RunTask
The task execution role
In an AWS-owned VPC outside your account
In a subnet you specify in the network configuration
On an EC2 instance you manage
In the ECR registry
26. How does Fargate isolate tasks from each other?
Every Fargate task runs in its own isolated compute environment built on Firecracker microVM technology. Tasks do not share a kernel, CPU, memory, or network interface with tasks from other customers, or with your other tasks.
That is a real difference from the EC2 launch type, where all tasks on an instance share the host kernel and rely on container namespaces and cgroups for separation.
Within a single task the boundary is softer. Containers in the same task share the task's kernel, ENI, and ephemeral storage, which is why localhost networking works between them. Put workloads with different trust levels in separate tasks, not as sidecars of each other.
Isolation also explains some Fargate limits. With no shared host there is nothing to mount or inspect, so privileged mode, host networking, and host path mounts are not offered.
Take quiz
Firecracker microVMs
Shared Docker bridge networks
One EC2 instance per cluster
A single shared kernel for all customers
Nothing at all
Only the Docker socket of the host
The kernel, the ENI, and ephemeral storage
The kernel of other customers' tasks
27. How do you troubleshoot CannotPullContainerError on Fargate?
The task failed to download its image before any container started. Start with aws ecs describe-tasks and read the stoppedReason. It usually names the cause.
Check these in order:
- Network path to the registry. A task in a private subnet needs a NAT gateway, or VPC endpoints for
ecr.api,ecr.dkr, an S3 gateway endpoint (layers live in S3), andlogsif you use awslogs. In a public subnet,assignPublicIpmust beENABLED. - Security group egress. Outbound HTTPS (443) must be allowed.
- Execution role. It needs
ecr:GetAuthorizationToken,ecr:BatchGetImage, andecr:GetDownloadUrlForLayer. - Image URI and tag. A typo, a deleted tag, or the wrong region or account causes a not-found error.
- External registries. Docker Hub rate limits or missing credentials in
repositoryCredentialswill fail pulls.
If it only fails sometimes, suspect rate limits or a flaky NAT path rather than permissions.
Take quiz
Assign an Elastic IP to the container
A NAT gateway, or ECR and S3 VPC endpoints
Raise the task CPU to 4 vCPU
Enable ECS Exec on the service
ECR stores task definitions in S3
S3 provides the load balancer for ECR
The S3 endpoint holds the IAM role
Image layers are served from S3
28. How do you troubleshoot a Fargate task stuck in PENDING?
A task sitting in PROVISIONING or PENDING has not yet started its containers. Work out which phase it is stuck in: networking, image pull, or secrets.
| Symptom | Likely cause | Check |
| Task fails right away, no ENI | Subnet has no free IPs | Free IPs in the subnet; add subnets or a larger CIDR |
Stays PENDING, then CannotPullContainerError |
No route to ECR or registry | NAT / VPC endpoints, security group egress |
ResourceInitializationError |
Cannot fetch secrets or reach SSM / Secrets Manager | Execution role permissions, endpoints, KMS access |
| Stays PENDING for minutes | Very large image or slow pull | Image size, region, SOCI, endpoint throughput |
| Service shows events but no tasks | Capacity or service-quota limit | Service events, Fargate vCPU quotas |
Run aws ecs describe-services and read the events list, then describe-tasks for the stopped task's stoppedReason. The service event log is often the fastest clue. If the reason says the ENI could not be created, look at the subnet and security group first.
Take quiz
The container image uses ARM64
The log retention period is too short
The subnet has run out of free IP addresses
The task role is too permissive
ResourceInitializationError
exec format error
OutOfMemoryError
TaskFailedToStart for ELB health check
29. How do you troubleshoot exit code 137 on Fargate?
Exit code 137 means the container received SIGKILL (128 + 9). On Fargate this is most often the out-of-memory killer, though it can also appear when a container ignores SIGTERM and ECS force-kills it after stopTimeout.
To tell which one it is:
- Run
describe-tasksand readstoppedReason. OutOfMemoryError: Container killed due to memory usage confirms OOM. - Check
MemoryUtilizationin CloudWatch Container Insights for the minutes before the kill. - If the memory graph is flat and the stop was a deployment, the app probably does not handle
SIGTERM.
Fixes for the OOM case:
- Raise the task memory, within a valid CPU and memory pairing.
- Set a hard container
memorylimit lower than the task size so a leaky sidecar cannot starve the app. - Tune the runtime. For Java, use
-XX:MaxRAMPercentage=75so the heap respects the container limit. - Profile for leaks and unbounded caches.
Take quiz
It exited cleanly after finishing its work
It failed to parse its config file
It was stopped by SIGTERM
It was killed with SIGKILL, often by the out-of-memory killer
-XX:+UseFargate
-XX:MaxRAMPercentage
-Xdebug
-XX:AutoScaleHeap
30. How does ECS service auto scaling work with Fargate?
ECS service auto scaling uses Application Auto Scaling to change the service's DesiredCount. Because Fargate has no instances, scaling the task count is the whole story. There is no cluster capacity to scale first.
| Policy type | How it works | Typical use |
| Target tracking | Holds a metric near a target, such as 60% CPU | Default choice for web services |
| Step scaling | CloudWatch alarm adds or removes tasks by thresholds | Fine control over bursts |
| Scheduled | Changes min/max or count at set times | Predictable daily peaks |
Built-in metrics are ECSServiceAverageCPUUtilization, ECSServiceAverageMemoryUtilization, and ALBRequestCountPerTarget. For queue consumers, publish a custom metric such as messages per task.
Scale-out is fast, but a new task still takes time to provision and pull its image. Keep a sensible minimum and use shorter scale-out than scale-in cooldowns so traffic spikes are handled without flapping.
Take quiz
The service's DesiredCount
The number of EC2 instances in the cluster
The CPU size of each running task
The Availability Zones in the VPC
ECSTaskDiskUsage
NetworkInterfaceCount
ALBRequestCountPerTarget
FargateSpotPrice
31. How can you reduce AWS Fargate costs?
Cost comes from task size multiplied by running time, so most savings come from smaller tasks, fewer hours, or cheaper rates.
- Right-size tasks. Compare actual CPU and memory in Container Insights with the configured size. Compute Optimizer also gives ECS on Fargate recommendations.
- Switch to ARM64 (Graviton). About 20% lower rates for compatible images.
- Use Fargate Spot for fault-tolerant tasks, mixed with a small On-Demand base.
- Buy Compute Savings Plans to cover the steady baseline.
- Scale down off-hours with scheduled scaling, especially in dev and test environments.
- Trim startup time. Billing starts at image pull, so smaller images save money on every launch.
- Cut the extras. Use VPC endpoints instead of NAT for ECR and S3 traffic, and set log retention.
Do the right-sizing first. It is free and usually the largest saving, since many services run at 10-20% of the memory they reserve.
Take quiz
Adding more sidecar containers
Moving from x86 to ARM64 Graviton
Using larger ephemeral storage
Switching to the bridge network mode
Smaller images use a cheaper CPU type
Image size is billed separately per request
Small images skip the execution role
Billing starts when the image pull begins
32. How do Savings Plans apply to Fargate?
Compute Savings Plans apply to Fargate usage automatically. You commit to a dollar amount of compute per hour for one or three years, and in return the covered usage is billed at a lower rate.
The same commitment covers EC2, Lambda, and Fargate. It is not tied to a region, instance family, or operating system, so you can move workloads between services and still use the discount.
Things to know:
- Usage above the commitment is billed at normal On-Demand rates.
- Fargate Spot is priced on its own and is not discounted further by a Savings Plan.
- The commitment is in dollars, not vCPUs, so size it to your steady baseline, not your peak.
- Unused commitment is lost for that hour. Do not cover bursty usage.
A common approach is to use the plan for baseline On-Demand tasks, run extra capacity on Spot, and review coverage and utilization in Cost Explorer.
Take quiz
As a fixed number of vCPUs
As a count of running tasks
As a dollar amount of compute spend per hour
As a specific instance family
It is billed at normal On-Demand rates
It is stopped automatically
It is billed at the Spot rate
It is free for the remainder of the hour
33. How do you handle Fargate Spot interruptions?
Treat interruption as a normal event, not a failure. When AWS reclaims capacity, ECS sends SIGTERM to the task's containers about two minutes before termination, and SIGKILL follows after stopTimeout (maximum 120 seconds on Fargate).
Make the application cooperate:
- Catch SIGTERM. Stop accepting new work, finish or checkpoint what is in flight, and exit with code 0.
- Raise
stopTimeoutin the container definition so the shutdown has time, up to the 120 second limit. - Make work idempotent. A queue message that was not deleted will return after its visibility timeout and be processed again.
- Keep an On-Demand base. Use a capacity provider strategy with
baseonFARGATEso the service never drops to zero. - Spread the risk. Use several subnets in different Availability Zones.
You can also watch the ECS Task State Change events in EventBridge. A stop caused by Spot reclaim is reported in the task's stoppedReason, which helps you track how often it happens.
Take quiz
Ignore it and keep processing
Restart itself immediately
Delete its own task definition
Stop taking new work, checkpoint in-flight work, and exit cleanly
Put every task on FARGATE_SPOT only
Set a base on the FARGATE capacity provider
Disable the load balancer health check
Increase the ephemeral storage
34. How do you attach persistent storage to Fargate tasks?
Ephemeral storage disappears with the task, so durable data needs another option. Fargate supports two: Amazon EFS and Amazon EBS.
| Ephemeral | Amazon EFS | Amazon EBS | |
| Lifetime | Task only | Independent of tasks | Tied to the task by default |
| Sharing | Containers in one task | Many tasks at once | One task at a time |
| Type | Local disk | NFS file system | Block volume |
| Best for | Scratch space, caches | Shared files, CMS uploads, ML model files | Single-writer databases, low-latency block IO |
| Size | 20 - 200 GiB | Elastic | Configured per volume |
For EFS, create mount targets in the task's subnets, allow NFS (port 2049) between the task and mount target security groups, and declare an efsVolumeConfiguration in the task definition. Access points and IAM authorization give per-service control.
Data stores like S3 or DynamoDB are often better than any disk. Prefer them when the application can use an API instead of a file system.
Take quiz
Amazon EFS
Ephemeral storage
Amazon EBS
The instance store
TCP 22 for SSH
TCP 3306 for MySQL
TCP 2049 for NFS
UDP 53 for DNS only
35. How do you attach an EBS volume to a Fargate task?
ECS can create and attach an encrypted EBS volume to each Fargate task when it launches. You mark the volume as configuredAtLaunch in the task definition, then supply the details when you run the task or create the service.
- In the task definition, declare a volume with
"configuredAtLaunch": trueand mount it in the container with amountPointsentry. - In
run-taskorcreate-service, passvolumeConfigurationswith amanagedEBSVolume: size, volume type, encryption, optional KMS key or snapshot. - Provide an ECS infrastructure IAM role that lets ECS manage the volume on your behalf.
"volumeConfigurations": [{ "name": "data", "managedEBSVolume": { "roleArn": "arn:aws:iam::111122223333:role/ecsInfrastructureRole", "volumeType": "gp3", "sizeInGiB": 50, "encrypted": true, "filesystemType": "ext4" } }]
Each volume attaches to a single task and cannot be shared. By default the volume is deleted when the task ends, so use a snapshot or terminationPolicy settings if the data must survive.
Take quiz
mountHost set to true
configuredAtLaunch set to true
persistAfterStop set to true
efsAuthorization set to enabled
Yes, up to ten tasks
Yes, if they share a security group
Yes, but only in the same subnet
No, it attaches to a single task
36. How do you load balance traffic to Fargate tasks?
Put an Application Load Balancer (HTTP/HTTPS) or Network Load Balancer (TCP/UDP/TLS) in front of the service and register tasks through a target group with target type ip. Instance targets do not work because tasks use awsvpc and have no host instance.
flowchart LR C[Client] --> L["ALB in public subnets"] L --> T1["Task 1 IP, private subnet AZ-a"] L --> T2["Task 2 IP, private subnet AZ-b"] L --> T3["Task 3 IP, private subnet AZ-c"]
The service definition ties the pieces together:
"loadBalancers": [{ "targetGroupArn": "arn:aws:elasticloadbalancing:...:targetgroup/api/abc123", "containerName": "api", "containerPort": 8080 }], "healthCheckGracePeriodSeconds": 60
ECS registers each new task IP when it becomes ready and deregisters it on shutdown. The task security group must allow inbound traffic from the load balancer's security group on the container port. Set a grace period so slow-starting apps are not killed by early failed health checks.
Also match the target group's deregistration delay to the container's stopTimeout. That way in-flight requests finish before the task receives SIGKILL, and rolling deployments don't drop connections.
Take quiz
instance
lambda
ip
host
Traffic from the load balancer's security group on the container port
All traffic from 0.0.0.0/0 on every port
SSH from the internet
Only traffic from other tasks' ENIs
37. What is the difference between Service Connect and Cloud Map discovery?
Both let services find each other by name. Cloud Map service discovery is plain DNS. ECS Service Connect adds a proxy sidecar that handles the connection, and it uses a Cloud Map namespace underneath.
| Aspect | Cloud Map service discovery | ECS Service Connect |
| Mechanism | DNS records for task IPs | ECS-managed proxy sidecar next to your container |
| Load balancing | Client-side, from DNS answers | Round robin by the proxy |
| Failure handling | Clients may cache stale IPs until TTL expires | Proxy retries and removes unhealthy tasks quickly |
| Metrics | None built in | Per-connection metrics in CloudWatch |
| Client changes | Resolve names yourself | Use short names like http://orders:8080 |
Service Connect is the better default for service-to-service calls inside a cluster, because it avoids DNS caching surprises and gives you request metrics. Plain Cloud Map still makes sense when you need DNS records for non-ECS clients or want no proxy.
Take quiz
A public IP for every task
A replacement for the VPC
An EC2 instance per service
A proxy sidecar that retries and tracks connections
It cannot resolve private IPs
Clients can keep using stale IPs until the TTL expires
It needs a GPU
It works only with Windows tasks
38. How do rolling deployments work for Fargate services?
In a rolling update (the default ECS deployment), the scheduler starts tasks from the new task definition revision and stops old ones step by step until all tasks are replaced.
Two settings control the pace:
minimumHealthyPercent(default 100): the floor of healthy tasks, as a percentage ofdesiredCount, during the rollout.maximumPercent(default 200): the ceiling of total tasks, including old and new ones.
With four tasks and the defaults, ECS can launch up to four new tasks beside the four old ones, wait for them to pass health checks, then drain the old ones. Capacity never drops below four. On Fargate there is no instance limit to block this, but you pay for the extra tasks for a few minutes and need spare subnet IPs.
Setting minimumHealthyPercent to 50 and maximumPercent to 100 replaces tasks in place with no extra cost, but at half capacity during the rollout. Pair either with the deployment circuit breaker to roll back failed releases. For traffic shifting and instant rollback, ECS also offers blue/green deployments.
Take quiz
Up to eight, with at least four healthy
Exactly four, with none healthy
Up to sixteen with no minimum
Only two at a time
Rollouts become instant with no risk
Tasks are billed double
No extra tasks are launched, but capacity halves during the rollout
Old tasks never get stopped
39. How does the ECS deployment circuit breaker work?
The deployment circuit breaker watches a rolling deployment and stops it if new tasks keep failing, so a bad release doesn't loop forever. With rollback enabled it also returns the service to the last good revision.
flowchart TD
A["New deployment starts"] --> B["Launch tasks from new revision"]
B --> C{Tasks reach steady state?}
C -- Yes --> D["Deployment completes"]
C -- No, launches or health checks keep failing --> E{Failure threshold reached?}
E -- No --> B
E -- Yes --> F["Deployment marked FAILED"]
F --> G["Roll back to last completed deployment"]
A failure is counted when a task cannot start or fails health checks. The threshold scales with the service's desired count, with a minimum of a few tasks. When it is crossed, the deployment is marked failed.
deploymentConfiguration={ "deploymentCircuitBreaker": {"enable": true, "rollback": true} }
Typical triggers on Fargate are a wrong image tag, a missing secret, an app that crashes at startup, or health check paths that return errors. Use it with CloudWatch alarms for application-level errors it cannot see.
Take quiz
Deletes the cluster
Marks the deployment failed and returns to the last good revision
Doubles the desired count
Switches the service to the EC2 launch type
Slow database queries after startup
High NAT gateway bills
Misspelled log group names after a successful deploy
New tasks that repeatedly fail to start or pass health checks
40. How do you configure health checks for Fargate tasks?
Fargate tasks can be checked at two levels, and they do different jobs. A container health check runs a command inside the container. A load balancer health check makes an HTTP or TCP request to the task's IP.
"healthCheck": { "command": ["CMD-SHELL", "curl -f http://localhost:8080/health || exit 1"], "interval": 30, "timeout": 5, "retries": 3, "startPeriod": 60 }
startPeriodgives the app time to boot before failures count.- If an essential container turns
UNHEALTHY, ECS stops the task and the service replaces it. - The target group health check decides whether the ALB sends traffic to the task. A failing target is deregistered and, in a service, replaced.
healthCheckGracePeriodSecondson the service stops ECS from replacing tasks that are still warming up.
The container command must exist in the image. A distroless or scratch image without curl will report unhealthy even if the app is fine.
Take quiz
Sets how long logs are kept
Defines the task's CPU units
Lets the app boot before failed checks are counted
Schedules the next deployment
The curl binary does not exist in the image
Distroless images cannot run on Fargate
Health checks need a public IP
The execution role lacks curl permissions
41. How do you secure Fargate workloads?
AWS secures the underlying host and platform. You secure what runs on it: identity, network, image, and secrets. This is the shared responsibility model.
| Area | What to do |
| Identity | One task role per service with least privilege. Keep the execution role narrow, limited to ECR, logs, and the specific secrets. |
| Network | Run tasks in private subnets. Give each service its own security group and allow inbound only from the load balancer's security group. |
| Image | Scan images in ECR, pin tags or digests, use minimal base images, and run as a non-root user. |
| Container | Set readonlyRootFilesystem to true and write only to explicit volumes. |
| Secrets | Use the secrets block with Secrets Manager or SSM, never plain environment values. |
| Access | Leave ECS Exec off by default. When on, log sessions to CloudWatch or S3. |
| Data | Encrypt EFS and EBS, and optionally use a customer-managed KMS key for ephemeral storage. |
Add VPC endpoints so traffic to AWS APIs stays off the internet, and use CloudTrail to audit who changed task definitions and services.
Take quiz
privileged set to false
essential set to true
assignPublicIp set to DISABLED
readonlyRootFilesystem set to true
In public subnets with open security groups
In private subnets behind a load balancer
Directly on the internet gateway
Inside the default VPC only
42. Why can't Fargate run privileged containers?
There is no host for a container to be privileged on. In Fargate you do not own the instance, and AWS keeps the platform locked down so tasks cannot affect it, or each other.
This shows up as concrete limits:
privileged: trueis rejected.- No
hostnetwork mode and no host path or/var/run/docker.sockmounts. - Linux capabilities cannot be freely added. Only
SYS_PTRACEis allowed to be added. - No custom kernel modules or sysctl changes beyond what the platform exposes.
Tools that need these rights, such as Docker-in-Docker, host monitoring agents, and eBPF tools, do not work directly. Workarounds are to build images in CodeBuild instead, use rootless build tools like Kaniko, or run the workload on the EC2 launch type.
Because the microVM boundary already isolates the task, the lost flexibility is also what gives Fargate its security guarantee.
Take quiz
SYS_PTRACE
SYS_ADMIN
NET_ADMIN
SYS_MODULE
Mount the host Docker socket
Enable privileged mode in the task definition
Use a rootless build tool or build in a service like CodeBuild
Switch the network mode to host
43. How does Fargate encrypt ephemeral storage?
On Linux platform version 1.4.0 and later, a task's ephemeral storage is encrypted at rest with AES-256 by default, using an encryption key managed by AWS Fargate. You do nothing to turn it on.
If your compliance rules require your own key, configure a customer-managed KMS key at the cluster level through managedStorageConfiguration, using the fargateEphemeralStorageKmsKeyId setting.
aws ecs create-cluster --cluster-name secure-prod \ --configuration 'managedStorageConfiguration={fargateEphemeralStorageKmsKeyId=arn:aws:kms:us-east-1:111122223333:key/1234abcd-12ab-34cd-56ef-1234567890ab}'
The key policy must allow the Fargate service to use the key. Only tasks launched after the key is configured use it. Existing tasks keep the old encryption.
This covers storage inside the task. EFS and EBS volumes are encrypted separately with their own key settings, and data in transit still needs TLS.
Take quiz
It is not encrypted unless you enable it
Encrypted at rest with AES-256 using an AWS-managed key
Encrypted only when a public IP is assigned
Encrypted with a key you must rotate manually
In each container's logConfiguration
On the load balancer listener
In the task's security group
In the cluster's managedStorageConfiguration
44. How can you speed up Fargate container startup?
Startup time is the sum of ENI creation, image pull, container start, and the health check delay before traffic arrives. Image pull and application boot are the parts you control.
- Shrink the image. Use multi-stage builds, slim or distroless bases, and remove build tools and caches.
- Use SOCI lazy loading. A Seekable OCI index lets Fargate start a container before the whole image is downloaded, which helps large images.
- Keep ECR in the same region and use VPC endpoints so pulls take a short, fast path.
- Cut application boot time. Lazy-load, trim startup work, and for Java consider class data sharing or AOT options.
- Tune health checks. A
startPeriodand grace period that match real boot time avoid false failures that restart tasks. - Scale ahead of demand. Scheduled scaling or a higher minimum hides the delay from users.
Do not rely on layer caching between tasks, since each new task may start on fresh infrastructure. Measure with the task's createdAt, pullStartedAt, pullStoppedAt, and startedAt timestamps to see where time goes.
Take quiz
Skip the ENI creation step
Run containers without an execution role
Start a container before the entire image has downloaded
Cache images permanently on your laptop
pullStartedAt and pullStoppedAt
connectivityAt and stopCode
healthStatus and lastStatus
platformVersion and capacityProviderName
45. How do you monitor Fargate tasks?
Monitoring uses four layers: metrics, logs, events, and traces. Since you cannot install agents on the host, everything is delivered through ECS features or sidecars.
| Layer | Tool | What you get |
| Service metrics | CloudWatch AWS/ECS namespace |
CPUUtilization, MemoryUtilization per service |
| Task and container metrics | Container Insights | Per-task CPU, memory, network, storage, and running task counts |
| Logs | awslogs driver or FireLens |
stdout/stderr in CloudWatch Logs or other targets |
| Events | EventBridge ECS Task State Change and Service Action | Alerts on stopped tasks, failed deployments, Spot reclaim |
| Traces and custom metrics | AWS Distro for OpenTelemetry sidecar, X-Ray | Request traces and app-level metrics |
Inside a task, the task metadata endpoint v4 (from the ECS_CONTAINER_METADATA_URI_V4 variable) returns the task ARN, container stats, and network details for the application to use.
Useful alarms: running task count below desired, memory above 80%, target group unhealthy hosts, and deployment failure events.
Take quiz
VPC Flow Logs
AWS Trusted Advisor
Amazon Inspector findings
CloudWatch Container Insights
The EC2 instance metadata service of the host
The task metadata endpoint v4
The load balancer access logs
The ECR image manifest
46. How do you use FireLens for logging on Fargate?
FireLens lets you route container logs to many destinations by running a Fluent Bit (or Fluentd) sidecar called the log router. Your app container sends logs with the awsfirelens driver, and the sidecar forwards them.
You need two parts in the task definition:
{ "name": "log_router", "image": "public.ecr.aws/aws-observability/aws-for-fluent-bit:stable", "essential": true, "firelensConfiguration": {"type": "fluentbit"} }, { "name": "api", "image": "...", "logConfiguration": { "logDriver": "awsfirelens", "options": { "Name": "firehose", "region": "us-east-1", "delivery_stream": "app-logs" } } }
The destination is chosen by the Name option. Fluent Bit supports CloudWatch, Amazon Data Firehose, S3, OpenSearch, and partner tools such as Datadog and Splunk.
You can also supply a custom config file for parsing, filtering, or multiline handling. Give the log router enough CPU and memory, because it shares the task's resources, and decide deliberately whether it is essential.
Take quiz
awsfirelens
awslogs
json-file
syslog-host
A second copy of the application
The ECS agent
Fluent Bit or Fluentd
A Prometheus server
47. How do Fargate profiles work in EKS?
A Fargate profile tells EKS which pods should run on Fargate. When a pod is created and matches a profile's selectors (a namespace and optional labels), EKS schedules it onto Fargate instead of an EC2 worker node.
flowchart TD
A["Pod created"] --> B{Matches a Fargate profile selector?}
B -- Yes --> C["Mutating webhook sets scheduler to fargate-scheduler"]
C --> D["Fargate starts a microVM sized to the pod"]
D --> E["Pod runs on its own virtual node"]
B -- No --> F["Default scheduler places it on EC2 nodes"]
What happens behind the scenes:
- A mutating admission webhook changes the pod's scheduler to
fargate-scheduler. - EKS launches a dedicated microVM for that pod, which appears as its own node.
- The pod is sized from the sum of its container resource requests, plus a small overhead for Kubernetes components, and rounded up to the next Fargate size.
- The profile supplies the pod execution role and the private subnets to use.
A pod that does not match any profile will not run on Fargate, even if a Fargate node exists. Always set resource requests, since they determine what you are billed.
Take quiz
Adding a GPU limit to the pod
Matching a Fargate profile's namespace and label selectors
Running kubectl on a bastion host
Using a node group with spot instances
The size of the cluster's node group
The pod name length
The EKS control plane version
The pod's resource requests plus a small overhead, rounded up
48. What are the limitations of Fargate on EKS?
Fargate on EKS suits stateless pods, but it removes features that rely on controlling the node. Know these before you move workloads.
| Limitation | Detail |
| No DaemonSets | There is no shared node to put a per-node pod on. Use sidecars instead. |
| No privileged pods | Also no hostNetwork, hostPort, or host path volumes. |
| No GPUs | Not supported on Fargate. |
| Private subnets only | Pods get ENIs in private subnets, so you need NAT or endpoints for outbound access. |
| Storage | EFS is supported for persistence. EBS volumes are not. |
| Load balancers | Use ALB or NLB with target type ip. |
| Per-pod cost floor | One microVM per pod, so many tiny pods cost more than packing them on a node. |
| System pods | CoreDNS needs a Fargate profile for kube-system or must run on EC2 nodes. |
DaemonSet-based tools, such as log collectors and security agents, need a replacement. Use sidecars for logging, or the platform's built-in log router for Fargate pods.
Take quiz
Deployments
Services
DaemonSets
ConfigMaps
Amazon EFS
Amazon EBS volumes
Instance store
hostPath volumes
49. How do you run scheduled jobs on Fargate?
A job on Fargate is just a task started on a timer. The usual setup is Amazon EventBridge Scheduler (or an EventBridge rule) with an ECS RunTask target.
- Create a task definition for the job with a command that exits when finished.
- Create a schedule, for example
cron(0 2 * * ? *)for 02:00 UTC daily. - Choose the ECS cluster, the task definition, a
FARGATElaunch type or capacity provider, and the network configuration. - Give the scheduler an IAM role allowed to call
ecs:RunTaskandiam:PassRolefor the task roles.
You pay only while the task runs. Add a FARGATE_SPOT strategy for jobs that can retry.
For heavier needs, there are two other options:
- AWS Batch with a Fargate compute environment, which adds queues, retries, array jobs, and dependencies.
- Step Functions with the
ecs:runTask.syncintegration, which waits for the task and lets you branch on success or failure.
Make the job idempotent, set an alarm on STOPPED tasks with a non-zero exit code, and keep a retry policy on the schedule.
Take quiz
AWS Config
Amazon Macie
AWS Shield
Amazon EventBridge Scheduler
An ECS service with desired count of one
AWS Batch with a Fargate compute environment
A bigger NAT gateway
A cross-region ECR replication rule
50. How would you architect a highly available Fargate service?
High availability on Fargate means removing single points of failure in the task layer, the network layer, and the deployment process. Fargate already manages the hosts, so the design focuses on the rest.
flowchart TD U[Users] --> ALB["ALB across 3 public subnets"] ALB --> S["ECS service, desiredCount 6"] S --> A["Tasks in private subnet AZ-a"] S --> B["Tasks in private subnet AZ-b"] S --> C["Tasks in private subnet AZ-c"] A --> D[(Database, Multi-AZ)] B --> D C --> D AS["Application Auto Scaling"] --> S
- Spread across at least two Availability Zones. List private subnets from each AZ in the service. ECS balances tasks across them.
- Run a minimum of two tasks, and more if losing one AZ must not hurt capacity. Size for N-1 AZs.
- Use an ALB with health checks and a grace period, so only ready tasks get traffic.
- Mix capacity. Put a base on
FARGATEand let extra tasks run onFARGATE_SPOT. - Auto scale on CPU or request count, with a floor equal to your availability minimum.
- Deploy safely. Rolling updates with
minimumHealthyPercent100, the circuit breaker with rollback, or blue/green. - Handle shutdown. Catch
SIGTERM, setstopTimeout, and keep deregistration delay in line. - Keep state outside the task in a Multi-AZ database, EFS, S3, or a managed cache.
For regional resilience, repeat the stack in a second region behind Route 53 health-checked routing, and replicate images with ECR cross-region replication.