Prev Next

Java / Azure Batch Interview questions

Last updated

1. What is Azure Batch? 2. What are the main components of Azure Batch? 3. What is a Batch account in Azure Batch? 4. What is a pool in Azure Batch? 5. What is a compute node in Azure Batch? 6. What is a job in Azure Batch? 7. What is a task in Azure Batch? 8. What are the common use cases for Azure Batch? 9. What is a start task in Azure Batch? 10. What are application packages in Azure Batch? 11. What are resource files in Azure Batch? 12. What are the pool allocation modes in Azure Batch? 13. What is a job manager task in Azure Batch? 14. What are the ways to interact with Azure Batch? 15. How do you submit a job to Azure Batch? 16. What are Spot VMs in Azure Batch? 17. What are the task states in Azure Batch? 18. What is a job schedule in Azure Batch? 19. What are output files in Azure Batch tasks? 20. What is the difference between Azure Batch and Azure Functions? 21. How does autoscaling work in Azure Batch? 22. How do you handle Spot VM preemption in Azure Batch? 23. When would you choose user subscription pool allocation mode? 24. How does Azure Batch schedule tasks on compute nodes? 25. What is the difference between a job preparation task and a start task? 26. How do task dependencies work in Azure Batch? 27. What are multi-instance tasks and when should you use them? 28. How does Azure Batch retry failed tasks? 29. How do you troubleshoot a failed task in Azure Batch? 30. What happens when a compute node becomes unusable? 31. How do you troubleshoot a Batch pool stuck in resizing? 32. How do you run Docker containers in Azure Batch? 33. How do you secure an Azure Batch account? 34. How do you put Azure Batch pools in a virtual network? 35. How can you optimize cost in Azure Batch? 36. How do you monitor Azure Batch workloads? 37. Why doesn't an Azure Batch command line expand environment variables? 38. What are the key environment variables available to Azure Batch tasks? 39. What is the difference between Azure Batch and Azure Kubernetes Service for batch workloads? 40. How do you run Azure Batch from Azure Data Factory? 41. Explain the lifecycle of a job in Azure Batch? 42. Explain the execution flow of a task on a compute node? 43. Explain the internal working of the Azure Batch service architecture? 44. How can you optimize task throughput in Azure Batch? 45. How do you reduce node startup time in Azure Batch? 46. How do you handle quotas in Azure Batch? 47. What happens when all tasks in an Azure Batch job complete? 48. How do you update a running Azure Batch pool? 49. What Azure Batch features have been retired? 50. Which is better for HPC on Azure, Batch or CycleCloud, and why?

1. What is Azure Batch?

Azure Batch is a managed Azure service for running large-scale parallel and high-performance computing (HPC) workloads in the cloud. You describe the work as jobs and tasks, and Batch provisions the virtual machines, schedules the tasks onto them, and scales the pool up or down.

You don't build or maintain a scheduler or cluster manager yourself. There is no extra charge for Batch itself, and you pay only for the VMs, storage and networking the workload consumes.

Typical fits are embarrassingly parallel jobs such as rendering, Monte Carlo simulations, media transcoding and data processing.

Take quiz
Which part of an Azure Batch workload do you pay for?
A per-task fee charged by the Batch scheduler
A monthly licence for each Batch account
Only the number of jobs submitted
The underlying VMs, storage and networking
Azure Batch is mainly intended for which kind of work?
Hosting always-on public websites
Large-scale parallel and HPC workloads
Managing relational database backups
Sending transactional email at scale

2. What are the main components of Azure Batch?

An Azure Batch solution is built from a small set of resources that sit inside a Batch account.

Component Role
Batch account Top-level resource that holds pools, jobs and quotas
Pool Collection of compute nodes that run tasks
Compute node A VM that executes tasks
Job Logical container for tasks, bound to one pool
Task A single command line or container run on a node

Around these sit supporting features such as start tasks, application packages, resource files and job schedules.

Take quiz
Which Batch resource is a collection of VMs that run your tasks?
Pool
Job
Batch account
Job schedule
Which resource is the top-level container that holds pools and jobs?
Task
Start task
Batch account
Application package

3. What is a Batch account in Azure Batch?

A Batch account is the Azure resource that identifies your Batch service instance in a specific region. All pools, jobs and tasks belong to one account, and every API call is made against the account's endpoint.

The account also defines the pool allocation mode (Batch service or user subscription) and carries the compute quotas that cap how many cores you can use. It can be linked to an Azure Storage account for application packages.

You authenticate to it with Microsoft Entra ID or with shared keys.

Take quiz
Which setting on a Batch account decides where pool VMs are created?
The task retry count
The pool allocation mode
The job priority
The autoscale evaluation interval
A Batch account is created in:
Every Azure region at once
Only the East US region
A job, not a subscription
A specific Azure region

4. What is a pool in Azure Batch?

A pool is the set of compute nodes on which a job's tasks run. When you create one you choose the VM size, the OS image, the node count (fixed or autoscaled), and optional settings such as a start task, virtual network and task slots per node.

Many jobs can share a pool, but a job is always attached to exactly one pool. Pools of the same account can use different VM sizes and images, so you can keep, say, a GPU pool beside a CPU pool.

Some properties, such as the VM size, can't be changed after creation. You create a new pool instead.

Take quiz
How many pools can a single Batch job run against?
Any number at the same time
None, jobs run on the account
Exactly one
Two, one dedicated and one Spot
Which pool property can't be changed after the pool is created?
VM size
Target node count
Pool metadata
Autoscale formula

5. What is a compute node in Azure Batch?

A compute node is a single Azure VM that belongs to a pool and runs tasks assigned to it. Each node runs the Batch node agent, which talks to the Batch service, downloads resource files, launches task command lines and reports state back.

A node has its own working directories and receives environment variables such as AZ_BATCH_NODE_ID. A node can be dedicated, or a Spot VM that Azure may reclaim.

Node states include idle, running, starting, startTaskFailed and unusable.

Take quiz
Which component on a node launches task command lines for Batch?
The Azure Load Balancer
The Batch account key
The pool's autoscale formula
The Batch node agent
Which node state means the node can't run tasks until it is recovered or replaced?
idle
unusable
running
starting

6. What is a job in Azure Batch?

A job is a logical collection of tasks that all run on the same pool. It carries shared settings such as priority, constraints, and the optional job manager, job preparation and job release tasks.

Job priority ranges from -1000 to 1000, and higher-priority jobs are scheduled first. By default a job stays active after its tasks finish unless you set onAllTasksComplete to terminateJob.

Take quiz
What happens by default when every task in a job has finished?
The job stays active until you terminate it or set onAllTasksComplete
The job is deleted automatically
The pool is deleted automatically
The job restarts all tasks
What range does Batch job priority use?
1 to 5
0 to 100
-1000 to 1000
1 to 10,000

7. What is a task in Azure Batch?

A task is the unit of work in Batch: a command line (or a container run) that executes on a compute node. Each task has an ID that is unique within its job and can declare resource files to download, output files to upload, environment variables and constraints.

A task finishes with an exit code. Batch records it, along with timing and failure details, so you can decide whether to retry or treat the task as failed.

Tasks in a job normally run independently and in parallel, unless you configure task dependencies.

Take quiz
What does Batch record when a task finishes?
Only the node's IP address
An exit code plus timing and failure details
A new pool definition
The Batch account key used
Task IDs must be unique within which scope?
The whole Azure subscription
The Azure region
The pool
The job

8. What are the common use cases for Azure Batch?

Batch suits workloads you can split into many independent pieces, or tightly coupled HPC runs that need several nodes at once.

  • Rendering of frames for film and visual effects.
  • Financial risk modelling and Monte Carlo simulation.
  • Media transcoding of large video libraries.
  • Genomics and scientific data processing.
  • Batch ML scoring of large datasets.
  • MPI-based engineering simulations using multi-instance tasks.

It is a poor fit for low-latency request/response services, where the node startup time and queueing model get in the way.

Take quiz
Which workload fits Azure Batch best?
Serving a latency-critical REST API
Hosting a static website
Rendering thousands of video frames in parallel
Running a single-row SQL update
Which Batch feature supports tightly coupled MPI workloads?
Multi-instance tasks
Job schedules
Application packages
Output files

9. What is a start task in Azure Batch?

A start task is a command that runs on every node each time it joins the pool, and again whenever the node is rebooted. It prepares the node, for example by installing software, mounting a share or copying reference data.

If waitForSuccess is true, Batch won't schedule tasks on the node until the start task completes successfully. A failing start task leaves the node in the startTaskFailed state.

Keep it idempotent, because it can run more than once on the same node.

Take quiz
When does a start task run?
Only once per job
Only when a task fails
Once per Batch account
Each time a node joins the pool or reboots
Why should a start task be idempotent?
It only runs on the first node
It can run again after a node reboot
Batch deletes it after success
It runs once per task

10. What are application packages in Azure Batch?

Application packages let you upload versioned zip archives of your executables to the Batch account and have them deployed to nodes automatically. The packages are stored in the linked Azure Storage account.

You reference a package and version on the pool (or task), and Batch downloads and unzips it on each node. The install location is exposed through an environment variable such as AZ_BATCH_APP_PACKAGE_<name>#<version>.

Use packages for application binaries, and resource files for task input data.

Take quiz
What do application packages store?
Versioned zip archives of your application binaries
Task output logs
Autoscale formulas
Pool network rules
Where does Batch keep the uploaded application package?
Inside the job object
On the user's laptop
In the linked Azure Storage account
In Microsoft Entra ID

11. What are resource files in Azure Batch?

Resource files are files that Batch downloads onto a node before a task (or start task) runs. Each entry points to a source such as a blob URL, a container prefix or an HTTP URL, and a destination path on the node.

Access to private blobs is granted through a SAS URL or, preferably, a managed identity on the pool. Large downloads are included in the task's start-up time, so use them for per-task inputs rather than big shared data sets.

Take quiz
What do resource files do?
Upload results after the task ends
Download input files to the node before the task runs
Resize the pool
Rotate the account keys
What is the preferred way to let nodes read private blobs?
A public anonymous container
Embedding the account key in the command line
Emailing a SAS token to the node
A managed identity on the pool

12. What are the pool allocation modes in Azure Batch?

There are two modes, chosen when you create the Batch account.

Batch service mode User subscription mode
VMs are created in subscriptions managed by Batch. VMs are created in your subscription.
Quotas are tracked at the Batch account. Quotas come from your subscription's regular VM quotas.
Simplest to set up, default choice. Needs a Key Vault and gives more control over the VM resources.

Most teams start with Batch service mode and move to user subscription mode only when they need subscription-level control.

Take quiz
In which mode are pool VMs created in your own subscription?
Batch service mode
Spot-only mode
User subscription mode
Account key mode
Which allocation mode is the default?
Batch service mode
User subscription mode
Hybrid mode
Regional mode

13. What is a job manager task in Azure Batch?

A job manager task is a special task that Batch starts first when a job is created. Its job is to create and coordinate the remaining tasks of the job, so the client doesn't have to stay connected.

It gets scheduling priority, can be relaunched if its node fails, and can end the job when the work is done. It is useful when the number of tasks depends on data discovered at run time.

It is optional, and a client can add all tasks directly instead.

Take quiz
What is the main purpose of a job manager task?
Resize the pool automatically
Delete finished jobs
Authenticate users to the account
Create and coordinate the other tasks of the job
When does Batch start the job manager task?
After all other tasks finish
First, when the job is created
Only on the last node
Once a day

14. What are the ways to interact with Azure Batch?

You can work with Batch through several interfaces, and they all call the same underlying REST API.

  • Azure portal for creating accounts and inspecting pools, jobs and tasks.
  • Azure CLI (az batch) for scripting.
  • SDKs for .NET, Python, Java and Node.js.
  • Batch Explorer, a desktop client for browsing and debugging.
  • ARM, Bicep and Terraform for provisioning accounts and pools as code.
Take quiz
Which tool is a desktop client for browsing Batch pools and tasks?
Batch Explorer
Azure Storage Explorer
Visual Studio Installer
Azure Monitor Agent
What do all Batch SDKs and tools ultimately call?
A SQL endpoint
The Azure Storage queue API only
The Batch REST API
The Windows registry

15. How do you submit a job to Azure Batch?

You create a pool (or reuse one), create a job that points to it, then add tasks to the job. Batch schedules the tasks as nodes become free.

A minimal Python flow looks like this:

from azure.batch import BatchServiceClient
import azure.batch.models as m

client = BatchServiceClient(credentials, batch_url=BATCH_URL)

client.job.add(m.JobAddParameter(
    id="render-job",
    pool_info=m.PoolInformation(pool_id="render-pool")))

client.task.add_collection("render-job", [
    m.TaskAddParameter(id=f"frame-{i}",
        command_line=f"/bin/bash -c 'render.sh {i}'")
    for i in range(100)])

Afterwards you poll task states or use the SDK's wait helpers.

Take quiz
What must a job reference when it is created?
A Spot VM size
The pool it will run on
A SQL connection string
A list of Azure regions
What is the order of steps to run work on Batch?
Tasks, then pool, then job
Job, then tasks, then pool
Account, then tasks, then delete pool
Pool, then job, then tasks

16. What are Spot VMs in Azure Batch?

Spot VMs are spare Azure capacity offered at a steep discount in exchange for the risk that Azure can reclaim (preempt) the VM when it needs the capacity back. In Batch you set a target number of Spot nodes beside the dedicated node target.

They suit interruptible, restartable workloads. When a node is preempted, tasks running on it are requeued and run elsewhere.

Spot VMs replace the older low-priority nodes, which were retired on 30 September 2025.

Take quiz
What happens to a task when its Spot node is preempted?
The task is marked successful
The whole job is deleted
The task is requeued to run on another node
The task keeps running on the lost VM
Which node type replaced low-priority nodes in Batch?
Spot VMs
Reserved nodes
Burstable nodes
Free-tier nodes

17. What are the task states in Azure Batch?

A task moves through four states.

State Meaning
active Queued and waiting to be scheduled, or waiting on dependencies or a retry
preparing Assigned to a node while the job preparation task runs
running Command line is executing on the node
completed Finished, with a success or failure result

A task that fails and has retries left returns to active for another attempt, while a successful one goes straight to completed.

Take quiz
Which state does a task return to when it fails and has retries left?
completed
preparing
deleted
active
In which state is a task's command line executing?
active
running
completed
terminating

18. What is a job schedule in Azure Batch?

A job schedule creates jobs on a recurring basis, for example a nightly report run. You define a recurrence (interval, start and end times) and a job specification, and Batch creates a new job from it at each occurrence.

Only one job from a schedule is active at a time. Use it instead of an external cron when the pool and tasks already live in Batch.

Take quiz
What does a job schedule create on each recurrence?
A new job from the job specification
A new Batch account
A new storage account
A new VM image
A job schedule is best suited to:
One-off migrations
Installing OS patches on nodes
A nightly report run
Rotating shared keys

19. What are output files in Azure Batch tasks?

Output files are a task setting that tells the node agent to upload files from the node to Azure Storage after the task runs. You give a file pattern, a destination container (SAS URL or managed identity) and an upload condition: taskSuccess, taskFailure or taskCompletion.

This matters because node disks are temporary. If the node is deleted or reimaged, anything not uploaded is lost. Uploading logs on failure also helps with later debugging.

Take quiz
Why are output files important?
They speed up the pool resize
Node disks are temporary, so results must be uploaded
They replace the start task
They encrypt the node OS disk
Which condition uploads files whether the task succeeded or failed?
taskSuccess
taskFailure
jobCompletion
taskCompletion

20. What is the difference between Azure Batch and Azure Functions?

Azure Functions is event-driven and short-lived, while Batch is built for long, compute-heavy jobs you control at the VM level.

Aspect Azure Batch Azure Functions
Trigger You submit jobs and tasks Events such as HTTP, queue or timer
Runtime Minutes to days per task Time limited per execution
Compute control You pick VM size, GPU, OS and images Platform-managed
Best for HPC, rendering, large parallel runs Light, bursty event handling

A common pattern is a Function that reacts to a blob upload and submits a Batch job for the heavy processing.

Take quiz
Which service lets you choose GPU VM sizes for long-running parallel work?
Azure Functions on a consumption plan
Azure Event Grid
Azure Batch
Azure Service Bus
What is a common way to combine the two?
A Function reacts to an event and submits a Batch job
Batch triggers the Function host to reboot
A Function replaces the Batch pool
Batch compiles the Function code

21. How does autoscaling work in Azure Batch?

With autoscale enabled, Batch evaluates a formula you supply at a fixed interval (minimum 5 minutes, default 15) and sets the target node counts from the result. It reads service metrics such as $PendingTasks, $ActiveTasks and $CPUPercent.

$samples = $PendingTasks.GetSamplePercent(TimeInterval_Minute * 5);
$tasks = $samples < 70 ? max(0, $PendingTasks.GetSample(1))
                      : max($PendingTasks.GetSample(1), avg($PendingTasks.GetSample(TimeInterval_Minute * 5)));
$TargetDedicatedNodes = min(20, $tasks);
$NodeDeallocationOption = taskcompletion;

While autoscale is on you can't resize the pool manually. Use the evaluate-formula API to test a formula before applying it.

Take quiz
What is the minimum autoscale evaluation interval in Batch?
1 second
30 seconds
24 hours
5 minutes
What must you do before you can manually resize an autoscaling pool?
Delete the job
Disable autoscale
Change the VM size
Restart the Batch account

22. How do you handle Spot VM preemption in Azure Batch?

Design tasks to be restartable. When a Spot node is preempted, Batch moves the node to the preempted state and requeues the tasks that were running on it, so they start again from the beginning on another node.

  1. Make tasks idempotent so a rerun produces the same result.
  2. Write checkpoints to Azure Storage, not the node disk.
  3. Keep a small dedicated node target for critical work, such as the job manager.
  4. Keep tasks short so little work is lost per eviction.

Batch replaces preempted nodes when capacity returns, as long as the pool target is unchanged.

Take quiz
Where should a long task save its checkpoints on a Spot pool?
In Azure Storage
On the node's temporary disk
In memory only
In the task ID
What does Batch do with tasks running on a preempted node?
Marks them succeeded
Deletes the job
Requeues them to run elsewhere
Pauses them until the node returns

23. When would you choose user subscription pool allocation mode?

Choose it when the VMs must live in your own subscription so existing governance applies. Typical reasons are using your subscription's reserved instance or savings plan discounts, applying your Azure Policy and network controls, or drawing on your subscription's VM quota instead of the Batch account quota.

The trade-off is extra setup. The Batch account needs access to an Azure Key Vault and the Batch service principal needs permissions on your subscription, and you take on more of the quota management.

If none of these needs apply, Batch service mode is simpler.

Take quiz
Which need most clearly points to user subscription mode?
Avoiding all Azure Storage usage
Applying your own reservation discounts and policies to the VMs
Running tasks without a pool
Removing the need for a Batch account
What extra requirement comes with user subscription mode?
A dedicated Active Directory forest
An on-premises gateway
A separate storage queue per task
The Batch account needs access to a Key Vault

24. How does Azure Batch schedule tasks on compute nodes?

Batch assigns tasks from active jobs to nodes that have free task slots. Higher-priority jobs are considered first, and inside a job, tasks with higher priority go first, then in order of creation.

Two pool settings shape placement:

  • taskSlotsPerNode sets how many task slots a node offers (up to 4 times the core count, capped at 256). A task can need several via requiredSlots.
  • nodeFillType is spread (default, even across nodes) or pack (fill one node before the next).

Use pack with autoscale so idle nodes can be released sooner.

Take quiz
What does the 'pack' node fill type do?
Spreads tasks evenly over every node
Runs each task twice
Fills one node's slots before using the next node
Places tasks on the newest node only
Which setting defines how many tasks a node can run concurrently?
taskSlotsPerNode
jobPriority
onAllTasksComplete
retentionTime

25. What is the difference between a job preparation task and a start task?

Both prepare a node, but they run at different scopes and times.

Aspect Start task Job preparation task
Scope Pool Job
Runs when Node joins the pool or reboots Before the first task of that job runs on a node
Typical use Install software shared by all jobs Copy data specific to one job
Cleanup None built in Pairs with a job release task

Put pool-wide setup in the start task, and job-specific setup in the job preparation task.

Take quiz
Which task runs once per node before the first task of a particular job?
Start task
Job manager task
Output file task
Job preparation task
Which task is paired with the job preparation task for cleanup?
Start task
Job release task
Multi-instance coordination task
Retry task

26. How do task dependencies work in Azure Batch?

Task dependencies let a task wait until other tasks complete successfully. You must enable them by setting usesTaskDependencies to true when the job is created, and you can't switch it on later.

A task can then declare dependsOn with specific task IDs or a range of IDs. Batch keeps the task in the active state until its dependencies succeed, and runs it as soon as they do. By default, if a dependency fails, the dependent task is never run, though exit conditions can change that with dependencyAction.

A typical use is a final merge step that waits for many parallel tasks.

Take quiz
When must usesTaskDependencies be set?
When the job is created
After the first task finishes
On the pool, at any time
At the Batch account level only
What state is a dependent task in while it waits?
running
completed
active
preparing

27. What are multi-instance tasks and when should you use them?

A multi-instance task runs across several nodes at once, which is how Batch supports MPI and other tightly coupled workloads. You set the number of instances and a coordination command line.

Batch picks a primary node and subtasks. Each node runs the coordination command first (for example, starting the MPI runtime), and once all are ready, the primary runs the application command line. The pool must have inter-node communication enabled.

Use them for simulations and solvers where processes exchange data. For independent work, use ordinary tasks.

Take quiz
Which pool setting is required for multi-instance MPI tasks?
Spot nodes only
Inter-node communication enabled
Autoscale disabled
A job schedule
What runs first on every node of a multi-instance task?
The output file upload
The job release task
The autoscale formula
The coordination command line

28. How does Azure Batch retry failed tasks?

Retries are controlled by maxTaskRetryCount in the task constraints. A task is retried when it exits with a non-zero code until the retry count is used up. Set it to 0 for no retries, or -1 for unlimited retries.

The count covers failures of the task itself. If a node is lost or preempted, the task is requeued independently of this count.

You can also use exit conditions to decide per exit code whether to retry, ignore the failure or terminate the job.

Take quiz
Which value of maxTaskRetryCount means unlimited retries?
0
1
-1
100
What controls the action taken for a specific exit code?
Exit conditions
Job priority
Pool allocation mode
Application package version

29. How do you troubleshoot a failed task in Azure Batch?

Start with the task's exit code and failureInfo, then read its output files on the node.

  1. Check the task's state, executionInfo.exitCode and failureInfo (category and message) in the portal or API.
  2. Read stdout.txt and stderr.txt from the task directory.
  3. If the task never started, check for resource file download errors and start task failures.
  4. If the node was lost, check the node's state and errors, and the pool's resize errors.
  5. Log in with Batch Explorer, or RDP/SSH where allowed, to inspect the node directly.

A UserError failure category usually means your command or inputs, while ServerError points to the service.

Take quiz
Which two files hold a task's console output on the node?
task.log and node.cfg
output.json and input.json
batch.out and batch.err only on the job
stdout.txt and stderr.txt
A failureInfo category of UserError most likely means:
A Batch service outage
A problem with your command line or inputs
A billing issue
A DNS failure in Azure

30. What happens when a compute node becomes unusable?

Batch marks the node unusable and stops scheduling work on it. Tasks that were running there are requeued and run on another node, and this doesn't consume their retry count.

Common causes are a VM fault, a platform issue or an OS-level failure. You can reboot or reimage the node, or remove it from the pool and let a resize or autoscale replace it.

A node in startTaskFailed is a separate case: the node is healthy but your start task failed, so fix the script rather than the VM.

Take quiz
What happens to tasks running on a node that becomes unusable?
They are requeued onto another node
They are deleted permanently
They are marked successful
They continue on the unusable node
Which state points to a failing start task script rather than a VM fault?
unusable
idle
startTaskFailed
leavingPool

31. How do you troubleshoot a Batch pool stuck in resizing?

Look at the pool's resizeErrors first; they usually name the cause.

  • Quota reached: not enough cores for the VM family or the account. Request a quota increase or use another size or region.
  • Start task failed: nodes are created but never become usable.
  • Resize timeout: the default timeout is 15 minutes and Batch stops allocating if it passes.
  • Capacity or networking: the VM size isn't available in that region or zone, or the subnet has too few free IPs.

After fixing the cause, resize the pool again.

Take quiz
Where do you find why a pool resize failed?
The job's priority field
The pool's resizeErrors
The application package list
The Batch account's access keys
Which issue can stop a resize even when your start task is fine?
Too many completed tasks
A long job ID
A missing job schedule
Insufficient core quota for the VM family

32. How do you run Docker containers in Azure Batch?

You can run tasks inside containers on a pool whose VM image supports containers, such as the Ubuntu container-ready images or a custom image with Docker installed.

  1. Set containerConfiguration on the pool, optionally listing images to pre-fetch and registry credentials.
  2. Set containerSettings on the task, with the image name and any run options.

Batch pulls the image, mounts the task working directory and runs the task command line inside the container. Pre-fetching images on the pool avoids a pull on each node's first task.

Take quiz
Which task property specifies the image a containerised task runs in?
resourceFiles
outputFiles
containerSettings
userIdentity
Why list images in the pool's containerConfiguration?
To pre-fetch them and avoid pulls at task time
To raise the node quota
To enable autoscale
To encrypt the OS disk

33. How do you secure an Azure Batch account?

Layer several controls rather than relying on one.

  • Authenticate with Microsoft Entra ID and Azure RBAC, and avoid shared keys where you can.
  • Use managed identities on pools so tasks reach Storage and Key Vault without secrets in command lines.
  • Deploy pools into a virtual network and restrict access with NSGs and a private endpoint for the account.
  • Store secrets in Azure Key Vault, not in task definitions.
  • Use customer-managed keys if your policy requires them, and enable diagnostic logs.
Take quiz
What is the safer alternative to shared keys for Batch API calls?
Anonymous access
Embedding keys in task command lines
A shared password file
Microsoft Entra ID authentication
Where should secrets used by tasks be kept?
In the task ID
Azure Key Vault
In job metadata
In the pool name

34. How do you put Azure Batch pools in a virtual network?

Specify a subnet ID in the pool's network configuration at creation time. The nodes then get private IPs from that subnet and can reach on-premises or peered resources.

Points to plan for:

  • The subnet needs enough free addresses for the maximum pool size.
  • Nodes must be able to talk to the Batch service, either through the default inbound rules or the newer node management private endpoint model.
  • Nodes need outbound access to Azure Storage and any registry you use.
  • You can disable public IP addresses on nodes, and then you need an outbound route such as NAT.

Network settings can't be changed on an existing pool, so create a new one.

Take quiz
What do you provide to place a pool in a VNet?
The subnet ID in the pool network configuration
The storage account name
A job schedule ID
The Batch account key
What must the subnet have for the pool to scale?
Only one free address
A public DNS zone per node
Enough free IP addresses for the maximum node count
No NSG at all

35. How can you optimize cost in Azure Batch?

Cost follows node-hours, so the goal is to run fewer and cheaper nodes for less time.

  1. Use Spot VMs for interruptible work, keeping dedicated nodes for critical parts.
  2. Enable autoscale and scale to zero when the queue is empty.
  3. Pick the right VM size by measuring CPU and memory use, and pack more tasks per node with task slots.
  4. Reduce startup time so nodes spend less billed time not working.
  5. Delete pools and old jobs you no longer need, and use storage lifecycle rules for outputs.
Take quiz
What should an autoscale pool do when the task queue is empty?
Keep the maximum nodes running
Scale down, ideally to zero nodes
Double the node count
Delete the Batch account
Which VM option gives a discount at the risk of eviction?
Dedicated VMs
Reimaged VMs
Single-instance VMs
Spot VMs

36. How do you monitor Azure Batch workloads?

Use a mix of metrics, logs and the task data Batch already stores.

  • Azure Monitor metrics for node counts by state, task counts and pool resize events, with alerts on them.
  • Diagnostic logs sent to Log Analytics for pool, task and service events.
  • Application Insights from inside tasks for application-level telemetry.
  • Batch Explorer for a heat map of node states and quick access to task logs.

A useful alert is on nodes stuck in startTaskFailed or unusable.

Take quiz
Which tool gives a heat map of node states?
Azure Advisor
Azure DevOps Boards
Batch Explorer
Azure Key Vault
Where would you send Batch diagnostic logs for querying?
A Log Analytics workspace
The task's command line
A pool metadata field
A job schedule

37. Why doesn't an Azure Batch command line expand environment variables?

Batch runs the command line directly, not inside a shell. Operators such as $VAR, %VAR%, pipes and redirection are therefore not interpreted.

To use them, invoke a shell explicitly:

/bin/bash -c "echo $AZ_BATCH_TASK_WORKING_DIR > out.txt"
cmd /c "echo %AZ_BATCH_TASK_WORKING_DIR% > out.txt"

The same applies to chaining commands with &&. If a Linux task works in your terminal but fails in Batch with a missing file or odd argument, this is a likely cause.

Take quiz
How do you make a Batch command line expand $VAR on Linux?
Add a second dollar sign
Set expandVars to true on the pool
Use a job schedule
Wrap it in /bin/bash -c
Why are pipes and redirects ignored in a plain Batch command?
Batch blocks all file output
The command is not run inside a shell
Pipes need a premium pool
Spot nodes disable redirection

38. What are the key environment variables available to Azure Batch tasks?

Batch sets a number of AZ_BATCH_* variables on every task.

Variable Meaning
AZ_BATCH_TASK_WORKING_DIR Task's working directory
AZ_BATCH_TASK_DIR Task directory holding stdout.txt and stderr.txt
AZ_BATCH_NODE_SHARED_DIR Directory shared by all tasks on the node
AZ_BATCH_JOB_ID / AZ_BATCH_TASK_ID IDs of the current job and task
AZ_BATCH_POOL_ID / AZ_BATCH_NODE_ID IDs of the pool and node

Use the shared directory to cache data that many tasks on the same node can reuse.

Take quiz
Which variable points to a directory shared by tasks on the same node?
AZ_BATCH_NODE_SHARED_DIR
AZ_BATCH_TASK_ID
AZ_BATCH_POOL_ID
AZ_BATCH_JOB_ID
Which directory holds a task's stdout.txt?
The pool's start task folder
The application package folder
The directory in AZ_BATCH_TASK_DIR
The Key Vault mount

39. What is the difference between Azure Batch and Azure Kubernetes Service for batch workloads?

Both can run containers at scale, but they differ in what they manage for you.

Aspect Azure Batch AKS
Model Jobs, tasks and pools managed by the service Kubernetes cluster you operate
Scheduling Built-in task queue, priorities, dependencies Kubernetes Jobs, CronJobs, or add-ons
Operations No cluster to upgrade You manage node pools and upgrades
Fits when Finite batch and HPC work You already run on Kubernetes or need long-running services

If the work is a finite queue of jobs, Batch is usually simpler. If the team already standardises on Kubernetes, AKS keeps one platform.

Take quiz
Which service has a built-in task queue with priorities and dependencies?
Azure DNS
Azure Batch
Azure Front Door
Azure Bastion
When does AKS make more sense than Batch?
When you have no containers
When you need no scheduler
When you want zero operations
When you already run workloads on Kubernetes

40. How do you run Azure Batch from Azure Data Factory?

Data Factory orchestrates; Batch does the compute. Use the Custom activity in a pipeline, which submits a task to a Batch pool through a linked service.

  1. Create a Batch linked service with the account, pool name and credentials.
  2. Create a storage linked service where the activity script and its folder live.
  3. Add the Custom activity with the command to run and optional reference objects.

Data Factory waits for the task to finish and picks up its exit code, so pipeline dependencies and retries still work.

Take quiz
Which ADF activity runs code on a Batch pool?
Copy activity
Lookup activity
Custom activity
Wait activity
What does Data Factory use to know the Batch task result?
The task's exit code
The pool's name
The number of nodes
The job's priority

41. Explain the lifecycle of a job in Azure Batch?

A job moves through a clear set of states, driven by your calls and by task completion.

stateDiagram-v2
  [*] --> active: create job
  active --> disabling: disable
  disabling --> disabled
  disabled --> enabling: enable
  enabling --> active
  active --> terminating: terminate or all tasks done
  terminating --> completed
  completed --> deleting: delete
  deleting --> [*]

While active the job accepts tasks and they are scheduled. Disabled stops new scheduling, and running tasks can be requeued, terminated or allowed to finish depending on the option chosen.

A job moves to completed only when terminated, or automatically when onAllTasksComplete is terminateJob. Deleting a job removes its tasks and their node data.

Take quiz
Which job state accepts new tasks and schedules them?
disabled
completed
deleting
active
What makes a job complete automatically when its tasks finish?
Setting taskSlotsPerNode to 1
Setting onAllTasksComplete to terminateJob
Disabling autoscale
Adding a start task

42. Explain the execution flow of a task on a compute node?

From the moment the scheduler assigns a task until its results are stored, the flow is as follows.

sequenceDiagram
  participant S as Batch Service
  participant A as Node Agent
  participant ST as Azure Storage
  S->>A: Assign task to node
  A->>ST: Download resource files
  A->>A: Run job preparation task if needed
  A->>A: Execute task command line
  A->>ST: Upload output files
  A->>S: Report exit code and state

If a resource file fails to download, the task fails before its command runs, and the failure is recorded in failureInfo. The task directory holding stdout and stderr stays on the node until the retention time passes or the task or job is deleted.

Take quiz
Which step happens before the task's command line runs?
Resource files are downloaded
Output files are uploaded
The pool is deleted
The job is terminated
Who reports the final exit code back to the Batch service?
The storage account
The virtual network
The node agent
The job schedule

43. Explain the internal working of the Azure Batch service architecture?

Batch is a control plane plus agents. Your client calls the Batch REST API, and the service keeps the state of accounts, pools, jobs and tasks.

flowchart LR
  C["Client or SDK"] --> API["Batch service API"]
  API --> SCH["Scheduler and pool manager"]
  SCH --> VM["Compute VMs from Azure"]
  VM --> AG["Node agent on each node"]
  AG --> API
  AG <--> ST["Azure Storage"]

The pool manager asks Azure compute for VMs matching the pool definition. Each VM boots an image that includes the node agent, which registers with the service, polls for assigned tasks, runs them and reports status.

The scheduler matches active tasks to free slots by job priority and task order. Because the agent initiates the communication, the service can track nodes without you managing any cluster software.

Take quiz
What runs on each node to receive and execute tasks?
A Kubernetes kubelet
The node agent
An Azure Function host
A SQL agent
What does the pool manager request from Azure compute?
Storage accounts for each task
A new Batch account per job
DNS zones for nodes
VMs matching the pool definition

44. How can you optimize task throughput in Azure Batch?

Throughput depends on how quickly nodes get busy and stay busy.

  • Right-size tasks: very short tasks waste time on scheduling and file downloads, so group tiny work items into one task.
  • Use task slots to match concurrency to CPU and memory.
  • Add tasks in bulk with the collection API rather than one call per task.
  • Cache shared inputs in the node shared directory or a start task instead of downloading per task.
  • Avoid long dependency chains that leave nodes idle.

Check the node state heat map to see whether nodes are idle or saturated before tuning.

Take quiz
Why group very short work items into one task?
To avoid using the Batch API
To remove the need for a pool
To reduce per-task scheduling and download overhead
To prevent exit codes
What is the efficient way to add thousands of tasks?
Use the add task collection operation
Add each task from a separate pool
Create one job per task
Use only the portal

45. How do you reduce node startup time in Azure Batch?

Startup time is the gap between a node being requested and it running your first task, and you pay for it.

  1. Bake software into a custom image in Azure Compute Gallery instead of installing it in the start task.
  2. Keep the start task short and avoid large downloads.
  3. Pre-fetch container images in the pool's container configuration.
  4. Use application packages of modest size, as they are downloaded to every new node.
  5. Keep a small warm pool if workloads arrive in bursts.

Measure startup from the node's creation time to its first task start to see which step dominates.

Take quiz
Which approach removes software installation time from node startup?
A larger job priority
More retry attempts
A longer task constraint
A custom image with the software pre-installed
Where are custom images for Batch commonly stored?
Azure DNS
Azure Compute Gallery
Azure Monitor
Azure Policy

46. How do you handle quotas in Azure Batch?

Quotas cap how many cores you can use, and they are the most common reason a pool fails to scale.

In Batch service mode, quotas are per account and per VM family, with separate limits for dedicated and Spot cores, plus limits on pools and jobs. In user subscription mode, the subscription's regular VM quotas apply.

  1. Check the quota page of the account before sizing pools.
  2. Request increases through a support request well ahead of big runs.
  3. Split workloads across VM families or regions if a family is exhausted.

A quota failure appears in the pool's resize errors.

Take quiz
In Batch service mode, quotas are tracked at which level?
Per Batch account, by VM family
Per task
Per storage blob
Per virtual network
Where does a quota problem show up for a pool?
In the job schedule
In the task's stdout only
In the pool's resize errors
In the application package

47. What happens when all tasks in an Azure Batch job complete?

By default, nothing. The job stays active and ready to accept more tasks, and it keeps showing in the job list.

To end it automatically, set onAllTasksComplete to terminateJob when you create the job, or update it afterwards. Be careful with this setting when you add tasks over time: if the job is momentarily empty or all current tasks are done, it will terminate before the next tasks arrive.

Task data stays on nodes until its retention time (7 days by default) passes, so save outputs to Storage first.

Take quiz
What is the default behaviour when every task in a job finishes?
The job is deleted
The job remains active
The pool is resized to zero
The job is cloned
What risk comes with onAllTasksComplete set to terminateJob?
Tasks run twice
Nodes are reimaged
Output files are skipped
The job can end before you add later tasks

48. How do you update a running Azure Batch pool?

You can change some settings in place, but others need a new pool.

Change in place Needs a new pool
Target node counts or autoscale formula VM size
Start task, application package references, metadata OS image
Resize, reboot or reimage of nodes Task slots per node and virtual network settings

Updates to the start task or packages apply to nodes when they next reboot or when new nodes join, so existing nodes aren't changed immediately. To roll out a change fully, reboot or reimage the nodes, or create a new pool and move jobs to it.

A new pool with the same pattern lets you switch jobs over without disrupting running work.

Take quiz
When does an updated start task reach existing nodes?
Immediately, without a reboot
Only after the pool is deleted
When they reboot or new nodes join
Never, it is read-only
Which change requires creating a new pool?
Changing the VM size
Changing the target node count
Changing pool metadata
Updating the autoscale formula

49. What Azure Batch features have been retired?

Several older features are gone, and you shouldn't use them in new designs.

Retired feature Retired on Replacement
Batch certificates 29 February 2024 Azure Key Vault and managed identities
Cloud Services configuration pools 29 February 2024 Virtual Machine configuration pools
Low-priority nodes 30 September 2025 Spot VMs

If you inherit older code, replace certificate-based secrets with Key Vault, and check that pools use Virtual Machine configuration and Spot nodes. Always confirm current status on Microsoft's retirement notices.

Take quiz
What replaced low-priority nodes in Batch?
Dedicated-only nodes
Reserved pools
Cloud Services roles
Spot VMs
What is the recommended alternative to Batch certificates?
Storing secrets in task IDs
Azure Key Vault with managed identities
Public blob containers
Job metadata strings

50. Which is better for HPC on Azure, Batch or CycleCloud, and why?

Neither is better in general. They answer different needs.

Aspect Azure Batch Azure CycleCloud
What it is Managed scheduler as a service Tool to build and manage HPC clusters
Scheduler Batch's own job and task model Your choice, such as Slurm, PBS or LSF
Effort No cluster to administer You manage the cluster software
Choose when Cloud-native or new parallel workloads Existing scripts built for a specific scheduler

Pick Batch for new, cloud-native parallel work with minimal operations. Pick CycleCloud when you need to lift an existing Slurm or PBS environment with its scripts and users unchanged.

Take quiz
Which service lets you run Slurm or PBS clusters on Azure?
Azure CycleCloud
Azure Batch's job schedule
Azure Functions
Azure Event Hubs
Why choose Batch over CycleCloud for a new project?
It supports Slurm natively
It needs no VMs
There is no cluster software to administer
It runs only on-premises
«
»

Comments & Discussions