Java / Azure Kubernetes Service (AKS) Interview questions
Last updated
1. What is Azure Kubernetes Service (AKS)?
Azure Kubernetes Service (AKS) is Microsoft's managed Kubernetes offering. Azure runs the control plane for you (API server, etcd, scheduler and controller manager), and you only manage the worker nodes that run your containers.
Day-to-day chores such as provisioning the cluster, patching, version upgrades and health monitoring of the control plane are handled by the platform. You keep control over node sizes, networking, scaling rules and the workloads you deploy.
AKS plugs into the rest of Azure: Microsoft Entra ID for sign-in, Azure Container Registry for images, Azure Monitor for telemetry, and virtual networks for private connectivity.
Take quiz
You operate it on VMs in your subscription
A pod in the kube-system namespace operates it
Azure operates it for you
Your container registry operates it
The API server and etcd by the hour
Only the Azure Container Registry
Only the pods that are running
The worker node VMs and attached resources
2. What are the main components of an AKS cluster?
An AKS cluster has two halves: a Microsoft-managed control plane and customer-owned nodes.
| Part | Component | Role |
| Control plane | kube-apiserver | Front door for every kubectl and controller request |
| Control plane | etcd | Key-value store holding all cluster state |
| Control plane | kube-scheduler | Picks a node for each new pod |
| Control plane | kube-controller-manager | Runs loops such as the ReplicaSet and node controllers |
| Node | kubelet | Agent that starts and watches pods on the node |
| Node | containerd | Container runtime that pulls images and runs containers |
| Node | kube-proxy | Programs service routing rules (unless Cilium replaces it) |
The nodes are Azure VMs grouped into node pools. AKS places them, along with load balancers and disks it creates, in a separate node resource group whose name starts with MC_.
Take quiz
kubelet
kube-proxy
containerd
etcd
The VM scale sets, load balancers and disks AKS creates for the cluster
The control plane virtual machines
Backup copies of etcd
Only your container registry
3. What are node pools in AKS?
A node pool is a group of worker nodes that share the same VM size, OS type and settings. On AKS each pool is backed by a Virtual Machine Scale Set, so scaling a pool just changes the scale set instance count.
Pools let you mix hardware in one cluster. You might keep a general-purpose pool for web pods, a GPU pool for model inference and a Windows pool for a legacy .NET service, then steer pods to the right pool with labels and taints.
az aks nodepool add \ --resource-group rg-demo --cluster-name aks-demo \ --name gpupool --node-count 2 \ --node-vm-size Standard_NC6s_v3 \ --node-taints sku=gpu:NoSchedule
Take quiz
A Virtual Machine Scale Set
An Azure App Service plan
A single Azure VM shared by all pods
An Azure Batch pool
To get a second control plane
To run workloads that need a different VM size, such as GPUs
To avoid paying for node VMs
To store container images closer to pods
4. What is the difference between system and user node pools?
System node pools host AKS's critical pods such as CoreDNS, metrics-server and the konnectivity agent. User node pools are where your own application pods run.
| Aspect | System pool | User pool |
| Purpose | Cluster-critical add-on pods | Application workloads |
| OS | Linux only | Linux or Windows |
| Minimum count | At least 1 pool, 2+ nodes for production | Can scale down to 0 |
| Typical taint | CriticalAddonsOnly=true:NoSchedule |
None by default |
Every cluster needs at least one system pool. Tainting it keeps application pods off those nodes, so a noisy app cannot starve CoreDNS.
Take quiz
The system node pool
A user node pool
Both, with no restriction
Neither pool type
kubernetes.io/os=windows:NoExecute
node.kubernetes.io/unreachable:NoSchedule
CriticalAddonsOnly=true:NoSchedule
sku=gpu:PreferNoSchedule
5. What are the AKS pricing tiers?
AKS charges nothing for the control plane on the Free tier, but there is no financially backed SLA, so it suits dev and test. You still pay for the nodes.
| Tier | What you get | Best for |
| Free | Managed control plane, no uptime SLA, lower scale limits | Learning, dev/test |
| Standard | Financially backed uptime SLA (higher with availability zones), up to 5,000 nodes | Production |
| Premium | Everything in Standard plus long-term support (LTS) for a Kubernetes version | Regulated or slow-upgrading estates |
Standard and Premium add a per-cluster hourly fee on top of node costs. Check the Azure pricing page for current rates.
Take quiz
Free
Standard
Premium
All tiers include it
Free
Developer
Basic
Standard
6. What is a pod in AKS?
A pod is the smallest deployable unit in Kubernetes. It wraps one or more containers that share a network namespace (one IP, talking over localhost) and can share mounted volumes.
Pods are disposable. If a node dies, the pod is not moved, a controller such as a Deployment creates a replacement. That is why production workloads should not be run as bare pods.
apiVersion: v1 kind: Pod metadata: name: web spec: containers: - name: nginx image: nginx:1.27 ports: - containerPort: 80
Take quiz
The node's root filesystem
A separate IP each
Nothing, they are fully isolated
The pod IP address and localhost networking
Nothing recreates them if the node fails
They cannot use environment variables
They cannot pull images from ACR
They always run as root
7. What are the networking models available in AKS?
AKS supports several CNI choices that differ mainly in where pod IPs come from.
| Model | Pod IP source | Notes |
| kubenet | Private range, NATed behind the node | Legacy; retiring on 31 March 2028 |
| Azure CNI Node Subnet | Same VNet subnet as nodes | Pods are directly routable; burns many IPs |
| Azure CNI Pod Subnet | Separate VNet subnet for pods | Routable pods with better IP planning |
| Azure CNI Overlay | Private overlay CIDR | Saves VNet space; scales to large clusters |
The Cilium dataplane can be layered on Overlay or Pod Subnet to get eBPF networking and network policy. For new clusters, Overlay is the usual starting point.
Take quiz
kubenet
Azure CNI Overlay
Azure CNI Pod Subnet
Azure CNI powered by Cilium
kubenet
Azure CNI Pod Subnet
Azure CNI Overlay
ExternalName networking
8. How do you create an AKS cluster using Azure CLI?
You need a resource group first, then a single az aks create call.
- Sign in and pick the subscription with
az login. - Create a resource group.
- Create the cluster with your node count, VM size and network plugin.
- Pull credentials and verify with kubectl.
az group create -n rg-demo -l eastus az aks create -g rg-demo -n aks-demo \ --node-count 3 --node-vm-size Standard_D4s_v5 \ --network-plugin azure --network-plugin-mode overlay \ --generate-ssh-keys az aks get-credentials -g rg-demo -n aks-demo kubectl get nodes
Provisioning usually takes a few minutes. Add --tier standard for a production SLA.
Take quiz
az kubernetes deploy
az aks create
az container create
az aks init
--replicas
--pod-count
--node-count
--scale-size
9. How do you connect to an AKS cluster using kubectl?
Run az aks get-credentials. It downloads the cluster's connection details and merges them into ~/.kube/config as a new context.
az aks get-credentials -g rg-demo -n aks-demo kubectl config current-context kubectl get nodes
If the cluster uses Microsoft Entra ID, kubectl needs kubelogin to fetch tokens (az aks install-cli installs both). The --admin flag fetches a local admin certificate, but it only works while local accounts are enabled and is best avoided in production.
Take quiz
The Azure Key Vault
/etc/kubernetes on the node
The kubeconfig file, ~/.kube/config
The cluster's etcd database
kubeadm
helm
kustomize
kubelogin
10. What is Azure Container Registry and how does AKS use it?
Azure Container Registry (ACR) is Azure's private, managed registry for container images and Helm charts. AKS pulls application images from it when pods start.
The cleanest integration is to attach the registry to the cluster. This grants the AcrPull role to the cluster's kubelet managed identity, so you do not need image pull secrets.
az aks update -g rg-demo -n aks-demo --attach-acr myregistry # reference the image in your manifest # image: myregistry.azurecr.io/shop/api:1.4.2
Take quiz
Contributor
AcrPush
Owner
AcrPull
The kubelet managed identity
The developer's user account
The ingress controller
The pod's service account
11. What are the types of Kubernetes Services?
A Service gives a stable virtual IP and DNS name to a changing set of pods. The type decides who can reach it.
| Type | Reachable from | On AKS |
| ClusterIP | Inside the cluster only (default) | Internal virtual IP |
| NodePort | Node IP on ports 30000-32767 | Opens that port on every node |
| LoadBalancer | Outside the cluster | Creates an Azure Load Balancer rule with a public IP |
| ExternalName | Inside the cluster | Returns a CNAME to an external DNS name |
For a private load balancer, add the annotation service.beta.kubernetes.io/azure-load-balancer-internal: "true" to the Service.
Take quiz
ClusterIP
NodePort
LoadBalancer
ExternalName
A new virtual network
An Azure Load Balancer rule and frontend IP
An Application Gateway WAF policy
A Traffic Manager profile
12. What is an Ingress in AKS?
An Ingress is a Kubernetes object that routes external HTTP and HTTPS traffic to Services based on host name and URL path, and can terminate TLS. It works at layer 7, so one public IP can front many apps.
The object does nothing on its own. An ingress controller must be running to read the rules and configure a proxy. On AKS that is commonly the application routing add-on, Application Gateway for Containers or a self-managed controller.
The community Ingress-NGINX project has been retired and the Ingress API is frozen, so newer designs lean towards Gateway API.
Take quiz
Layer 2 (Ethernet)
Layer 7 (HTTP/HTTPS)
Layer 3 only (IP)
Layer 4 only (TCP)
A NodePort on every service
A second node pool
An ingress controller
A private DNS zone
13. What is a managed identity in AKS?
A managed identity is an Azure-managed credential that AKS uses to call Azure APIs, so you never store or rotate a service principal secret.
A cluster uses two of them. The control plane identity lets AKS create and manage resources such as load balancers, routes and disks. The kubelet identity is used by nodes, mainly to pull images from ACR.
You can let AKS create them (system-assigned for the control plane) or bring your own user-assigned identities, which is handy when you must pre-grant roles before the cluster exists.
Take quiz
They make pods start faster
They remove the need for RBAC
No client secret to store or rotate
They replace the container runtime
The control plane identity
The ingress identity
The Entra admin group
The kubelet identity
14. What are ConfigMaps and Secrets in Kubernetes?
A ConfigMap holds non-sensitive settings such as feature flags or a log level. A Secret holds sensitive values such as passwords and tokens. Both can reach a pod as environment variables or as files in a mounted volume.
Secret values are only base64-encoded, which is encoding, not encryption. Anyone with read access to the Secret can decode it. AKS encrypts etcd at rest, and you can add KMS encryption with Key Vault, but tight RBAC matters just as much.
kubectl create configmap app-config --from-literal=LOG_LEVEL=info kubectl create secret generic db-pass --from-literal=password='S3cret!'
Take quiz
A Secret
A PersistentVolume
A ServiceAccount
A ConfigMap
Encoding only, not protection from readers
Strong AES encryption
Automatic rotation
Hiding the value from kubectl
15. What is the cluster autoscaler in AKS?
The cluster autoscaler adds nodes to a node pool when pods cannot be scheduled for lack of CPU or memory, and removes nodes that stay underused.
It is enabled per node pool with a minimum and maximum count. By default a node becomes a scale-down candidate after it has been unneeded for about 10 minutes and its utilization is below 50%.
az aks nodepool update -g rg-demo --cluster-name aks-demo -n userpool \ --enable-cluster-autoscaler --min-count 2 --max-count 8
Take quiz
Pods that cannot be scheduled due to insufficient resources
High CPU on the control plane
A new Docker image push
A failed liveness probe
With --replicas and --pods
With --min-count and --max-count
With a ResourceQuota
With --node-taints
16. What is the Horizontal Pod Autoscaler?
The Horizontal Pod Autoscaler (HPA) changes the replica count of a Deployment based on observed metrics, usually CPU or memory from metrics-server, which AKS ships by default.
The controller computes desired = ceil(current replicas x current metric / target metric). For CPU-based scaling, pods must declare resource requests, otherwise utilization cannot be calculated.
kubectl autoscale deployment web --cpu-percent=70 --min=3 --max=15
Take quiz
The VM size of nodes
The number of pod replicas
The container image tag
The Kubernetes version
A NodePort
An init container
Resource requests
An ingress annotation
17. How do you scale an AKS cluster manually?
Scale nodes with az aks scale (or az aks nodepool scale for one pool) and scale pods with kubectl scale.
# nodes az aks nodepool scale -g rg-demo --cluster-name aks-demo -n userpool --node-count 5 # pods kubectl scale deployment web --replicas=10
If the cluster autoscaler is enabled on that pool, you cannot set a fixed count. Change the min and max values, or disable the autoscaler first.
Take quiz
kubectl resize nodes
az aks pods scale
az aks nodepool scale
az vmss reimage
Pools can never be resized
The control plane is locked
kubectl must be used instead
The autoscaler owns the node count
18. What are persistent volumes in AKS?
A PersistentVolume (PV) is storage that outlives any single pod. A pod asks for it through a PersistentVolumeClaim (PVC), and a StorageClass tells Kubernetes how to provision it dynamically.
AKS ships CSI-based classes such as managed-csi (Azure Disk) and azurefile-csi (Azure Files). When a PVC is created, the driver provisions the Azure resource and binds it automatically.
apiVersion: v1 kind: PersistentVolumeClaim metadata: name: data-pvc spec: accessModes: [ReadWriteOnce] storageClassName: managed-csi resources: requests: storage: 20Gi
Take quiz
A ConfigMap
A NodePort
A ServiceAccount
A PersistentVolumeClaim
managed-csi
azurefile-csi
local-path
ephemeral-os
19. What is Container Insights in AKS?
Container Insights is the Azure Monitor feature that collects container logs, node and pod performance data and Kubernetes events, and sends them to a Log Analytics workspace.
Enabling the monitoring add-on deploys the ama-logs agent as a DaemonSet. You then query data with KQL or use the built-in workbooks in the portal.
az aks enable-addons -g rg-demo -n aks-demo -a monitoring
For metrics, Azure Monitor managed service for Prometheus and managed Grafana are the usual companions.
Take quiz
In a Log Analytics workspace
In the cluster's etcd
In Azure Container Registry
On each node's local disk only
As a StatefulSet
As a DaemonSet
As a CronJob
As a static VM extension only
20. What is AKS Automatic?
AKS Automatic is an opinionated cluster SKU where Azure manages far more than the control plane, including node provisioning, scaling, upgrades and many security defaults.
It comes pre-configured with node auto-provisioning (Karpenter-based), Azure CNI Overlay with Cilium, managed Prometheus and Grafana, Entra ID with Azure RBAC, workload identity and deployment safeguards. You focus on deploying apps.
az aks create -g rg-demo -n aks-auto --sku automatic
The trade-off is less freedom to tweak low-level settings than on AKS Standard.
Take quiz
--mode auto-pilot
--sku automatic
--tier premium
--managed-nodes true
You hand-pick each VM size
Only fixed-size pools are allowed
They are provisioned automatically based on pod needs
Nodes are never patched
21. How do you upgrade an AKS cluster?
AKS upgrades the control plane first, then the node pools. You pick a supported target and AKS performs a rolling upgrade.
- List what is available with
az aks get-upgrades. - Test the new version in a lower environment and check for removed APIs.
- Run
az aks upgrade(use--control-plane-onlyto stage the nodes later). - Confirm node versions with
kubectl get nodes.
az aks get-upgrades -g rg-demo -n aks-demo -o table az aks upgrade -g rg-demo -n aks-demo --kubernetes-version 1.33.2
You cannot skip minor versions, so 1.31 goes to 1.32 before 1.33. To automate this, set an auto-upgrade channel such as patch, stable or rapid and define a planned maintenance window.
Take quiz
Yes, go directly to the latest
Only on the Free tier
No, minor versions must be upgraded one at a time
Only for node pools
Upgrades only system pods
Skips the API server
Reboots every node first
Upgrades only the control plane, leaving nodes for later
22. What is the difference between kubenet and Azure CNI?
With kubenet nodes get VNet IPs while pods get addresses from a separate private range and are NATed through the node. With Azure CNI pods receive routable VNet IPs directly.
| Aspect | kubenet | Azure CNI (node subnet) |
| Pod IP | Private range, NATed | VNet IP |
| VNet IP use | One per node | One per node plus one per pod |
| Default max pods/node | 110 | 30 (up to 250) |
| Network policy | Calico only | Azure NPM, Calico or Cilium |
| Virtual nodes, Windows pools | Not supported | Supported |
| Future | Retires 31 March 2028 | Supported |
Plan to move kubenet clusters to Azure CNI Overlay, which keeps the low IP footprint without kubenet's limits.
Take quiz
A dedicated VNet IP each
From Azure DHCP per pod
From the control plane subnet
From a private range, NATed through the node
Windows node pools
Linux node pools
Services of type ClusterIP
Persistent volume claims
23. When would you choose Azure CNI Overlay?
Choose Azure CNI Overlay when VNet address space is scarce or you expect a large cluster. Pods draw IPs from a private overlay CIDR (default 10.244.0.0/16), so only nodes consume VNet addresses.
It supports up to 250 pods per node and clusters of up to 5,000 nodes, and the same pod CIDR can be reused across clusters because it is never advertised on the VNet.
The catch is that pods are not directly routable from outside the cluster. External systems must reach workloads through Services, ingress or load balancers. If something must address pod IPs directly, pick Pod Subnet instead.
az aks create -g rg-demo -n aks-demo \ --network-plugin azure --network-plugin-mode overlay \ --pod-cidr 192.168.0.0/16
Take quiz
Pods do not consume VNet IP addresses
Pods become directly routable from on-premises
It removes the need for node pools
It disables network policy
It only runs Windows pods
Pod IPs are not directly reachable from outside the cluster
It caps clusters at 10 nodes
It forbids load balancers
24. How does AKS integrate with Microsoft Entra ID?
With AKS-managed Entra integration, users sign in with their Entra identity and the API server validates the token. You nominate Entra groups as cluster admins, and access is then decided by Kubernetes RBAC or Azure RBAC.
sequenceDiagram participant U as User participant E as Microsoft Entra ID participant A as AKS API server U->>E: az login / kubelogin E-->>U: Access token U->>A: kubectl request + token A->>E: Validate token A-->>U: Allow or deny via RBAC
az aks create -g rg-demo -n aks-demo \ --enable-aad --aad-admin-group-object-ids <group-id> \ --disable-local-accounts
Disabling local accounts removes shared admin certificates, so every action is tied to a named identity.
Take quiz
Stores it in etcd forever
Validates it against Microsoft Entra ID
Forwards it to the container registry
Ignores it and uses a certificate
To speed up node pool scaling
To enable Windows containers
To force every request to use a named Entra identity
To remove the need for RBAC
25. What is the difference between Kubernetes RBAC and Azure RBAC in AKS?
Both answer the question "what can this user do?", but they store the rules in different places.
| Aspect | Kubernetes RBAC | Azure RBAC for Kubernetes |
| Rules live in | Role, ClusterRole and bindings inside the cluster | Azure role assignments in IAM |
| Managed with | kubectl and YAML | Azure portal, CLI, policy |
| Roles | Custom, any verbs you define | Built-in such as RBAC Reader, Writer, Admin, Cluster Admin |
| Scope | Namespace or cluster | Cluster, namespace or subscription hierarchy |
Azure RBAC for Kubernetes authorization is switched on with --enable-azure-rbac. Do not confuse it with the Azure roles that control the AKS resource itself (for example Contributor), which govern who can scale or delete the cluster, not who can read pods.
Take quiz
In Azure IAM only
In the container registry
Inside the cluster as Role and RoleBinding objects
On each node's disk
--enable-rbac-v2
--azure-roles-only
--enable-entra-roles
--enable-azure-rbac
26. How does Workload Identity work in AKS?
Workload Identity lets a pod call Azure services (Key Vault, Storage, SQL) as a managed identity without any secret. It uses OIDC federation between the cluster and Microsoft Entra ID.
flowchart LR A["Pod with annotated service account"] --> B["Projected service account token"] B --> C["Microsoft Entra ID checks federated credential"] C --> D["Entra issues Azure access token"] D --> E["Pod calls Key Vault or Storage"]
- Enable
--enable-oidc-issuer --enable-workload-identityon the cluster. - Create a user-assigned managed identity and grant it the needed Azure role.
- Create a federated credential tied to
system:serviceaccount:<ns>:<sa>. - Annotate the service account with the identity's client ID and label the pod
azure.workload.identity/use: "true".
It replaces the retired AAD Pod Identity.
Take quiz
A shared client secret in a ConfigMap
A node-level SSH key
An ACR token
A federated identity credential
azure.workload.identity/use: "true"
aks.azure.com/identity: on
workload.pod.io/enabled: true
kubernetes.io/entra: enabled
27. How do you use Azure Key Vault secrets in AKS?
Use the Secrets Store CSI Driver with the Azure Key Vault provider. It mounts secrets, keys or certificates from Key Vault into the pod as files, so values never sit in your manifests.
- Enable the add-on:
az aks enable-addons -a azure-keyvault-secrets-provider. - Give the pod an identity (Workload Identity is recommended) with Get permission on the vault.
- Define a
SecretProviderClassnaming the vault and objects. - Mount it as a CSI volume in the pod spec.
Secrets can optionally be synced to a native Kubernetes Secret if an app insists on environment variables. Enable rotation polling if you want updated vault values to appear in running pods.
Take quiz
Mounted as files in a volume
Copied into the container image
Baked into the node VM image
Sent by email to the developer
StorageClass
SecretProviderClass
IngressClass
PriorityClass
28. What is the application routing add-on in AKS?
The application routing add-on is a managed ingress solution. It deploys NGINX ingress controllers and wires up Azure DNS records and Key Vault-backed TLS certificates for you.
az aks approuting enable -g rg-demo -n aks-demo # Ingress uses this class # ingressClassName: webapprouting.kubernetes.azure.com
Upstream Ingress-NGINX reached end of life in March 2026. Microsoft committed to critical security patches for the add-on's NGINX implementation only through November 2026, so existing users should plan a move to the add-on's Gateway API implementation or another Gateway API option.
Take quiz
alb.azure.com
webapprouting.kubernetes.azure.com
istio.io/ingress
nginx.ingress.azure.net
March 2028
June 2030
November 2026
There is no committed date
29. Why is Gateway API replacing Ingress in AKS?
The Ingress API is frozen and relies on controller-specific annotations for anything beyond basic host and path routing. Gateway API is its successor and models traffic with separate, role-based resources.
| Aspect | Ingress | Gateway API |
| Resources | One Ingress object | GatewayClass, Gateway, HTTPRoute, GRPCRoute |
| Ownership | Single team edits everything | Platform team owns Gateway, app teams own routes |
| Advanced routing | Vendor annotations | Header matches, weighted traffic splitting in the spec |
| Portability | Varies by controller | Standard across conformant implementations |
On AKS the main options are Application Gateway for Containers, the Istio-based Gateway API mode in the application routing add-on (in preview when announced, so verify its status) and the Istio service mesh add-on. With Ingress-NGINX retired, new designs should start on Gateway API.
Take quiz
GatewayClass
StorageClass
HTTPRoute
NetworkPolicy
It cannot terminate TLS
It only supports UDP
It cannot run on AKS
Reliance on controller-specific annotations for advanced routing
30. When should you use Azure Disk versus Azure Files in AKS?
Use Azure Disk for single-writer, latency-sensitive data such as a database. Use Azure Files when several pods on different nodes must read and write the same files.
| Aspect | Azure Disk | Azure Files |
| Access mode | ReadWriteOnce (one node at a time) | ReadWriteMany |
| Protocol | Block device | SMB or NFS share |
| Performance | Low latency, high IOPS on Premium/Ultra | Higher latency, shared throughput |
| Zone behavior | Zonal unless ZRS disk is used | ZRS shares survive a zone failure |
| Typical use | Databases, queues, single-instance stateful apps | Shared uploads, CMS content, config sharing |
Also keep the per-VM disk attach limit in mind when packing many stateful pods on one node.
Take quiz
Azure Disk
Node local temp disk
Azure Disk with LRS
Azure Files
Azure Disk
Azure Files over SMB
An emptyDir volume
A ConfigMap volume
31. What is KEDA and when would you use it in AKS?
KEDA (Kubernetes Event-driven Autoscaling) scales workloads based on external event sources such as Service Bus queue length, Event Hubs lag, Kafka topics or a Prometheus query. AKS offers it as a managed add-on with --enable-keda.
Unlike plain HPA, KEDA can scale a Deployment down to zero when there is no work and back up when events arrive. For one-or-more replicas it feeds an HPA behind the scenes.
apiVersion: keda.sh/v1alpha1 kind: ScaledObject metadata: name: orders-worker spec: scaleTargetRef: name: orders-worker minReplicaCount: 0 maxReplicaCount: 30 triggers: - type: azure-servicebus metadata: queueName: orders messageCount: "20"
Take quiz
Scale to zero based on queue or event depth
Resize node VMs
Upgrade the control plane
Encrypt etcd
SecretProviderClass
ScaledObject
PodDisruptionBudget
NginxIngressController
32. How do you implement network policies in AKS?
By default every pod can talk to every other pod. A NetworkPolicy restricts that, but it only works if the cluster was created with a policy engine: azure, calico or cilium set through --network-policy.
A common pattern is a default-deny rule per namespace, then explicit allows. Policies are additive, so any matching allow rule opens the path.
apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: allow-web-to-api namespace: shop spec: podSelector: matchLabels: {app: api} policyTypes: [Ingress] ingress: - from: - podSelector: matchLabels: {app: web} ports: - port: 8080
For new clusters Cilium (eBPF) is generally the direction Microsoft favors, but check current docs before choosing.
Take quiz
All traffic is denied
All pod-to-pod traffic is allowed
Only same-namespace traffic is allowed
Only HTTPS is allowed
--pod-firewall
--policy-engine
--network-policy
--enable-acl
33. How do you enforce policies on AKS with Azure Policy?
Enable the Azure Policy add-on. It installs Gatekeeper (OPA) admission webhooks and syncs Azure Policy definitions into the cluster as constraints.
az aks enable-addons -g rg-demo -n aks-demo -a azure-policy
You then assign built-in definitions or initiatives at subscription, resource group or cluster scope. Typical rules block privileged containers, require resource limits, allow images only from approved registries or forbid hostPath mounts.
Each policy runs in Audit mode (report only) or Deny mode (reject at admission). Starting in Audit lets you see what would break before enforcing. Compliance shows up in the Azure Policy dashboard.
Take quiz
Calico
Flux
Gatekeeper (OPA)
CoreDNS
Deny
Disabled-by-admission
Mutate-all
Audit
34. How do you create a private AKS cluster?
A private cluster exposes the API server only on a private IP inside your VNet, through Private Link, so there is no public endpoint.
az aks create -g rg-demo -n aks-private \ --enable-private-cluster \ --private-dns-zone system \ --network-plugin azure --network-plugin-mode overlay
Because the endpoint is private, kubectl must run from somewhere with a network path: a VM in the VNet or a peered one, a VPN or ExpressRoute link, or a self-hosted CI agent. For quick one-off commands, az aks command invoke runs kubectl through the Azure control plane.
An alternative is API Server VNet Integration, which places the API server behind an internal load balancer in a delegated subnet and avoids Private Link.
Take quiz
On a public IP with a firewall rule
Through a NodePort on each node
Only through the Azure portal
On a private IP inside the VNet via Private Link
az aks command invoke
az aks proxy
kubectl --public
az network bastion
35. What is the difference between node image upgrades and Kubernetes version upgrades?
A Kubernetes version upgrade moves the control plane and kubelet to a new Kubernetes release. A node image upgrade keeps the Kubernetes version but replaces the node OS image with a newer one containing OS and runtime security patches.
| Aspect | Kubernetes upgrade | Node image upgrade |
| Changes | Kubernetes minor or patch version | OS, containerd and kernel patches |
| Cadence | Few times a year | New images roughly weekly |
| Command | az aks upgrade |
az aks nodepool upgrade --node-image-only |
| API impact | Possible API deprecations | None |
Many teams automate node images with the NodeImage or SecurityPatch auto-upgrade channel and reserve Kubernetes upgrades for planned windows.
Take quiz
A node image upgrade
A minor version upgrade
A control-plane-only upgrade
A CNI migration
--os-only-reboot
--node-image-only
--no-kubelet
--patch-os
36. How do you use spot node pools in AKS?
A spot node pool runs on Azure Spot VMs, which are much cheaper but can be evicted with about 30 seconds' notice when Azure needs the capacity. Use them for interruption-tolerant work such as batch jobs, CI runners and stateless workers.
az aks nodepool add -g rg-demo --cluster-name aks-demo -n spotpool \ --priority Spot --eviction-policy Delete --spot-max-price -1 \ --enable-cluster-autoscaler --min-count 0 --max-count 10
AKS automatically taints these nodes with kubernetes.azure.com/scalesetpriority=spot:NoSchedule, so pods must carry a matching toleration. Spot pools cannot be the system pool, and critical or stateful services should stay on regular nodes.
Take quiz
A user node pool
The system node pool
A GPU pool
A Windows pool
They request a NodePort
They set a ConfigMap flag
They tolerate the spot taint AKS applies
They run as privileged
37. What are taints and tolerations in AKS?
A taint on a node repels pods that do not tolerate it. A toleration on a pod allows, but does not force, scheduling onto a matching tainted node.
Effects are NoSchedule (do not place new pods), PreferNoSchedule (try to avoid) and NoExecute (also evict running pods).
# taint at pool creation az aks nodepool add ... --node-taints sku=gpu:NoSchedule # pod spec tolerations: - key: sku operator: Equal value: gpu effect: NoSchedule
Because a toleration only permits placement, pair it with a nodeSelector or node affinity if the pod must land on that pool.
Take quiz
Forces the pod onto that node
Deletes the taint
Permits a pod to schedule on a tainted node
Adds a second container
NoSchedule
PreferNoSchedule
AllowOnly
NoExecute
38. Which is better: AKS or Azure Container Apps?
Neither wins outright, it depends on how much Kubernetes you actually need. Azure Container Apps (ACA) is a serverless platform built on Kubernetes, KEDA and Envoy that hides the Kubernetes API. AKS hands you the full API.
| Aspect | AKS | Azure Container Apps |
| Kubernetes API access | Full (CRDs, operators, DaemonSets) | None |
| Operations burden | Higher: nodes, upgrades, networking | Low: no nodes to manage |
| Scale to zero | With KEDA | Built in |
| Customization | Very high | Limited to platform features |
| Best for | Platform teams, complex or multi-tenant estates | Microservices, APIs, event-driven apps, jobs |
Pick ACA when you want containers running fast without cluster operations. Pick AKS when you need operators, custom networking, specific schedulers or portability of existing Kubernetes tooling.
Take quiz
Azure Container Apps
Azure Container Instances
Azure Functions
AKS
When you want to run containers without managing a cluster
When you need custom operators and CRDs
When you require DaemonSets
When you must run your own CNI plugin
39. How do availability zones improve AKS resilience?
Availability zones are physically separate datacenters in a region. Spreading nodes across them lets the cluster survive the loss of one zone.
- Create node pools with
--zones 1 2 3. This must be set at pool creation. - On Standard and Premium tiers, the control plane is also zone-redundant in supported regions.
- Spread replicas using
topologySpreadConstraintsontopology.kubernetes.io/zone.
Watch storage: a standard Azure Disk lives in one zone, so a pod using it cannot restart in another zone. Use ZRS disks or Azure Files ZRS for stateful workloads that must fail over.
Take quiz
At node pool creation time
Any time using a ConfigMap
Only after the first upgrade
Never, AKS decides
Disks cannot be detached
The disk is tied to a single zone
Zones share one network card
The autoscaler blocks it
40. How does the cluster autoscaler differ from node auto-provisioning?
The cluster autoscaler grows and shrinks existing node pools. Every pool has a fixed VM size, so you must design the pools in advance. Node auto-provisioning (NAP), built on Karpenter, looks at pending pods and creates nodes of whatever VM size fits best, without predefined pools.
| Aspect | Cluster autoscaler | Node auto-provisioning |
| Unit of scaling | Node pool (VM scale set) | Individual nodes |
| VM size choice | Fixed per pool | Chosen per workload need |
| Config | Min/max per pool | NodePool and AKSNodeClass resources |
| Consolidation | Removes idle nodes | Actively repacks and replaces nodes to cut cost |
| Used by default in | AKS Standard | AKS Automatic |
NAP is attractive for mixed or bursty workloads. The classic autoscaler remains simpler and more predictable when you already run well-sized, purpose-built pools.
Take quiz
The Kubernetes version
The VM size that fits pending pods
The container registry
The ingress class
KEDA
Flux
Karpenter
Gatekeeper
41. How does the AKS control plane communicate with nodes?
Traffic flows in both directions, using different paths. Kubelets and kube-proxy initiate outbound connections to the API server. The API server needs to reach kubelets for kubectl logs, exec and port-forward, so AKS uses a reverse tunnel.
The konnectivity-agent pods in kube-system open an outbound, authenticated tunnel to the managed control plane. The API server sends kubelet requests through it, so no inbound firewall holes are needed on your nodes.
flowchart LR K["Kubelet on node"] -->|outbound 443| API["API server"] AG["konnectivity-agent pod"] -->|outbound tunnel| CP["Control plane konnectivity-server"] API --> CP CP -->|logs, exec, port-forward| AG AG --> K
This is why locked-down clusters must allow outbound access to the API server FQDN and required Microsoft endpoints. With API Server VNet Integration, nodes reach the API server directly over the VNet instead.
Take quiz
An open SSH port on every node
A public IP on each pod
An outbound konnectivity tunnel from the nodes
A VPN built into etcd
The API server, inbound to the node
The ingress controller
Azure Monitor
The kubelet, outbound from the node
42. Explain the execution flow of an AKS node pool upgrade?
AKS upgrades a node pool in a rolling fashion so pods keep running. The speed and capacity are controlled by max surge, which defaults to 10% of the pool for new pools (33% is a common production setting).
flowchart TD
A["Start pool upgrade"] --> B["Create surge node with new version"]
B --> C["Cordon an old node"]
C --> D["Drain: evict pods respecting PDBs"]
D --> E["Replace or reimage old node"]
E --> F{More old nodes?}
F -- Yes --> C
F -- No --> G["Remove surge nodes, upgrade complete"]
- Extra surge nodes are created first.
- A node is cordoned so nothing new is scheduled on it.
- It is drained; evictions respect PodDisruptionBudgets and termination grace periods.
- The node is replaced and rejoins, and the loop repeats.
If a PodDisruptionBudget cannot be satisfied within the drain timeout (30 minutes by default), the upgrade fails. Use --max-surge, --drain-timeout and --node-soak-duration to tune it.
Take quiz
How many pods a node can run
The size of the control plane
The number of namespaces
How many extra nodes can be added during the upgrade
A PodDisruptionBudget that cannot be satisfied
A missing ConfigMap
An unused Service
A public IP quota on ingress
43. What happens when an AKS node becomes NotReady?
When the kubelet stops reporting, the node controller marks the node Unknown after roughly 40 seconds and adds the node.kubernetes.io/unreachable and not-ready taints.
Pods carry a default toleration for these taints of 300 seconds. After five minutes without recovery they are evicted, and their Deployments or ReplicaSets create replacements on healthy nodes. Services stop routing to endpoints whose readiness fails.
Stateful pods are slower to recover because Azure Disks must be detached from the dead node and attached to another one.
AKS node auto-repair also watches for this. If a node stays NotReady for around 10 minutes, it escalates through reboot, reimage and redeploy. Check kubectl describe node for conditions such as memory pressure, disk pressure or a kubelet crash.
Take quiz
300 seconds
5 seconds
24 hours
They are evicted immediately
Deletes the whole cluster
Escalates through reboot, reimage and redeploy
Switches the CNI plugin
Re-creates the control plane
44. How do you troubleshoot pods stuck in Pending state?
Pending means the scheduler has not placed the pod yet. Start with the events, they almost always give the reason.
kubectl describe pod <pod> -n <ns> kubectl get events -n <ns> --sort-by=.lastTimestamp
Look for the FailedScheduling message and map it to a cause:
| Message | Likely cause | Fix |
| Insufficient cpu / memory | Requests exceed free capacity | Add nodes, lower requests, check autoscaler max |
| untolerated taint | Pool is tainted | Add toleration or use another pool |
| didn't match node selector / affinity | Wrong label or selector | Correct labels or the selector |
| unbound PersistentVolumeClaim | PVC cannot provision | Check StorageClass and zone |
| Too many pods | Node max-pods reached | Add nodes or raise max-pods on a new pool |
If the cluster autoscaler should have reacted, check its status ConfigMap and logs, and confirm the pod fits a node of the pool's VM size at all.
Take quiz
kubectl top node
kubectl describe pod
kubectl cordon
kubectl rollout undo
The image tag is wrong
DNS is down
The pod lacks a toleration for a tainted node
The PVC is too small
45. How do you troubleshoot ImagePullBackOff when pulling from ACR?
Run kubectl describe pod and read the pull error. It usually points to one of four things: authorization, a wrong image reference, networking, or registry throttling.
- Image name: confirm the registry, repository and tag exist (
az acr repository show-tags). - Auth: run
az aks check-acr -g rg -n aks --acr myregistry.azurecr.io, and confirm the kubelet identity holds AcrPull. - Network: with ACR firewall, private endpoint or UDR egress, check DNS resolution and rules from the node subnet.
- Throttling: pulls from Docker Hub can hit rate limits, so import the image into ACR.
A 401 or 403 points to permissions, manifest unknown to a missing tag, and a timeout to the network path. Also check for an architecture mismatch, such as an arm64-only image on amd64 nodes.
Take quiz
az acr purge
kubectl get acr
az aks check-acr
az aks nodepool upgrade
The node is out of memory
CoreDNS crashed
The pod lacks a toleration
The image tag does not exist in the repository
46. How can you optimize AKS costs?
Nodes are the biggest line item, so most savings come from running fewer, better-packed and cheaper nodes.
| Lever | How it saves |
| Right-size requests and limits | Stops over-reserving CPU and memory; the VPA add-on can recommend values |
| Autoscaling | Cluster autoscaler or NAP removes idle nodes; HPA and KEDA scale pods |
| Spot pools | Deep discounts for batch and stateless workloads |
| Reservations / savings plans | Commitment discounts on steady baseline nodes |
| Stop dev clusters | az aks stop pauses compute outside work hours |
| Efficient SKUs | ARM-based VMs and Azure Hybrid Benefit for Windows nodes |
| Right tier | Free tier for dev, Standard for production only |
Enable the AKS cost analysis add-on to attribute spend to namespaces and workloads, so you know where to look first. Review it monthly, because new deployments quietly change the picture.
Take quiz
az aks delete --soft
kubectl suspend cluster
az aks tier pause
az aks stop
Attributing spend to namespaces and workloads
Encrypting etcd
Upgrading node images
Creating ingress rules
47. How do you design a multi-region AKS architecture?
Run an independent cluster in each region (typically a paired region), and put a global entry point in front. Keep clusters stateless and rebuildable so any one of them can be lost.
flowchart TD U[Users] --> FD["Azure Front Door"] FD --> A["AKS East US"] FD --> B["AKS West US"] A --> DB[(Replicated data tier)] B --> DB ACR["ACR geo-replicated"] --> A ACR --> B G["Git repo with GitOps"] --> A G --> B
- Traffic: Azure Front Door or Traffic Manager with health probes for failover.
- Images: ACR Premium with geo-replication so each region pulls locally.
- Config parity: GitOps so both clusters apply the same manifests.
- Data: Cosmos DB multi-region writes or SQL failover groups, since Kubernetes does not replicate state.
- Fleet operations: Azure Kubernetes Fleet Manager to coordinate upgrades and propagate workloads.
Decide between active-active (lower failover time, higher cost) and active-passive (cheaper, slower to take over) based on your RTO and RPO.
Take quiz
Azure Front Door
Azure Bastion
Azure Key Vault
Azure DevTest Labs
Node image upgrades
ACR geo-replication
Pod topology spread
kubelet garbage collection
48. Explain how GitOps with Flux works in AKS?
With GitOps, a Git repository is the source of truth and an in-cluster agent continuously pulls and applies it. AKS offers Flux v2 as a managed extension (microsoft.flux).
flowchart LR
Dev["Developer pushes to Git"] --> Repo[(Git repo)]
Repo --> SC["source-controller fetches"]
SC --> KC["kustomize-controller / helm-controller"]
KC --> K["Cluster resources applied"]
K --> D{Drift from Git?}
D -- Yes --> KC
az k8s-configuration flux create -g rg-demo -c aks-demo -t managedClusters \ -n platform --namespace flux-system --scope cluster \ --url https://github.com/contoso/platform --branch main \ --kustomization name=apps path=./apps prune=true
The source-controller fetches the repo, then the kustomize or helm controller renders and applies it on every interval. Because it reconciles, manual changes made with kubectl are reverted, and prune=true deletes objects removed from Git.
Take quiz
The kubectl history
The Git repository
The node's local disk
The Azure portal blade
It is committed back to Git automatically
It is kept forever
Flux reverts it during reconciliation
Flux deletes the cluster
49. How do you achieve zero-downtime deployments in AKS?
It takes several settings working together, not just a rolling update.
- Run at least 2-3 replicas spread across nodes or zones.
- Use a rolling strategy with
maxUnavailable: 0so capacity never drops. - Add a readiness probe so traffic reaches only pods that are ready.
- Add a
preStopsleep and a sufficientterminationGracePeriodSecondsso in-flight requests finish and endpoints are removed first. - Create a PodDisruptionBudget so node drains and upgrades never take down too many pods.
strategy: type: RollingUpdate rollingUpdate: {maxSurge: 1, maxUnavailable: 0} template: spec: terminationGracePeriodSeconds: 45 containers: - name: api readinessProbe: httpGet: {path: /ready, port: 8080} lifecycle: preStop: exec: {command: ["sleep", "10"]}
Without the preStop delay, a pod can receive requests after it starts shutting down, because endpoint removal and SIGTERM race each other.
Take quiz
To disable readiness probes
To stop HPA from scaling
So serving capacity never drops during rollout
To run a single replica
Image pull failures
Expired TLS certificates
Slow DNS lookups
Too many pods being evicted at once during drains
50. How do you troubleshoot DNS resolution failures in AKS?
Work from the pod outward: confirm the symptom, then check CoreDNS, then the upstream path.
- Test from a pod:
kubectl run dnstest --rm -it --image=registry.k8s.io/e2e-test-images/agnhost:2.39 -- nslookup kubernetes.default. - Check CoreDNS:
kubectl get pods -n kube-system -l k8s-app=kube-dns, then read its logs for SERVFAIL or timeouts. - If only external names fail, look at the upstream: custom VNet DNS servers must forward to Azure DNS
168.63.129.16. - Check NSGs, firewalls or UDRs that block UDP/TCP 53 between nodes and the DNS servers.
- Review
/etc/resolv.confin the pod;ndots:5makes short names trigger many extra lookups.
Intermittent failures under load often mean the per-VM limit on Azure DNS (about 500 queries per second) is being hit. Deploying NodeLocal DNSCache and scaling CoreDNS reduces that pressure. Custom rules belong in the coredns-custom ConfigMap, not in edits to the managed CoreDNS config.