Prev Next

Java / Azure Kubernetes Service (AKS) Interview questions

Last updated

1. What is Azure Kubernetes Service (AKS)? 2. What are the main components of an AKS cluster? 3. What are node pools in AKS? 4. What is the difference between system and user node pools? 5. What are the AKS pricing tiers? 6. What is a pod in AKS? 7. What are the networking models available in AKS? 8. How do you create an AKS cluster using Azure CLI? 9. How do you connect to an AKS cluster using kubectl? 10. What is Azure Container Registry and how does AKS use it? 11. What are the types of Kubernetes Services? 12. What is an Ingress in AKS? 13. What is a managed identity in AKS? 14. What are ConfigMaps and Secrets in Kubernetes? 15. What is the cluster autoscaler in AKS? 16. What is the Horizontal Pod Autoscaler? 17. How do you scale an AKS cluster manually? 18. What are persistent volumes in AKS? 19. What is Container Insights in AKS? 20. What is AKS Automatic? 21. How do you upgrade an AKS cluster? 22. What is the difference between kubenet and Azure CNI? 23. When would you choose Azure CNI Overlay? 24. How does AKS integrate with Microsoft Entra ID? 25. What is the difference between Kubernetes RBAC and Azure RBAC in AKS? 26. How does Workload Identity work in AKS? 27. How do you use Azure Key Vault secrets in AKS? 28. What is the application routing add-on in AKS? 29. Why is Gateway API replacing Ingress in AKS? 30. When should you use Azure Disk versus Azure Files in AKS? 31. What is KEDA and when would you use it in AKS? 32. How do you implement network policies in AKS? 33. How do you enforce policies on AKS with Azure Policy? 34. How do you create a private AKS cluster? 35. What is the difference between node image upgrades and Kubernetes version upgrades? 36. How do you use spot node pools in AKS? 37. What are taints and tolerations in AKS? 38. Which is better: AKS or Azure Container Apps? 39. How do availability zones improve AKS resilience? 40. How does the cluster autoscaler differ from node auto-provisioning? 41. How does the AKS control plane communicate with nodes? 42. Explain the execution flow of an AKS node pool upgrade? 43. What happens when an AKS node becomes NotReady? 44. How do you troubleshoot pods stuck in Pending state? 45. How do you troubleshoot ImagePullBackOff when pulling from ACR? 46. How can you optimize AKS costs? 47. How do you design a multi-region AKS architecture? 48. Explain how GitOps with Flux works in AKS? 49. How do you achieve zero-downtime deployments in AKS? 50. How do you troubleshoot DNS resolution failures in AKS?

1. What is Azure Kubernetes Service (AKS)?

Azure Kubernetes Service (AKS) is Microsoft's managed Kubernetes offering. Azure runs the control plane for you (API server, etcd, scheduler and controller manager), and you only manage the worker nodes that run your containers.

Day-to-day chores such as provisioning the cluster, patching, version upgrades and health monitoring of the control plane are handled by the platform. You keep control over node sizes, networking, scaling rules and the workloads you deploy.

AKS plugs into the rest of Azure: Microsoft Entra ID for sign-in, Azure Container Registry for images, Azure Monitor for telemetry, and virtual networks for private connectivity.

Take quiz
In AKS, who operates the Kubernetes control plane?
You operate it on VMs in your subscription
A pod in the kube-system namespace operates it
Azure operates it for you
Your container registry operates it
Which resources do you pay for in the AKS Free tier?
The API server and etcd by the hour
Only the Azure Container Registry
Only the pods that are running
The worker node VMs and attached resources

2. What are the main components of an AKS cluster?

An AKS cluster has two halves: a Microsoft-managed control plane and customer-owned nodes.

Part Component Role
Control plane kube-apiserver Front door for every kubectl and controller request
Control plane etcd Key-value store holding all cluster state
Control plane kube-scheduler Picks a node for each new pod
Control plane kube-controller-manager Runs loops such as the ReplicaSet and node controllers
Node kubelet Agent that starts and watches pods on the node
Node containerd Container runtime that pulls images and runs containers
Node kube-proxy Programs service routing rules (unless Cilium replaces it)

The nodes are Azure VMs grouped into node pools. AKS places them, along with load balancers and disks it creates, in a separate node resource group whose name starts with MC_.

Take quiz
Which component stores the cluster's desired and current state?
kubelet
kube-proxy
containerd
etcd
What does the MC_ resource group contain?
The VM scale sets, load balancers and disks AKS creates for the cluster
The control plane virtual machines
Backup copies of etcd
Only your container registry

3. What are node pools in AKS?

A node pool is a group of worker nodes that share the same VM size, OS type and settings. On AKS each pool is backed by a Virtual Machine Scale Set, so scaling a pool just changes the scale set instance count.

Pools let you mix hardware in one cluster. You might keep a general-purpose pool for web pods, a GPU pool for model inference and a Windows pool for a legacy .NET service, then steer pods to the right pool with labels and taints.

az aks nodepool add \
  --resource-group rg-demo --cluster-name aks-demo \
  --name gpupool --node-count 2 \
  --node-vm-size Standard_NC6s_v3 \
  --node-taints sku=gpu:NoSchedule

Take quiz
What Azure resource backs an AKS node pool?
A Virtual Machine Scale Set
An Azure App Service plan
A single Azure VM shared by all pods
An Azure Batch pool
Why would you add a second node pool to an existing cluster?
To get a second control plane
To run workloads that need a different VM size, such as GPUs
To avoid paying for node VMs
To store container images closer to pods

4. What is the difference between system and user node pools?

System node pools host AKS's critical pods such as CoreDNS, metrics-server and the konnectivity agent. User node pools are where your own application pods run.

Aspect System pool User pool
Purpose Cluster-critical add-on pods Application workloads
OS Linux only Linux or Windows
Minimum count At least 1 pool, 2+ nodes for production Can scale down to 0
Typical taint CriticalAddonsOnly=true:NoSchedule None by default

Every cluster needs at least one system pool. Tainting it keeps application pods off those nodes, so a noisy app cannot starve CoreDNS.

Take quiz
Which node pool type can be scaled to zero nodes?
The system node pool
A user node pool
Both, with no restriction
Neither pool type
What taint is commonly used to dedicate a pool to system pods?
kubernetes.io/os=windows:NoExecute
node.kubernetes.io/unreachable:NoSchedule
CriticalAddonsOnly=true:NoSchedule
sku=gpu:PreferNoSchedule

5. What are the AKS pricing tiers?

AKS charges nothing for the control plane on the Free tier, but there is no financially backed SLA, so it suits dev and test. You still pay for the nodes.

Tier What you get Best for
Free Managed control plane, no uptime SLA, lower scale limits Learning, dev/test
Standard Financially backed uptime SLA (higher with availability zones), up to 5,000 nodes Production
Premium Everything in Standard plus long-term support (LTS) for a Kubernetes version Regulated or slow-upgrading estates

Standard and Premium add a per-cluster hourly fee on top of node costs. Check the Azure pricing page for current rates.

Take quiz
Which AKS tier adds long-term support for Kubernetes versions?
Free
Standard
Premium
All tiers include it
Which tier should a production cluster use for a financially backed uptime SLA?
Free
Developer
Basic
Standard

6. What is a pod in AKS?

A pod is the smallest deployable unit in Kubernetes. It wraps one or more containers that share a network namespace (one IP, talking over localhost) and can share mounted volumes.

Pods are disposable. If a node dies, the pod is not moved, a controller such as a Deployment creates a replacement. That is why production workloads should not be run as bare pods.

apiVersion: v1
kind: Pod
metadata:
  name: web
spec:
  containers:
  - name: nginx
    image: nginx:1.27
    ports:
    - containerPort: 80

Take quiz
What do containers inside the same pod share?
The node's root filesystem
A separate IP each
Nothing, they are fully isolated
The pod IP address and localhost networking
Why avoid running bare pods in production?
Nothing recreates them if the node fails
They cannot use environment variables
They cannot pull images from ACR
They always run as root

7. What are the networking models available in AKS?

AKS supports several CNI choices that differ mainly in where pod IPs come from.

Model Pod IP source Notes
kubenet Private range, NATed behind the node Legacy; retiring on 31 March 2028
Azure CNI Node Subnet Same VNet subnet as nodes Pods are directly routable; burns many IPs
Azure CNI Pod Subnet Separate VNet subnet for pods Routable pods with better IP planning
Azure CNI Overlay Private overlay CIDR Saves VNet space; scales to large clusters

The Cilium dataplane can be layered on Overlay or Pod Subnet to get eBPF networking and network policy. For new clusters, Overlay is the usual starting point.

Take quiz
Which AKS networking model retires on 31 March 2028?
kubenet
Azure CNI Overlay
Azure CNI Pod Subnet
Azure CNI powered by Cilium
Which model gives pods VNet IPs from a subnet separate from the nodes?
kubenet
Azure CNI Pod Subnet
Azure CNI Overlay
ExternalName networking

8. How do you create an AKS cluster using Azure CLI?

You need a resource group first, then a single az aks create call.

  1. Sign in and pick the subscription with az login.
  2. Create a resource group.
  3. Create the cluster with your node count, VM size and network plugin.
  4. Pull credentials and verify with kubectl.
az group create -n rg-demo -l eastus

az aks create -g rg-demo -n aks-demo \
  --node-count 3 --node-vm-size Standard_D4s_v5 \
  --network-plugin azure --network-plugin-mode overlay \
  --generate-ssh-keys

az aks get-credentials -g rg-demo -n aks-demo
kubectl get nodes

Provisioning usually takes a few minutes. Add --tier standard for a production SLA.

Take quiz
Which Azure CLI command creates an AKS cluster?
az kubernetes deploy
az aks create
az container create
az aks init
Which flag sets the initial number of nodes?
--replicas
--pod-count
--node-count
--scale-size

9. How do you connect to an AKS cluster using kubectl?

Run az aks get-credentials. It downloads the cluster's connection details and merges them into ~/.kube/config as a new context.

az aks get-credentials -g rg-demo -n aks-demo
kubectl config current-context
kubectl get nodes

If the cluster uses Microsoft Entra ID, kubectl needs kubelogin to fetch tokens (az aks install-cli installs both). The --admin flag fetches a local admin certificate, but it only works while local accounts are enabled and is best avoided in production.

Take quiz
Where does az aks get-credentials write cluster access details?
The Azure Key Vault
/etc/kubernetes on the node
The kubeconfig file, ~/.kube/config
The cluster's etcd database
Which tool handles Entra ID token login for kubectl?
kubeadm
helm
kustomize
kubelogin

10. What is Azure Container Registry and how does AKS use it?

Azure Container Registry (ACR) is Azure's private, managed registry for container images and Helm charts. AKS pulls application images from it when pods start.

The cleanest integration is to attach the registry to the cluster. This grants the AcrPull role to the cluster's kubelet managed identity, so you do not need image pull secrets.

az aks update -g rg-demo -n aks-demo --attach-acr myregistry

# reference the image in your manifest
# image: myregistry.azurecr.io/shop/api:1.4.2

Take quiz
Which role does --attach-acr grant?
Contributor
AcrPush
Owner
AcrPull
Which identity receives that role?
The kubelet managed identity
The developer's user account
The ingress controller
The pod's service account

11. What are the types of Kubernetes Services?

A Service gives a stable virtual IP and DNS name to a changing set of pods. The type decides who can reach it.

Type Reachable from On AKS
ClusterIP Inside the cluster only (default) Internal virtual IP
NodePort Node IP on ports 30000-32767 Opens that port on every node
LoadBalancer Outside the cluster Creates an Azure Load Balancer rule with a public IP
ExternalName Inside the cluster Returns a CNAME to an external DNS name

For a private load balancer, add the annotation service.beta.kubernetes.io/azure-load-balancer-internal: "true" to the Service.

Take quiz
Which Service type is the default?
ClusterIP
NodePort
LoadBalancer
ExternalName
What does a LoadBalancer Service create on AKS?
A new virtual network
An Azure Load Balancer rule and frontend IP
An Application Gateway WAF policy
A Traffic Manager profile

12. What is an Ingress in AKS?

An Ingress is a Kubernetes object that routes external HTTP and HTTPS traffic to Services based on host name and URL path, and can terminate TLS. It works at layer 7, so one public IP can front many apps.

The object does nothing on its own. An ingress controller must be running to read the rules and configure a proxy. On AKS that is commonly the application routing add-on, Application Gateway for Containers or a self-managed controller.

The community Ingress-NGINX project has been retired and the Ingress API is frozen, so newer designs lean towards Gateway API.

Take quiz
At which OSI layer does an Ingress route traffic?
Layer 2 (Ethernet)
Layer 7 (HTTP/HTTPS)
Layer 3 only (IP)
Layer 4 only (TCP)
What must exist for an Ingress resource to take effect?
A NodePort on every service
A second node pool
An ingress controller
A private DNS zone

13. What is a managed identity in AKS?

A managed identity is an Azure-managed credential that AKS uses to call Azure APIs, so you never store or rotate a service principal secret.

A cluster uses two of them. The control plane identity lets AKS create and manage resources such as load balancers, routes and disks. The kubelet identity is used by nodes, mainly to pull images from ACR.

You can let AKS create them (system-assigned for the control plane) or bring your own user-assigned identities, which is handy when you must pre-grant roles before the cluster exists.

Take quiz
What is the main benefit of managed identities over service principals?
They make pods start faster
They remove the need for RBAC
No client secret to store or rotate
They replace the container runtime
Which AKS identity is typically used to pull images from ACR?
The control plane identity
The ingress identity
The Entra admin group
The kubelet identity

14. What are ConfigMaps and Secrets in Kubernetes?

A ConfigMap holds non-sensitive settings such as feature flags or a log level. A Secret holds sensitive values such as passwords and tokens. Both can reach a pod as environment variables or as files in a mounted volume.

Secret values are only base64-encoded, which is encoding, not encryption. Anyone with read access to the Secret can decode it. AKS encrypts etcd at rest, and you can add KMS encryption with Key Vault, but tight RBAC matters just as much.

kubectl create configmap app-config --from-literal=LOG_LEVEL=info
kubectl create secret generic db-pass --from-literal=password='S3cret!'

Take quiz
Which object should hold a non-sensitive feature flag?
A Secret
A PersistentVolume
A ServiceAccount
A ConfigMap
What does base64 encoding of a Secret provide?
Encoding only, not protection from readers
Strong AES encryption
Automatic rotation
Hiding the value from kubectl

15. What is the cluster autoscaler in AKS?

The cluster autoscaler adds nodes to a node pool when pods cannot be scheduled for lack of CPU or memory, and removes nodes that stay underused.

It is enabled per node pool with a minimum and maximum count. By default a node becomes a scale-down candidate after it has been unneeded for about 10 minutes and its utilization is below 50%.

az aks nodepool update -g rg-demo --cluster-name aks-demo -n userpool \
  --enable-cluster-autoscaler --min-count 2 --max-count 8

Take quiz
What triggers the cluster autoscaler to add a node?
Pods that cannot be scheduled due to insufficient resources
High CPU on the control plane
A new Docker image push
A failed liveness probe
How do you bound the autoscaler for a pool?
With --replicas and --pods
With --min-count and --max-count
With a ResourceQuota
With --node-taints

16. What is the Horizontal Pod Autoscaler?

The Horizontal Pod Autoscaler (HPA) changes the replica count of a Deployment based on observed metrics, usually CPU or memory from metrics-server, which AKS ships by default.

The controller computes desired = ceil(current replicas x current metric / target metric). For CPU-based scaling, pods must declare resource requests, otherwise utilization cannot be calculated.

kubectl autoscale deployment web --cpu-percent=70 --min=3 --max=15

Take quiz
What does the HPA change?
The VM size of nodes
The number of pod replicas
The container image tag
The Kubernetes version
What must pods define for CPU-percentage scaling to work?
A NodePort
An init container
Resource requests
An ingress annotation

17. How do you scale an AKS cluster manually?

Scale nodes with az aks scale (or az aks nodepool scale for one pool) and scale pods with kubectl scale.

# nodes
az aks nodepool scale -g rg-demo --cluster-name aks-demo -n userpool --node-count 5

# pods
kubectl scale deployment web --replicas=10

If the cluster autoscaler is enabled on that pool, you cannot set a fixed count. Change the min and max values, or disable the autoscaler first.

Take quiz
Which command resizes a single node pool?
kubectl resize nodes
az aks pods scale
az aks nodepool scale
az vmss reimage
Why does a manual node count change fail on an autoscaled pool?
Pools can never be resized
The control plane is locked
kubectl must be used instead
The autoscaler owns the node count

18. What are persistent volumes in AKS?

A PersistentVolume (PV) is storage that outlives any single pod. A pod asks for it through a PersistentVolumeClaim (PVC), and a StorageClass tells Kubernetes how to provision it dynamically.

AKS ships CSI-based classes such as managed-csi (Azure Disk) and azurefile-csi (Azure Files). When a PVC is created, the driver provisions the Azure resource and binds it automatically.

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: data-pvc
spec:
  accessModes: [ReadWriteOnce]
  storageClassName: managed-csi
  resources:
    requests:
      storage: 20Gi

Take quiz
What does a pod use to request persistent storage?
A ConfigMap
A NodePort
A ServiceAccount
A PersistentVolumeClaim
Which built-in StorageClass provisions Azure Disks?
managed-csi
azurefile-csi
local-path
ephemeral-os

19. What is Container Insights in AKS?

Container Insights is the Azure Monitor feature that collects container logs, node and pod performance data and Kubernetes events, and sends them to a Log Analytics workspace.

Enabling the monitoring add-on deploys the ama-logs agent as a DaemonSet. You then query data with KQL or use the built-in workbooks in the portal.

az aks enable-addons -g rg-demo -n aks-demo -a monitoring

For metrics, Azure Monitor managed service for Prometheus and managed Grafana are the usual companions.

Take quiz
Where does Container Insights store its data?
In a Log Analytics workspace
In the cluster's etcd
In Azure Container Registry
On each node's local disk only
How is the Container Insights agent deployed to nodes?
As a StatefulSet
As a DaemonSet
As a CronJob
As a static VM extension only

20. What is AKS Automatic?

AKS Automatic is an opinionated cluster SKU where Azure manages far more than the control plane, including node provisioning, scaling, upgrades and many security defaults.

It comes pre-configured with node auto-provisioning (Karpenter-based), Azure CNI Overlay with Cilium, managed Prometheus and Grafana, Entra ID with Azure RBAC, workload identity and deployment safeguards. You focus on deploying apps.

az aks create -g rg-demo -n aks-auto --sku automatic

The trade-off is less freedom to tweak low-level settings than on AKS Standard.

Take quiz
Which flag creates an AKS Automatic cluster?
--mode auto-pilot
--sku automatic
--tier premium
--managed-nodes true
How are nodes managed in AKS Automatic?
You hand-pick each VM size
Only fixed-size pools are allowed
They are provisioned automatically based on pod needs
Nodes are never patched

21. How do you upgrade an AKS cluster?

AKS upgrades the control plane first, then the node pools. You pick a supported target and AKS performs a rolling upgrade.

  1. List what is available with az aks get-upgrades.
  2. Test the new version in a lower environment and check for removed APIs.
  3. Run az aks upgrade (use --control-plane-only to stage the nodes later).
  4. Confirm node versions with kubectl get nodes.
az aks get-upgrades -g rg-demo -n aks-demo -o table
az aks upgrade -g rg-demo -n aks-demo --kubernetes-version 1.33.2

You cannot skip minor versions, so 1.31 goes to 1.32 before 1.33. To automate this, set an auto-upgrade channel such as patch, stable or rapid and define a planned maintenance window.

Take quiz
Can you skip a minor version during an AKS upgrade?
Yes, go directly to the latest
Only on the Free tier
No, minor versions must be upgraded one at a time
Only for node pools
What does --control-plane-only do?
Upgrades only system pods
Skips the API server
Reboots every node first
Upgrades only the control plane, leaving nodes for later

22. What is the difference between kubenet and Azure CNI?

With kubenet nodes get VNet IPs while pods get addresses from a separate private range and are NATed through the node. With Azure CNI pods receive routable VNet IPs directly.

Aspect kubenet Azure CNI (node subnet)
Pod IP Private range, NATed VNet IP
VNet IP use One per node One per node plus one per pod
Default max pods/node 110 30 (up to 250)
Network policy Calico only Azure NPM, Calico or Cilium
Virtual nodes, Windows pools Not supported Supported
Future Retires 31 March 2028 Supported

Plan to move kubenet clusters to Azure CNI Overlay, which keeps the low IP footprint without kubenet's limits.

Take quiz
How do pods get IP addresses under kubenet?
A dedicated VNet IP each
From Azure DHCP per pod
From the control plane subnet
From a private range, NATed through the node
Which is supported with Azure CNI but not with kubenet?
Windows node pools
Linux node pools
Services of type ClusterIP
Persistent volume claims

23. When would you choose Azure CNI Overlay?

Choose Azure CNI Overlay when VNet address space is scarce or you expect a large cluster. Pods draw IPs from a private overlay CIDR (default 10.244.0.0/16), so only nodes consume VNet addresses.

It supports up to 250 pods per node and clusters of up to 5,000 nodes, and the same pod CIDR can be reused across clusters because it is never advertised on the VNet.

The catch is that pods are not directly routable from outside the cluster. External systems must reach workloads through Services, ingress or load balancers. If something must address pod IPs directly, pick Pod Subnet instead.

az aks create -g rg-demo -n aks-demo \
  --network-plugin azure --network-plugin-mode overlay \
  --pod-cidr 192.168.0.0/16

Take quiz
What is the key advantage of Azure CNI Overlay?
Pods do not consume VNet IP addresses
Pods become directly routable from on-premises
It removes the need for node pools
It disables network policy
What limitation comes with Overlay networking?
It only runs Windows pods
Pod IPs are not directly reachable from outside the cluster
It caps clusters at 10 nodes
It forbids load balancers

24. How does AKS integrate with Microsoft Entra ID?

With AKS-managed Entra integration, users sign in with their Entra identity and the API server validates the token. You nominate Entra groups as cluster admins, and access is then decided by Kubernetes RBAC or Azure RBAC.

sequenceDiagram
  participant U as User
  participant E as Microsoft Entra ID
  participant A as AKS API server
  U->>E: az login / kubelogin
  E-->>U: Access token
  U->>A: kubectl request + token
  A->>E: Validate token
  A-->>U: Allow or deny via RBAC
az aks create -g rg-demo -n aks-demo \
  --enable-aad --aad-admin-group-object-ids <group-id> \
  --disable-local-accounts

Disabling local accounts removes shared admin certificates, so every action is tied to a named identity.

Take quiz
What does the AKS API server do with the token from kubelogin?
Stores it in etcd forever
Validates it against Microsoft Entra ID
Forwards it to the container registry
Ignores it and uses a certificate
Why disable local accounts on an Entra-integrated cluster?
To speed up node pool scaling
To enable Windows containers
To force every request to use a named Entra identity
To remove the need for RBAC

25. What is the difference between Kubernetes RBAC and Azure RBAC in AKS?

Both answer the question "what can this user do?", but they store the rules in different places.

Aspect Kubernetes RBAC Azure RBAC for Kubernetes
Rules live in Role, ClusterRole and bindings inside the cluster Azure role assignments in IAM
Managed with kubectl and YAML Azure portal, CLI, policy
Roles Custom, any verbs you define Built-in such as RBAC Reader, Writer, Admin, Cluster Admin
Scope Namespace or cluster Cluster, namespace or subscription hierarchy

Azure RBAC for Kubernetes authorization is switched on with --enable-azure-rbac. Do not confuse it with the Azure roles that control the AKS resource itself (for example Contributor), which govern who can scale or delete the cluster, not who can read pods.

Take quiz
Where are Kubernetes RBAC rules stored?
In Azure IAM only
In the container registry
Inside the cluster as Role and RoleBinding objects
On each node's disk
Which flag enables Azure RBAC authorization for Kubernetes objects?
--enable-rbac-v2
--azure-roles-only
--enable-entra-roles
--enable-azure-rbac

26. How does Workload Identity work in AKS?

Workload Identity lets a pod call Azure services (Key Vault, Storage, SQL) as a managed identity without any secret. It uses OIDC federation between the cluster and Microsoft Entra ID.

flowchart LR
  A["Pod with annotated service account"] --> B["Projected service account token"]
  B --> C["Microsoft Entra ID checks federated credential"]
  C --> D["Entra issues Azure access token"]
  D --> E["Pod calls Key Vault or Storage"]
  1. Enable --enable-oidc-issuer --enable-workload-identity on the cluster.
  2. Create a user-assigned managed identity and grant it the needed Azure role.
  3. Create a federated credential tied to system:serviceaccount:<ns>:<sa>.
  4. Annotate the service account with the identity's client ID and label the pod azure.workload.identity/use: "true".

It replaces the retired AAD Pod Identity.

Take quiz
What links a Kubernetes service account to an Azure managed identity?
A shared client secret in a ConfigMap
A node-level SSH key
An ACR token
A federated identity credential
Which pod label opts a pod in to Workload Identity?
azure.workload.identity/use: "true"
aks.azure.com/identity: on
workload.pod.io/enabled: true
kubernetes.io/entra: enabled

27. How do you use Azure Key Vault secrets in AKS?

Use the Secrets Store CSI Driver with the Azure Key Vault provider. It mounts secrets, keys or certificates from Key Vault into the pod as files, so values never sit in your manifests.

  1. Enable the add-on: az aks enable-addons -a azure-keyvault-secrets-provider.
  2. Give the pod an identity (Workload Identity is recommended) with Get permission on the vault.
  3. Define a SecretProviderClass naming the vault and objects.
  4. Mount it as a CSI volume in the pod spec.

Secrets can optionally be synced to a native Kubernetes Secret if an app insists on environment variables. Enable rotation polling if you want updated vault values to appear in running pods.

Take quiz
How do Key Vault secrets reach the pod with the CSI driver?
Mounted as files in a volume
Copied into the container image
Baked into the node VM image
Sent by email to the developer
Which resource describes which vault objects to mount?
StorageClass
SecretProviderClass
IngressClass
PriorityClass

28. What is the application routing add-on in AKS?

The application routing add-on is a managed ingress solution. It deploys NGINX ingress controllers and wires up Azure DNS records and Key Vault-backed TLS certificates for you.

az aks approuting enable -g rg-demo -n aks-demo

# Ingress uses this class
# ingressClassName: webapprouting.kubernetes.azure.com

Upstream Ingress-NGINX reached end of life in March 2026. Microsoft committed to critical security patches for the add-on's NGINX implementation only through November 2026, so existing users should plan a move to the add-on's Gateway API implementation or another Gateway API option.

Take quiz
Which ingress class does the application routing add-on register?
alb.azure.com
webapprouting.kubernetes.azure.com
istio.io/ingress
nginx.ingress.azure.net
Until when does Microsoft promise security patches for its managed NGINX path?
March 2028
June 2030
November 2026
There is no committed date

29. Why is Gateway API replacing Ingress in AKS?

The Ingress API is frozen and relies on controller-specific annotations for anything beyond basic host and path routing. Gateway API is its successor and models traffic with separate, role-based resources.

Aspect Ingress Gateway API
Resources One Ingress object GatewayClass, Gateway, HTTPRoute, GRPCRoute
Ownership Single team edits everything Platform team owns Gateway, app teams own routes
Advanced routing Vendor annotations Header matches, weighted traffic splitting in the spec
Portability Varies by controller Standard across conformant implementations

On AKS the main options are Application Gateway for Containers, the Istio-based Gateway API mode in the application routing add-on (in preview when announced, so verify its status) and the Istio service mesh add-on. With Ingress-NGINX retired, new designs should start on Gateway API.

Take quiz
Which Gateway API object do application teams typically define for routing?
GatewayClass
StorageClass
HTTPRoute
NetworkPolicy
What is a weakness of Ingress that Gateway API fixes?
It cannot terminate TLS
It only supports UDP
It cannot run on AKS
Reliance on controller-specific annotations for advanced routing

30. When should you use Azure Disk versus Azure Files in AKS?

Use Azure Disk for single-writer, latency-sensitive data such as a database. Use Azure Files when several pods on different nodes must read and write the same files.

Aspect Azure Disk Azure Files
Access mode ReadWriteOnce (one node at a time) ReadWriteMany
Protocol Block device SMB or NFS share
Performance Low latency, high IOPS on Premium/Ultra Higher latency, shared throughput
Zone behavior Zonal unless ZRS disk is used ZRS shares survive a zone failure
Typical use Databases, queues, single-instance stateful apps Shared uploads, CMS content, config sharing

Also keep the per-VM disk attach limit in mind when packing many stateful pods on one node.

Take quiz
Which storage type supports ReadWriteMany across nodes?
Azure Disk
Node local temp disk
Azure Disk with LRS
Azure Files
A single-instance database needing low latency fits best on:
Azure Disk
Azure Files over SMB
An emptyDir volume
A ConfigMap volume

31. What is KEDA and when would you use it in AKS?

KEDA (Kubernetes Event-driven Autoscaling) scales workloads based on external event sources such as Service Bus queue length, Event Hubs lag, Kafka topics or a Prometheus query. AKS offers it as a managed add-on with --enable-keda.

Unlike plain HPA, KEDA can scale a Deployment down to zero when there is no work and back up when events arrive. For one-or-more replicas it feeds an HPA behind the scenes.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: orders-worker
spec:
  scaleTargetRef:
    name: orders-worker
  minReplicaCount: 0
  maxReplicaCount: 30
  triggers:
  - type: azure-servicebus
    metadata:
      queueName: orders
      messageCount: "20"

Take quiz
What can KEDA do that a CPU-based HPA cannot?
Scale to zero based on queue or event depth
Resize node VMs
Upgrade the control plane
Encrypt etcd
Which custom resource defines a KEDA scaling rule?
SecretProviderClass
ScaledObject
PodDisruptionBudget
NginxIngressController

32. How do you implement network policies in AKS?

By default every pod can talk to every other pod. A NetworkPolicy restricts that, but it only works if the cluster was created with a policy engine: azure, calico or cilium set through --network-policy.

A common pattern is a default-deny rule per namespace, then explicit allows. Policies are additive, so any matching allow rule opens the path.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-web-to-api
  namespace: shop
spec:
  podSelector:
    matchLabels: {app: api}
  policyTypes: [Ingress]
  ingress:
  - from:
    - podSelector:
        matchLabels: {app: web}
    ports:
    - port: 8080

For new clusters Cilium (eBPF) is generally the direction Microsoft favors, but check current docs before choosing.

Take quiz
What happens to pod traffic when no NetworkPolicy exists?
All traffic is denied
All pod-to-pod traffic is allowed
Only same-namespace traffic is allowed
Only HTTPS is allowed
Which flag selects the policy engine at cluster creation?
--pod-firewall
--policy-engine
--network-policy
--enable-acl

33. How do you enforce policies on AKS with Azure Policy?

Enable the Azure Policy add-on. It installs Gatekeeper (OPA) admission webhooks and syncs Azure Policy definitions into the cluster as constraints.

az aks enable-addons -g rg-demo -n aks-demo -a azure-policy

You then assign built-in definitions or initiatives at subscription, resource group or cluster scope. Typical rules block privileged containers, require resource limits, allow images only from approved registries or forbid hostPath mounts.

Each policy runs in Audit mode (report only) or Deny mode (reject at admission). Starting in Audit lets you see what would break before enforcing. Compliance shows up in the Azure Policy dashboard.

Take quiz
Which engine does the Azure Policy add-on use inside the cluster?
Calico
Flux
Gatekeeper (OPA)
CoreDNS
Which effect only reports violations without blocking them?
Deny
Disabled-by-admission
Mutate-all
Audit

34. How do you create a private AKS cluster?

A private cluster exposes the API server only on a private IP inside your VNet, through Private Link, so there is no public endpoint.

az aks create -g rg-demo -n aks-private \
  --enable-private-cluster \
  --private-dns-zone system \
  --network-plugin azure --network-plugin-mode overlay

Because the endpoint is private, kubectl must run from somewhere with a network path: a VM in the VNet or a peered one, a VPN or ExpressRoute link, or a self-hosted CI agent. For quick one-off commands, az aks command invoke runs kubectl through the Azure control plane.

An alternative is API Server VNet Integration, which places the API server behind an internal load balancer in a delegated subnet and avoids Private Link.

Take quiz
How does a private cluster expose its API server?
On a public IP with a firewall rule
Through a NodePort on each node
Only through the Azure portal
On a private IP inside the VNet via Private Link
Which command lets you run kubectl against a private cluster without VNet access?
az aks command invoke
az aks proxy
kubectl --public
az network bastion

35. What is the difference between node image upgrades and Kubernetes version upgrades?

A Kubernetes version upgrade moves the control plane and kubelet to a new Kubernetes release. A node image upgrade keeps the Kubernetes version but replaces the node OS image with a newer one containing OS and runtime security patches.

Aspect Kubernetes upgrade Node image upgrade
Changes Kubernetes minor or patch version OS, containerd and kernel patches
Cadence Few times a year New images roughly weekly
Command az aks upgrade az aks nodepool upgrade --node-image-only
API impact Possible API deprecations None

Many teams automate node images with the NodeImage or SecurityPatch auto-upgrade channel and reserve Kubernetes upgrades for planned windows.

Take quiz
Which upgrade replaces the node OS image without changing the Kubernetes version?
A node image upgrade
A minor version upgrade
A control-plane-only upgrade
A CNI migration
Which flag limits a node pool upgrade to the OS image?
--os-only-reboot
--node-image-only
--no-kubelet
--patch-os

36. How do you use spot node pools in AKS?

A spot node pool runs on Azure Spot VMs, which are much cheaper but can be evicted with about 30 seconds' notice when Azure needs the capacity. Use them for interruption-tolerant work such as batch jobs, CI runners and stateless workers.

az aks nodepool add -g rg-demo --cluster-name aks-demo -n spotpool \
  --priority Spot --eviction-policy Delete --spot-max-price -1 \
  --enable-cluster-autoscaler --min-count 0 --max-count 10

AKS automatically taints these nodes with kubernetes.azure.com/scalesetpriority=spot:NoSchedule, so pods must carry a matching toleration. Spot pools cannot be the system pool, and critical or stateful services should stay on regular nodes.

Take quiz
Which pool type cannot be a spot pool?
A user node pool
The system node pool
A GPU pool
A Windows pool
How do pods get scheduled onto spot nodes?
They request a NodePort
They set a ConfigMap flag
They tolerate the spot taint AKS applies
They run as privileged

37. What are taints and tolerations in AKS?

A taint on a node repels pods that do not tolerate it. A toleration on a pod allows, but does not force, scheduling onto a matching tainted node.

Effects are NoSchedule (do not place new pods), PreferNoSchedule (try to avoid) and NoExecute (also evict running pods).

# taint at pool creation
az aks nodepool add ... --node-taints sku=gpu:NoSchedule

# pod spec
tolerations:
- key: sku
  operator: Equal
  value: gpu
  effect: NoSchedule

Because a toleration only permits placement, pair it with a nodeSelector or node affinity if the pod must land on that pool.

Take quiz
What does a toleration do?
Forces the pod onto that node
Deletes the taint
Permits a pod to schedule on a tainted node
Adds a second container
Which taint effect also evicts pods already running?
NoSchedule
PreferNoSchedule
AllowOnly
NoExecute

38. Which is better: AKS or Azure Container Apps?

Neither wins outright, it depends on how much Kubernetes you actually need. Azure Container Apps (ACA) is a serverless platform built on Kubernetes, KEDA and Envoy that hides the Kubernetes API. AKS hands you the full API.

Aspect AKS Azure Container Apps
Kubernetes API access Full (CRDs, operators, DaemonSets) None
Operations burden Higher: nodes, upgrades, networking Low: no nodes to manage
Scale to zero With KEDA Built in
Customization Very high Limited to platform features
Best for Platform teams, complex or multi-tenant estates Microservices, APIs, event-driven apps, jobs

Pick ACA when you want containers running fast without cluster operations. Pick AKS when you need operators, custom networking, specific schedulers or portability of existing Kubernetes tooling.

Take quiz
Which service exposes the full Kubernetes API?
Azure Container Apps
Azure Container Instances
Azure Functions
AKS
When does Azure Container Apps fit better than AKS?
When you want to run containers without managing a cluster
When you need custom operators and CRDs
When you require DaemonSets
When you must run your own CNI plugin

39. How do availability zones improve AKS resilience?

Availability zones are physically separate datacenters in a region. Spreading nodes across them lets the cluster survive the loss of one zone.

  • Create node pools with --zones 1 2 3. This must be set at pool creation.
  • On Standard and Premium tiers, the control plane is also zone-redundant in supported regions.
  • Spread replicas using topologySpreadConstraints on topology.kubernetes.io/zone.

Watch storage: a standard Azure Disk lives in one zone, so a pod using it cannot restart in another zone. Use ZRS disks or Azure Files ZRS for stateful workloads that must fail over.

Take quiz
When must node pool zones be chosen?
At node pool creation time
Any time using a ConfigMap
Only after the first upgrade
Never, AKS decides
Why can a pod with a standard Azure Disk fail to restart in another zone?
Disks cannot be detached
The disk is tied to a single zone
Zones share one network card
The autoscaler blocks it

40. How does the cluster autoscaler differ from node auto-provisioning?

The cluster autoscaler grows and shrinks existing node pools. Every pool has a fixed VM size, so you must design the pools in advance. Node auto-provisioning (NAP), built on Karpenter, looks at pending pods and creates nodes of whatever VM size fits best, without predefined pools.

Aspect Cluster autoscaler Node auto-provisioning
Unit of scaling Node pool (VM scale set) Individual nodes
VM size choice Fixed per pool Chosen per workload need
Config Min/max per pool NodePool and AKSNodeClass resources
Consolidation Removes idle nodes Actively repacks and replaces nodes to cut cost
Used by default in AKS Standard AKS Automatic

NAP is attractive for mixed or bursty workloads. The classic autoscaler remains simpler and more predictable when you already run well-sized, purpose-built pools.

Take quiz
What does node auto-provisioning choose on its own?
The Kubernetes version
The VM size that fits pending pods
The container registry
The ingress class
Which technology is NAP in AKS built on?
KEDA
Flux
Karpenter
Gatekeeper

41. How does the AKS control plane communicate with nodes?

Traffic flows in both directions, using different paths. Kubelets and kube-proxy initiate outbound connections to the API server. The API server needs to reach kubelets for kubectl logs, exec and port-forward, so AKS uses a reverse tunnel.

The konnectivity-agent pods in kube-system open an outbound, authenticated tunnel to the managed control plane. The API server sends kubelet requests through it, so no inbound firewall holes are needed on your nodes.

flowchart LR
  K["Kubelet on node"] -->|outbound 443| API["API server"]
  AG["konnectivity-agent pod"] -->|outbound tunnel| CP["Control plane konnectivity-server"]
  API --> CP
  CP -->|logs, exec, port-forward| AG
  AG --> K

This is why locked-down clusters must allow outbound access to the API server FQDN and required Microsoft endpoints. With API Server VNet Integration, nodes reach the API server directly over the VNet instead.

Take quiz
What lets the API server reach kubelets for kubectl exec without inbound rules?
An open SSH port on every node
A public IP on each pod
An outbound konnectivity tunnel from the nodes
A VPN built into etcd
Who initiates the connection from kubelet to the API server?
The API server, inbound to the node
The ingress controller
Azure Monitor
The kubelet, outbound from the node

42. Explain the execution flow of an AKS node pool upgrade?

AKS upgrades a node pool in a rolling fashion so pods keep running. The speed and capacity are controlled by max surge, which defaults to 10% of the pool for new pools (33% is a common production setting).

flowchart TD
  A["Start pool upgrade"] --> B["Create surge node with new version"]
  B --> C["Cordon an old node"]
  C --> D["Drain: evict pods respecting PDBs"]
  D --> E["Replace or reimage old node"]
  E --> F{More old nodes?}
  F -- Yes --> C
  F -- No --> G["Remove surge nodes, upgrade complete"]
  1. Extra surge nodes are created first.
  2. A node is cordoned so nothing new is scheduled on it.
  3. It is drained; evictions respect PodDisruptionBudgets and termination grace periods.
  4. The node is replaced and rejoins, and the loop repeats.

If a PodDisruptionBudget cannot be satisfied within the drain timeout (30 minutes by default), the upgrade fails. Use --max-surge, --drain-timeout and --node-soak-duration to tune it.

Take quiz
What does max surge control during a node pool upgrade?
How many pods a node can run
The size of the control plane
The number of namespaces
How many extra nodes can be added during the upgrade
What can cause a node drain to stall and the upgrade to fail?
A PodDisruptionBudget that cannot be satisfied
A missing ConfigMap
An unused Service
A public IP quota on ingress

43. What happens when an AKS node becomes NotReady?

When the kubelet stops reporting, the node controller marks the node Unknown after roughly 40 seconds and adds the node.kubernetes.io/unreachable and not-ready taints.

Pods carry a default toleration for these taints of 300 seconds. After five minutes without recovery they are evicted, and their Deployments or ReplicaSets create replacements on healthy nodes. Services stop routing to endpoints whose readiness fails.

Stateful pods are slower to recover because Azure Disks must be detached from the dead node and attached to another one.

AKS node auto-repair also watches for this. If a node stays NotReady for around 10 minutes, it escalates through reboot, reimage and redeploy. Check kubectl describe node for conditions such as memory pressure, disk pressure or a kubelet crash.

Take quiz
How long do pods tolerate unreachable or not-ready taints by default?
300 seconds
5 seconds
24 hours
They are evicted immediately
What does AKS auto-repair do for a node stuck in NotReady?
Deletes the whole cluster
Escalates through reboot, reimage and redeploy
Switches the CNI plugin
Re-creates the control plane

44. How do you troubleshoot pods stuck in Pending state?

Pending means the scheduler has not placed the pod yet. Start with the events, they almost always give the reason.

kubectl describe pod <pod> -n <ns>
kubectl get events -n <ns> --sort-by=.lastTimestamp

Look for the FailedScheduling message and map it to a cause:

Message Likely cause Fix
Insufficient cpu / memory Requests exceed free capacity Add nodes, lower requests, check autoscaler max
untolerated taint Pool is tainted Add toleration or use another pool
didn't match node selector / affinity Wrong label or selector Correct labels or the selector
unbound PersistentVolumeClaim PVC cannot provision Check StorageClass and zone
Too many pods Node max-pods reached Add nodes or raise max-pods on a new pool

If the cluster autoscaler should have reacted, check its status ConfigMap and logs, and confirm the pod fits a node of the pool's VM size at all.

Take quiz
Which command gives the scheduling reason for a Pending pod?
kubectl top node
kubectl describe pod
kubectl cordon
kubectl rollout undo
A message about an untolerated taint means:
The image tag is wrong
DNS is down
The pod lacks a toleration for a tainted node
The PVC is too small

45. How do you troubleshoot ImagePullBackOff when pulling from ACR?

Run kubectl describe pod and read the pull error. It usually points to one of four things: authorization, a wrong image reference, networking, or registry throttling.

  1. Image name: confirm the registry, repository and tag exist (az acr repository show-tags).
  2. Auth: run az aks check-acr -g rg -n aks --acr myregistry.azurecr.io, and confirm the kubelet identity holds AcrPull.
  3. Network: with ACR firewall, private endpoint or UDR egress, check DNS resolution and rules from the node subnet.
  4. Throttling: pulls from Docker Hub can hit rate limits, so import the image into ACR.

A 401 or 403 points to permissions, manifest unknown to a missing tag, and a timeout to the network path. Also check for an architecture mismatch, such as an arm64-only image on amd64 nodes.

Take quiz
Which command validates AKS-to-ACR connectivity and permissions?
az acr purge
kubectl get acr
az aks check-acr
az aks nodepool upgrade
A 'manifest unknown' pull error usually means:
The node is out of memory
CoreDNS crashed
The pod lacks a toleration
The image tag does not exist in the repository

46. How can you optimize AKS costs?

Nodes are the biggest line item, so most savings come from running fewer, better-packed and cheaper nodes.

Lever How it saves
Right-size requests and limits Stops over-reserving CPU and memory; the VPA add-on can recommend values
Autoscaling Cluster autoscaler or NAP removes idle nodes; HPA and KEDA scale pods
Spot pools Deep discounts for batch and stateless workloads
Reservations / savings plans Commitment discounts on steady baseline nodes
Stop dev clusters az aks stop pauses compute outside work hours
Efficient SKUs ARM-based VMs and Azure Hybrid Benefit for Windows nodes
Right tier Free tier for dev, Standard for production only

Enable the AKS cost analysis add-on to attribute spend to namespaces and workloads, so you know where to look first. Review it monthly, because new deployments quietly change the picture.

Take quiz
Which command pauses compute for a non-production cluster?
az aks delete --soft
kubectl suspend cluster
az aks tier pause
az aks stop
What does the AKS cost analysis add-on help with?
Attributing spend to namespaces and workloads
Encrypting etcd
Upgrading node images
Creating ingress rules

47. How do you design a multi-region AKS architecture?

Run an independent cluster in each region (typically a paired region), and put a global entry point in front. Keep clusters stateless and rebuildable so any one of them can be lost.

flowchart TD
  U[Users] --> FD["Azure Front Door"]
  FD --> A["AKS East US"]
  FD --> B["AKS West US"]
  A --> DB[(Replicated data tier)]
  B --> DB
  ACR["ACR geo-replicated"] --> A
  ACR --> B
  G["Git repo with GitOps"] --> A
  G --> B
  • Traffic: Azure Front Door or Traffic Manager with health probes for failover.
  • Images: ACR Premium with geo-replication so each region pulls locally.
  • Config parity: GitOps so both clusters apply the same manifests.
  • Data: Cosmos DB multi-region writes or SQL failover groups, since Kubernetes does not replicate state.
  • Fleet operations: Azure Kubernetes Fleet Manager to coordinate upgrades and propagate workloads.

Decide between active-active (lower failover time, higher cost) and active-passive (cheaper, slower to take over) based on your RTO and RPO.

Take quiz
Which service steers users between regional clusters with health probes?
Azure Front Door
Azure Bastion
Azure Key Vault
Azure DevTest Labs
What keeps container images available locally in each region?
Node image upgrades
ACR geo-replication
Pod topology spread
kubelet garbage collection

48. Explain how GitOps with Flux works in AKS?

With GitOps, a Git repository is the source of truth and an in-cluster agent continuously pulls and applies it. AKS offers Flux v2 as a managed extension (microsoft.flux).

flowchart LR
  Dev["Developer pushes to Git"] --> Repo[(Git repo)]
  Repo --> SC["source-controller fetches"]
  SC --> KC["kustomize-controller / helm-controller"]
  KC --> K["Cluster resources applied"]
  K --> D{Drift from Git?}
  D -- Yes --> KC
az k8s-configuration flux create -g rg-demo -c aks-demo -t managedClusters \
  -n platform --namespace flux-system --scope cluster \
  --url https://github.com/contoso/platform --branch main \
  --kustomization name=apps path=./apps prune=true

The source-controller fetches the repo, then the kustomize or helm controller renders and applies it on every interval. Because it reconciles, manual changes made with kubectl are reverted, and prune=true deletes objects removed from Git.

Take quiz
Who is the source of truth in a GitOps workflow?
The kubectl history
The Git repository
The node's local disk
The Azure portal blade
What happens to a manual kubectl change that differs from Git?
It is committed back to Git automatically
It is kept forever
Flux reverts it during reconciliation
Flux deletes the cluster

49. How do you achieve zero-downtime deployments in AKS?

It takes several settings working together, not just a rolling update.

  1. Run at least 2-3 replicas spread across nodes or zones.
  2. Use a rolling strategy with maxUnavailable: 0 so capacity never drops.
  3. Add a readiness probe so traffic reaches only pods that are ready.
  4. Add a preStop sleep and a sufficient terminationGracePeriodSeconds so in-flight requests finish and endpoints are removed first.
  5. Create a PodDisruptionBudget so node drains and upgrades never take down too many pods.
strategy:
  type: RollingUpdate
  rollingUpdate: {maxSurge: 1, maxUnavailable: 0}
template:
  spec:
    terminationGracePeriodSeconds: 45
    containers:
    - name: api
      readinessProbe:
        httpGet: {path: /ready, port: 8080}
      lifecycle:
        preStop:
          exec: {command: ["sleep", "10"]}

Without the preStop delay, a pod can receive requests after it starts shutting down, because endpoint removal and SIGTERM race each other.

Take quiz
Why set maxUnavailable to 0 in a rolling update?
To disable readiness probes
To stop HPA from scaling
So serving capacity never drops during rollout
To run a single replica
What does a PodDisruptionBudget protect against?
Image pull failures
Expired TLS certificates
Slow DNS lookups
Too many pods being evicted at once during drains

50. How do you troubleshoot DNS resolution failures in AKS?

Work from the pod outward: confirm the symptom, then check CoreDNS, then the upstream path.

  1. Test from a pod: kubectl run dnstest --rm -it --image=registry.k8s.io/e2e-test-images/agnhost:2.39 -- nslookup kubernetes.default.
  2. Check CoreDNS: kubectl get pods -n kube-system -l k8s-app=kube-dns, then read its logs for SERVFAIL or timeouts.
  3. If only external names fail, look at the upstream: custom VNet DNS servers must forward to Azure DNS 168.63.129.16.
  4. Check NSGs, firewalls or UDRs that block UDP/TCP 53 between nodes and the DNS servers.
  5. Review /etc/resolv.conf in the pod; ndots:5 makes short names trigger many extra lookups.

Intermittent failures under load often mean the per-VM limit on Azure DNS (about 500 queries per second) is being hit. Deploying NodeLocal DNSCache and scaling CoreDNS reduces that pressure. Custom rules belong in the coredns-custom ConfigMap, not in edits to the managed CoreDNS config.

Take quiz
Which IP should a custom VNet DNS server forward unresolved queries to?
10.244.0.1
127.0.0.53
8.8.8.8 only
168.63.129.16
Where should you add custom CoreDNS rules in AKS?
The coredns-custom ConfigMap
Directly on the node's resolv.conf
The Azure Policy add-on
A NodePort Service
«
»

Comments & Discussions