BigData / Apache Airflow Interview Questions
1. What is Apache Airflow?
Apache Airflow is an open-source platform for programmatically authoring, scheduling, and monitoring workflows. Originally created at Airbnb in 2014 and donated to the Apache Software Foundation, Airflow lets you define workflows as Directed Acyclic Graphs (DAGs) written in Python....
2. What is a DAG in Apache Airflow?
A DAG (Directed Acyclic Graph) is the core concept in Airflow. It is a collection of tasks organized with dependencies and relationships that define how they should run....
3. What are Operators in Apache Airflow?
Operators are the building blocks of a DAG. Each operator represents a single, idempotent unit of work — a task....
4. What is the Airflow Scheduler?
The Airflow Scheduler is a core component that monitors DAGs and triggers task instances when their dependencies are met and the schedule interval fires. It continuously parses DAG files, checks the DAG schedule and any data interval conditions, and submits...
5. What are the main components of Apache Airflow?
Apache Airflow consists of five main components: Component Role Webserver Provides the UI to monitor, trigger, and debug DAGs Scheduler Triggers task instances per schedule and dependency rules Executor Determines how and where tasks run (locally, on Celery workers, on...
6. What is an Executor in Airflow and what types are available?
The Executor defines how tasks are run. Airflow supports several executor types suited for different scale and infrastructure needs: SequentialExecutor — runs one task at a time; only for development/testing....
7. What is an Airflow Connection?
A Connection stores credentials and endpoint information for external systems such as databases, cloud providers, and APIs. Rather than hardcoding credentials in DAG code, you store them in Airflow's metadata database (or a secrets backend) and reference them by a...
8. What is an Airflow Variable?
Variables are key-value pairs stored in Airflow's metadata database. They provide a way to pass configuration or runtime values into DAGs without hardcoding them....
9. What is XCom in Airflow and how is it used?
XCom (cross-communication) is a mechanism that lets tasks exchange small amounts of data. A task can push a value into XCom and a downstream task can pull it....
10. What are Hooks in Apache Airflow?
Hooks are interfaces to external platforms and databases. They abstract connection handling, authentication, and API calls so operators and tasks don't need to manage those details directly....
11. What is the difference between a DAG Run and a Task Instance in Airflow?
A DAG Run is an instantiation of the whole DAG for a specific data interval (logical date). It tracks the overall execution state (running, success, failed) for that interval....
12. What are Sensors in Apache Airflow?
Sensors are a special type of operator that wait for a certain condition to become true before proceeding. They poke the condition at a configurable interval ( poke_interval ) and either block (poke mode) or release the worker slot between...
13. What is catchup in Airflow and how does it work?
Catchup is a DAG parameter that controls whether Airflow should backfill all DAG Runs from the start_date up to the current date when a DAG is first enabled or its schedule is changed. When catchup=True (default), Airflow will create a...
14. What is backfilling in Apache Airflow?
Backfilling is the process of running a DAG for historical date ranges that were not previously executed. You can trigger a backfill from the CLI: airflow dags backfill --start-date 2024-01-01 --end-date 2024-03-31 my_dag This creates DAG Runs for every schedule...
15. What is the TaskFlow API in Airflow?
Introduced in Airflow 2.0, the TaskFlow API is a decorator-based approach to writing DAGs and tasks that reduces boilerplate. Instead of instantiating operator objects explicitly, you decorate plain Python functions with @task and define the DAG with @dag ....
16. What is the difference between schedule_interval and timetable in Airflow?
schedule_interval (deprecated from Airflow 2.4) accepted cron strings or timedelta objects to define how frequently a DAG should run. schedule (the replacement) accepts the same values but also supports Timetables ....
17. What is a SubDAG and why is it generally discouraged?
A SubDAG is a pattern where one DAG embeds another DAG as a single task using the SubDagOperator . It was used to group related tasks and visualize them as a unit in the UI....
18. What is a TaskGroup in Airflow?
TaskGroup nests related tasks in the Airflow UI without a SubDAG. It is a layout and naming helper, not a separate DAG execution unit.
19. What is branching in Airflow and how is BranchPythonOperator used?
BranchPythonOperator returns the next task_id (or list) to run. Downstream tasks not returned are skipped, so skipped/none_skipped rules matter.
20. What are trigger rules in Airflow?
Trigger rules define when a task should be triggered relative to its upstream tasks. The default is all_success , but many alternatives exist: Rule Description all_success All upstream tasks succeeded (default) all_failed All upstream tasks failed all_done All upstream tasks...
21. What is the Airflow metadata database and what does it store?
The metadata database (metastore) is a relational database (PostgreSQL or MySQL recommended for production; SQLite for local testing) that is the central state store for Airflow. It persists: DAG definitions and schedules DAG Runs and their states Task Instances and...
22. How does the CeleryExecutor work in Airflow?
The CeleryExecutor uses Celery (a distributed task queue) to distribute Airflow task execution across multiple worker nodes. The workflow is: The Scheduler submits tasks to a message broker (Redis or RabbitMQ)....
23. What is the KubernetesExecutor and what are its benefits?
The KubernetesExecutor launches a dedicated Kubernetes pod for every task instance. When a task is scheduled, the executor calls the Kubernetes API to create a pod; the pod runs the task and then terminates....
24. What are Pools in Apache Airflow?
Pools are a mechanism to limit the number of concurrently running tasks that use shared resources (e.g., database connections, API rate limits). You define a pool with a maximum number of slots, and assign tasks to it....
25. What are Airflow Providers?
Providers (formerly contrib) are installable packages that extend Airflow with operators, hooks, sensors, and connections for third-party services. They are published separately from the core Airflow package, so you install only what you need.
26. What is dynamic task mapping in Airflow?
Dynamic task mapping, introduced in Airflow 2.3, allows you to create a variable number of task instances at runtime based on data rather than defining a fixed set of tasks at DAG parse time. from airflow.decorators import dag, task @dag...
27. What is the difference between depends_on_past and wait_for_downstream in Airflow?
Both parameters control inter-run dependencies for a task: depends_on_past=True : The task will only run if the same task in the previous DAG Run succeeded. Useful for tasks that process sequential data....
28. What is the Airflow Web UI and what can you do with it?
The Airflow Web UI (powered by Flask and the FAB — Flask-AppBuilder security framework) is a browser-based dashboard for managing and monitoring pipelines. Key features include: DAG list view — see all DAGs, toggle pause/unpause, check recent run states....
29. What are Airflow task states and what do they mean?
Each Task Instance in Airflow goes through a lifecycle represented by states: State Meaning none Not yet scheduled scheduled Dependency met, waiting for executor queued Sent to executor, waiting for a worker slot running Currently executing success Completed without error...
30. What are retries and retry_delay in Airflow tasks?
Task-level retry parameters control behavior after a task failure: retries — number of times Airflow will retry the task before marking it as failed . Default is 0....
31. What is a Deferrable Operator (async operator) in Airflow?
Deferrable Operators (introduced in Airflow 2.2) allow tasks to suspend themselves, release the worker slot, and wait for an external event via a lightweight Trigger component. This is more efficient than sensors in poke mode because no worker thread is...
32. What are Airflow Plugins?
Plugins are a way to extend Airflow with custom operators, hooks, sensors, macros, UI views, and Flask blueprints without forking the core codebase. You place a Python module in the plugins folder (configured as plugins_folder ) and Airflow automatically discovers...
33. How does Airflow handle templating and macros?
Airflow supports Jinja2 templating inside operator parameters that are listed in template_fields . You can inject runtime values such as execution date, data interval, and custom macros....
34. What is idempotency in the context of Airflow tasks?
An idempotent task produces the same result regardless of how many times it is executed for the same input. In Airflow, tasks are often retried or re-run (via clear), so designing them to be idempotent prevents duplicate data or inconsistent...
35. What are best practices for writing efficient Airflow DAGs?
Key best practices for production-quality DAGs: Keep DAG files lightweight — avoid heavy imports or database calls at parse time; the scheduler parses DAG files continuously. Use top-level constants only — don't call APIs or read files at module level;...
36. What is the ExternalTaskSensor in Airflow?
The ExternalTaskSensor waits for a task in another DAG (or the DAG itself) to reach a target state (default: success) before proceeding. This enables cross-DAG dependencies without tight coupling.
37. What is the KubernetesPodOperator in Airflow?
The KubernetesPodOperator runs any command inside a Docker container launched as a Kubernetes pod. It is independent of the executor type — even with CeleryExecutor you can use KubernetesPodOperator to run specific tasks in isolated pods.
38. What are SLAs in Apache Airflow and how are they configured?
SLAs (Service Level Agreements) in Airflow allow you to define the maximum time by which a task or DAG should complete after the scheduled execution date. If the deadline is missed, Airflow sends an email alert and logs an SLA...
39. How does Airflow handle task concurrency and parallelism?
Airflow provides several levels of concurrency control: parallelism (airflow.cfg) — global maximum number of tasks running across all DAGs. dag_concurrency / max_active_tasks (per DAG) — max tasks running within a single DAG at once....
40. What is an Airflow Dataset and how does data-driven scheduling work?
Datasets (introduced in Airflow 2.4) are logical references to data assets identified by a URI. A DAG can produce a Dataset via an outlet, and another DAG can be scheduled to run automatically when that Dataset is updated....
41. What is the difference between Airflow and Apache Spark?
Airflow and Spark serve different purposes and are frequently used together: Aspect Apache Airflow Apache Spark Purpose Workflow orchestration — define, schedule, and monitor pipelines Distributed data processing — transform and analyze large datasets in memory Data Does not process...
42. How do you deploy Apache Airflow using Docker Compose?
The official Airflow Docker Compose setup ( docker-compose.yaml ) is the quickest way to run a production-like local environment with Postgres, Redis, and CeleryExecutor. Steps: Download the official compose file: curl -LfO 'https://airflow.apache.org/docs/apache-airflow/stable/docker-compose.yam
43. What is Airflow on Kubernetes (KEDA) autoscaling?
When running CeleryExecutor on Kubernetes, KEDA (Kubernetes-based Event Driven Autoscaler) can automatically scale Celery worker pods based on the number of tasks in the queue. This removes the need for static worker counts and makes the cluster cost-efficient....
44. What is the SparkSubmitOperator in Airflow?
The SparkSubmitOperator submits a Spark application to a Spark cluster (Standalone, YARN, or Kubernetes). It wraps the spark-submit command and reads Spark connection details from an Airflow Connection ( conn_id='spark_default' ).
45. What is Managed Airflow (MWAA) on AWS?
Amazon Managed Workflows for Apache Airflow (MWAA) is a fully managed Airflow service on AWS. It handles infrastructure provisioning, scaling, patching, and HA configuration....
46. How does Airflow handle secrets management?
Airflow supports a pluggable Secrets Backend to retrieve Connections and Variables from external secret stores rather than the metadata DB. Supported backends include AWS Secrets Manager, HashiCorp Vault, GCP Secret Manager, and Azure Key Vault....
47. What is the difference between PythonOperator and PythonVirtualenvOperator?
Both operators run Python callables, but differ in execution environment: PythonOperator — runs the callable in the same Python environment as the Airflow worker. All packages available to the worker are accessible....
48. What is the Grid view in Airflow 2.x?
The Grid view (introduced in Airflow 2.3 as a replacement for the Tree view) shows a matrix of DAG Runs on the X-axis and tasks on the Y-axis. Each cell represents a task instance colored by state (green=success, red=failed, yellow=running,...
49. What are common Airflow anti-patterns to avoid?
Common pitfalls that hurt reliability and performance: Top-level DB calls — calling Variable.get() or Connection.get_connection_from_secrets() at parse time stresses the scheduler. Non-idempotent tasks — retries create duplicate data.
50. What is Airflow 2 vs Airflow 1 — key differences?
Airflow 2.0 (released December 2020) was a major overhaul. Key changes: Feature Airflow 1.x Airflow 2.x Scheduler HA Single scheduler only Multiple schedulers supported (HA) TaskFlow API Not available @dag / @task decorators Provider packages All in monolithic airflow package...