BigData / Apache Airflow Interview Questions
What is an Airflow Dataset and how does data-driven scheduling work?
Datasets (introduced in Airflow 2.4) are logical references to data assets identified by a URI. A DAG can produce a Dataset via an outlet, and another DAG can be scheduled to run automatically when that Dataset is updated.
from airflow.datasets import Dataset my_dataset = Dataset('s3://my-bucket/output/daily.csv') # Producer DAG @dag(schedule='@daily', ...) def producer(): @task(outlets=[my_dataset]) def write_data(): ... # write to S3 # Consumer DAG â triggered whenever my_dataset is updated @dag(schedule=[my_dataset], ...) def consumer(): ...
This replaces fragile time-based scheduling with event-driven, data-dependency-aware scheduling.
More Related questions...