Cloud / AWS Glue Interview questions
Last updated
1. What is AWS Glue?
AWS Glue is a fully managed, serverless data integration service that discovers, prepares, moves, and combines data from multiple sources for analytics, machine learning, and application development.
It centers on three pieces working together: the Data Catalog, which stores metadata about your datasets; crawlers, which scan data stores and populate that catalog; and ETL jobs, which run on a managed Apache Spark, Python shell, or Ray environment to transform and load data.
Because Glue is serverless, there is no infrastructure to provision or patch - AWS spins up compute (measured in DPUs) only while a crawler or job is running, and you pay for that usage rather than for idle servers.
Take quiz
a managed relational database
a container orchestration platform
a serverless data integration service
DPU usage while a job or crawler actually runs
a flat monthly subscription
a fixed number of reserved EC2 instances
2. What are the core components of AWS Glue?
AWS Glue is built from a handful of components that cover metadata management, data discovery, and data processing.
- Data Catalog - the central, persistent metadata store
- Crawlers - connect to data stores and populate catalog tables automatically
- Classifiers - determine the schema and format a crawler assigns to a file
- ETL jobs - Spark, Python shell, or Ray runtimes that transform data
- Triggers & workflows - orchestrate crawlers and jobs together
- Connections - network and credential details for reaching data stores in a VPC
Glue Studio, DataBrew, and Data Quality are higher-level tools layered on top of this same foundation rather than separate systems.
Take quiz
Classifier
Trigger
Crawler
Classifier and trigger
Job and connection
Crawler and workflow
3. What is the AWS Glue Data Catalog?
The Data Catalog is AWS Glue's persistent metadata repository - it stores table definitions (schema, location, format, partitions) for datasets that live in Amazon S3, JDBC databases, Redshift, DynamoDB, and other sources, without storing the data itself.
Because the catalog uses the same metadata store that Amazon Athena, Redshift Spectrum, and EMR can query, a table cataloged once by Glue becomes immediately queryable from any of those services without redefining its schema.
The catalog is priced separately from ETL compute: the first million stored objects and first million requests each month are free, with a small per-unit charge beyond that, which makes it cheap to catalog even large numbers of tables.
Take quiz
IAM policies
Table metadata such as schema and location
The actual data files
AWS CloudTrail
Amazon SNS
Amazon Athena
4. What is a Glue crawler?
A crawler is a Glue component that connects to a source data store - such as an S3 bucket, JDBC database, or DynamoDB table - infers the schema of the data it finds, and writes or updates the corresponding table definitions in the Data Catalog.
Internally, a crawler samples files, runs them through a prioritized list of classifiers to detect format (CSV, JSON, Parquet, and so on) and schema, and then groups similar files into partitions of a single table rather than creating a new table per file.
Crawlers can run on demand, on a schedule, or be started by a trigger, and billing follows the same DPU-hour model as ETL jobs, with a 10-minute minimum per run.
Take quiz
Transform data in place
Encrypt S3 buckets
Infer schema and populate the Data Catalog
10 minutes
1 hour
1 minute
5. What are the types of AWS Glue jobs?
AWS Glue supports three job types, each suited to a different workload.
- Spark ETL jobs - run on a managed Apache Spark environment (Spark 3.5.4 on Glue 5.0) for large-scale batch or streaming transformations.
- Python shell jobs - run plain Python scripts without a Spark cluster, suited to lightweight tasks on datasets up to roughly 10 GB.
- Ray jobs - run on a managed Ray cluster for distributed Python workloads such as data preprocessing for machine learning.
Spark streaming jobs are a variant of the Spark job type that process data continuously from sources like Kinesis or Kafka instead of running to completion.
Take quiz
Python shell
Ray
Spark ETL
Relational joins expressed in SQL
Distributed Python workloads like ML preprocessing
Scheduling triggers
6. What is the purpose of a Glue classifier?
A classifier examines a sample of data and determines its schema and format so a crawler knows how to catalog it - for example, recognizing that a file is CSV with a header row, JSON, Parquet, Avro, or a specific log format.
Glue ships with built-in classifiers for common formats and runs them in a fixed order of priority; the first classifier that returns a confident match wins for that data store.
When built-in classifiers cannot correctly parse a custom or unusual format - a proprietary delimited log format, for instance - you can write a custom classifier using a grok pattern, XML row tag, or JSON path expression, and place it ahead of the built-in ones in the crawler's classifier list.
Take quiz
Encrypt data at rest
Determine a data store's schema and format
Schedule ETL jobs
You want to change IAM permissions
The data is a standard CSV file
The data is in a proprietary or unusual format built-in classifiers cannot parse
7. Define a DynamicFrame in AWS Glue?
A DynamicFrame is Glue's own distributed data structure, similar to a Spark DataFrame, but designed for semi-structured data where the schema is not fully known or consistent up front.
Unlike a DataFrame, a DynamicFrame can represent a column with multiple possible types across records (a "choice" type) without failing, and it tracks schema inconsistencies so you can resolve them explicitly later with transforms like resolveChoice.
You can convert freely between the two: dynamic_frame.toDF() produces a Spark DataFrame, and DynamicFrame.fromDF(df, glue_context, "name") converts it back, which lets you mix Glue-specific transforms with native Spark or Spark SQL in the same script.
datasource = glueContext.create_dynamic_frame.from_catalog( database="sales_db", table_name="orders" ) df = datasource.toDF() # convert to Spark DataFrame
Take quiz
Requires a fixed schema up front
Only works with SQL queries
Can represent inconsistent or ambiguous column types safely
.toDF()
.toPandas()
.toRDD()
8. Describe an AWS Glue trigger?
A trigger starts one or more Glue jobs or crawlers, either immediately or based on a condition, letting you orchestrate a pipeline without writing custom scheduling code.
Glue supports three trigger types: scheduled triggers that fire on a cron-style expression, on-demand triggers that a user or API call starts manually, and conditional (event-based) triggers that fire when one or more specified jobs or crawlers complete successfully.
Conditional triggers are what make it possible to chain steps - for instance, starting a transform job only after its upstream crawler finishes refreshing the catalog - without any code outside the Glue console or API.
Take quiz
Scheduled trigger
Conditional trigger
On-demand trigger
On-demand trigger
Conditional trigger
Scheduled trigger
9. What is an AWS Glue workflow?
A workflow is a container that groups a set of related crawlers, jobs, and triggers into a single, visual, trackable unit - useful once a pipeline grows beyond a single job with one trigger.
Instead of managing each trigger independently, you design the whole sequence in the workflow's graph, and Glue exposes shared run properties that one job in the workflow can set and a downstream job can read, making it easy to pass metadata like a processed row count or a file path along the chain.
Each execution of a workflow gets its own run ID, and the console's visual graph shows the real-time state of every crawler and job in that run, which is far easier to monitor than checking each component separately.
flowchart LR A["Scheduled trigger"] --> B["Crawler: raw zone"] B --> C["Conditional trigger"] C --> D["ETL job: transform"] D --> E["Conditional trigger"] E --> F["Crawler: curated zone"]
Take quiz
Store raw data
Group and orchestrate related crawlers, jobs, and triggers
Replace the Data Catalog
Hardcoded S3 paths only
IAM policies
Workflow run properties
10. List the worker types supported by AWS Glue Spark jobs?
| Worker type | DPU | vCPU / Memory |
| G.1X | 1 | 4 vCPU / 16 GB |
| G.2X | 2 | 8 vCPU / 32 GB |
| G.4X | 4 | 16 vCPU / 64 GB |
| G.8X | 8 | 32 vCPU / 128 GB |
| G.025X | 0.25 | 2 vCPU / 4 GB |
G.1X and G.2X cover most transforms, joins, and queries; G.4X and G.8X target the heaviest aggregations and joins; and G.025X exists only for low-volume streaming jobs on Glue 3.0+. Ray jobs use a separate worker type, Z.2X, rather than any of these.
Take quiz
G.1X
G.8X
G.025X
4 vCPU and 16 GB memory
8 vCPU and 32 GB memory
2 vCPU and 8 GB memory
11. What is a Glue connection?
A connection is a Data Catalog object that stores the properties Glue needs to reach a data store that is not plain S3 - things like a JDBC URL, VPC subnet and security group, and credentials (often via a Secrets Manager reference).
Because many source databases (RDS, Redshift, on-prem via VPN, self-managed databases on EC2) live inside a VPC, a connection tells Glue which subnet and security group to launch elastic network interfaces in so its Spark or Python shell environment can route traffic to that database.
Connections are reusable: once defined, the same connection can be attached to any number of crawlers or jobs that need to reach that data store, so credentials and network details are maintained in one place instead of being duplicated in every script.
Take quiz
Network and credential details for reaching a data store
Quiz feedback text
ETL transformation logic
In a public S3 bucket
Inside a VPC
In the Data Catalog itself
12. What is the purpose of job bookmarks in AWS Glue?
Job bookmarks let a Glue job track which data it has already processed across runs, so a scheduled or repeated job run only picks up new or changed data instead of reprocessing everything from scratch.
For S3 sources, Glue tracks bookmarks by file path, size, and last-modified timestamp; for JDBC sources, it tracks a monotonically increasing column you specify, since JDBC has no equivalent of file metadata.
Bookmarks are enabled per job (not on by default for every job type) and are stored internally by Glue, so you do not manage the tracking state yourself - though you can reset a bookmark through the console or API if you need to reprocess historical data.
Take quiz
Always reprocess all data from scratch
Skip data that was already processed in prior runs
Delete the Data Catalog
The crawler's classifier
File last-modified timestamps
A monotonically increasing column you specify
13. What is AWS Glue Studio?
Glue Studio is a graphical interface for authoring, running, and monitoring Glue ETL jobs without hand-writing all the Spark code, alongside a code editor and notebook interface for teams that prefer to work in script form.
In the visual editor, you drag in a source node, a chain of transform nodes (join, filter, apply mapping, and so on), and a target node, and Glue Studio generates the underlying PySpark script automatically, which you can still open and edit directly.
Glue Studio also includes a job monitoring dashboard showing run history, DPU-hours consumed, and any errors, which makes it a common entry point even for teams that ultimately maintain their ETL logic as hand-written scripts.
Take quiz
A static PDF report
IAM policies
The underlying PySpark script automatically
Run history and DPU-hours consumed
Source code repositories
EC2 instance health
14. What is AWS Glue DataBrew?
DataBrew is a visual data-preparation tool for cleaning and normalizing data without writing code, aimed at analysts and data scientists rather than engineers writing Spark scripts.
You point DataBrew at a dataset, and it profiles the data and suggests over 250 pre-built transformations - handling missing values, standardizing formats, removing duplicates - which you apply through a spreadsheet-like interface and combine into a reusable "recipe."
DataBrew is billed separately from core Glue ETL: interactive sessions are billed per 30-minute session and recipe jobs are billed per node-hour, so it sits alongside Glue jobs as a distinct, complementary service rather than a mode of the Spark ETL engine.
Take quiz
Analysts who want to clean data visually without code
Network administrators
Spark engineers writing custom code
Workflow
Recipe
Classifier
15. Define AWS Glue Data Quality?
Glue Data Quality is a managed capability, built on the open-source Deequ framework, that measures and monitors the quality of data in the Data Catalog or inside an ETL job, without you having to install or manage anything.
You define checks using Data Quality Definition Language (DQDL), a simple domain-specific syntax for rules like completeness, uniqueness, or value-range checks, and Glue can also auto-recommend a starting ruleset by profiling your data.
Running a ruleset produces a data quality score along with the specific records that failed each rule, so you can quarantine or fix bad data before it reaches downstream consumers, and you can invoke this from Glue Studio, notebooks, interactive sessions, or the API.
Take quiz
Plain SQL only
Data Quality Definition Language (DQDL)
YAML
Apache Airflow
Terraform
Deequ framework
16. How do you use the AWS Glue Schema Registry?
The Glue Schema Registry is a serverless store for Avro, JSON Schema, and Protobuf schemas used by streaming applications - typically Kinesis Data Streams or Kafka producers and consumers - to enforce a shared, versioned contract on message format.
A producer registers (or references) a schema, serializes messages with a small schema-id header instead of embedding the full schema in every message, and a consumer looks up that schema-id in the registry to deserialize correctly, which shrinks message size and prevents mismatched formats from silently breaking a pipeline.
Compatibility modes (backward, forward, full, or none) control whether a new schema version is allowed to be registered at all - for example, backward compatibility blocks a change that would break consumers still running the previous schema version - and Glue ETL jobs can consume registry-managed streams directly through the Kafka or Kinesis connector.
Take quiz
Share IAM roles
Share disk storage
Agree on a versioned message schema
Backward
Full
None
17. How do you apply a custom classifier in a Glue crawler?
You create the custom classifier first - as a grok pattern for text logs, an XML row-tag classifier, or a JSON path expression - giving it a name, then attach it to a crawler in the crawler's configuration alongside (or ahead of) the built-in classifiers.
Order matters: Glue runs classifiers in the sequence listed and stops at the first one that returns a match with sufficient certainty, so a custom classifier meant to override a built-in one (say, a delimited log format that looks superficially like CSV) needs to be placed before the built-in CSV classifier in that list.
If none of the listed classifiers - custom or built-in - reach sufficient certainty on a file, the crawler falls back to treating it as an unknown or unstructured format, so it is worth testing a new custom classifier against representative sample files before relying on it in production.
Take quiz
Placed before it in the list
Placed after the built-in classifier in the list
Deleted from the crawler
Fails the entire crawl
Marks the file's format as unknown
Deletes the file
18. What is the purpose of the resolveChoice transform in AWS Glue?
resolveChoice resolves columns in a DynamicFrame that Glue flagged as having ambiguous or inconsistent types across records - a "choice" type, like a column that is sometimes an integer and sometimes a string - into a single, well-defined type.
You can resolve a choice several ways: cast to force a specific target type, make_struct to keep every observed type by wrapping the value in a struct, project to keep only one of the observed types and drop records of other types, or match_catalog to conform the DynamicFrame to the schema already registered for that table in the Data Catalog.
Leaving a choice type unresolved will cause a later toDF() conversion or a write to a strongly-typed sink to fail or produce a malformed schema, so resolveChoice is typically one of the first transforms applied right after reading messy source data.
resolved = dynamic_frame.resolveChoice( specs=[("customer_id", "cast:long")] )
Take quiz
Schedule a trigger
Resolve ambiguous or inconsistent column types in a DynamicFrame
Partition output files
A hardcoded Python dictionary
The source file's original schema
The schema already registered in the Data Catalog
19. How is a DynamicFrame different from a Spark DataFrame?
Both are distributed collections of records, but they serve different points in a Glue ETL pipeline and behave differently when a schema is not clean.
| DynamicFrame | Spark DataFrame |
| Tolerates inconsistent or ambiguous types per column (choice types) | Requires one fixed type per column |
| Ships Glue-specific transforms like resolveChoice, relationalize, unnest | Uses native Spark / Spark SQL transforms only |
| Created via GlueContext (from_catalog, from_options) | Created via SparkSession (read, sql) |
| Generally slower for pure Spark-native operations | Faster for joins and aggregations once schema is clean |
A common pattern is to read with a DynamicFrame to take advantage of schema flexibility and Glue's built-in transforms, resolve any ambiguity, convert to a DataFrame with toDF() for the performance-sensitive core logic, and convert back with DynamicFrame.fromDF() only if a Glue-specific sink needs one.
Take quiz
Neither structure
Spark DataFrame
DynamicFrame
.toDF()
.collect()
.toPandas()
20. What is the difference between AWS Glue and AWS Glue DataBrew?
Both sit in the same Glue family and share the Data Catalog, but they target different users and different stages of data preparation.
Glue (crawlers plus Spark, Python shell, or Ray jobs) is code-first: engineers write or generate scripts to build repeatable, often complex, production ETL pipelines that move and reshape data at scale. DataBrew is code-free: analysts click through a visual profiling and transformation interface to clean up a dataset - deduplicate rows, standardize date formats, handle nulls - without writing Spark code.
In practice, teams often use DataBrew early, during exploration and one-off cleanup, and use Glue ETL jobs for the pipeline that has to run reliably and repeatedly in production, sometimes even exporting a DataBrew recipe to reuse its transformation steps inside a Glue job later.
Take quiz
A code-free, analyst-facing cleanup tool
Only usable for streaming data
A replacement for the Data Catalog
Production-scale nightly pipelines
Exploration and one-off cleanup
IAM role management
21. What is the difference between a Glue crawler and a Glue job?
A crawler and a job operate on opposite sides of the same pipeline: a crawler reads metadata about your data and writes it into the Data Catalog, while a job reads and writes the actual data itself.
A crawler never touches the data's contents beyond sampling enough to infer schema and format; a job (Spark, Python shell, or Ray) processes every record - filtering, joining, aggregating, or reshaping it - and writes transformed output to a target.
A typical pipeline uses both in sequence: a crawler catalogs a raw S3 landing zone, a job reads from that cataloged table, transforms the data, and writes it to a curated location, and a second crawler then catalogs that curated output so it is queryable by Athena or Redshift Spectrum.
Take quiz
Transformed data files
Catalog metadata (table definitions)
A Step Functions state machine
Only reads metadata
Can only run on a schedule
Reads and writes the actual data records
22. Why do we use job bookmarks instead of manually tracking processed files?
You could track "what is already processed" yourself - writing processed file names or a max timestamp to DynamoDB or S3 and checking it at the start of every run - but that adds a whole extra piece of state you have to build, secure, and keep consistent with the actual job logic.
Job bookmarks push that responsibility into Glue itself: enabling bookmarks on a job is a single job parameter, and Glue transparently records what was read (by file path, size, and timestamp for S3, or by a specified column for JDBC) after each successful run, so the very next run skips already-processed input automatically.
This matters most for cost and correctness at scale: without bookmarks, a daily incremental job re-scans and reprocesses the entire dataset every single run, multiplying DPU-hours and risking duplicate rows in an append-only sink, whereas bookmarks keep each run's cost proportional to only the new data.
Take quiz
Only work for JDBC sources
Require you to build and maintain separate tracking infrastructure
Are built into Glue and require no separate state store
Re-scans and reprocesses the full dataset every run
Stops running after one day
Processes only the new rows each run
23. What is the difference between AWS Glue Standard and Flex execution class?
Both run the exact same Spark engine and produce identical results; the difference is purely about how quickly compute is provisioned and what that costs.
| Standard | Flex |
| Starts immediately with dedicated capacity | Starts using spare capacity, may be delayed |
| $0.44 per DPU-hour | About $0.29 per DPU-hour (roughly 34% cheaper) |
| Suited to time-sensitive jobs | Suited to batch jobs that can tolerate a delayed start |
| Supported for all job and worker combinations | Not supported for Python shell, streaming, or ML jobs |
Flex is set per job as an execution class, so migrating an existing non-time-critical batch job to Flex is usually just changing that one setting, with no code changes required.
Take quiz
A possible delayed start for a lower DPU-hour rate
Memory for CPU
Accuracy for speed
Spark ETL batch jobs
Python shell, streaming, and ML jobs
G.1X worker jobs
24. When should you use a Glue Python shell job instead of a Spark job?
Reach for a Python shell job when the workload does not need distributed processing at all - a script that calls an API, does light file manipulation, moves small files around, or runs a one-off transformation on a dataset up to roughly 10 GB.
Python shell jobs start faster and cost less than Spark jobs (billed down to 0.0625 DPU) because there is no cluster to provision, and they support plain Python plus common libraries (pandas, requests, boto3) without any Spark API to learn.
The tradeoff is that a Python shell job runs on a single node with no distributed execution, so it does not scale as data volume grows - once a job is doing real joins or aggregations across tens of gigabytes or more, moving to a Spark job with an appropriate worker type becomes the better choice.
Take quiz
Streaming Kafka ingestion
Lightweight scripts on small-to-medium datasets
Multi-terabyte distributed joins
1 full DPU
8 DPU
0.0625 DPU
25. How does AWS Glue integrate with AWS Lake Formation?
Lake Formation sits on top of the same Data Catalog that Glue populates, and adds fine-grained permissions - table, column, row, and cell-level access control - on top of what would otherwise be plain IAM access to S3 and catalog resources.
When Lake Formation manages a database or table, Glue crawlers still discover and register schema as usual, but read and write access for jobs, Athena, and Redshift Spectrum queries is now governed by Lake Formation grants instead of (or in addition to) S3 bucket policies, and access is enforced either through temporary credential vending or, in Glue 5.0, through Spark-native fine-grained access control for read paths.
This combination is what lets an organization maintain one shared catalog and enforce different column-level permissions per consumer - for example, masking a salary column for one analyst group while leaving it visible to another - without duplicating the underlying data.
Take quiz
Automatic schema inference
Spark cluster provisioning
Fine-grained, column and row-level access control
Spark-native fine-grained access control
Job bookmarks
Python shell jobs
26. What is the difference between AWS Glue crawlers and manually created Data Catalog tables?
A crawler discovers schema automatically by sampling the underlying data, while a manually created table is one you (or a script using the Glue API, CLI, or CloudFormation) define explicitly, specifying every column, type, and location yourself.
Manual definition gives you exact, stable control - useful when you already know the schema, want it version-controlled as infrastructure-as-code, or need to avoid a crawler's inference quirks (like collapsing similar-but-different schemas into one table, or splitting one logical dataset into several tables).
Crawler-based cataloging is faster to set up and adapts automatically as new partitions or files with a compatible schema appear, which suits datasets that grow continuously and do not need pinpoint schema control - but it can silently redefine a table's schema on a later run if the underlying files' shape drifts.
Take quiz
Exact, stable schema control defined as code
Automatic schema-drift detection
The crawler to run faster
Never update after creation
Silently redefine the table's schema if source files drift
Require manual IAM policy changes every week
27. How do you troubleshoot a Glue crawler that creates duplicate or fragmented tables?
This almost always traces back to the crawler seeing files under one logical dataset that do not look similar enough to it, so it creates a separate table per distinct-looking group instead of one partitioned table.
- Check the S3 prefix structure - inconsistent folder layouts (some files directly under the dataset prefix, others nested an extra level deeper) will split into separate tables.
- Compare schemas and file formats across the "duplicate" files - a stray CSV mixed into a set of Parquet files, or a renamed or reordered column, is a common trigger.
- Check the crawler's table-level grouping configuration, which controls how aggressively similar-looking paths are merged into one table.
- Delete the fragmented tables and rerun the crawler after fixing the underlying inconsistency, rather than repeatedly re-crawling on top of the bad state.
Enabling the crawler's schema-change logging also shows exactly which files or partitions it treated as incompatible on the run that caused the split, which is the fastest way to confirm the root cause before changing configuration.
Take quiz
Too many IAM permissions
Inconsistent folder structure or schema across files in the same dataset
Running the crawler too frequently
Job bookmarks
Classifier priority
Table-level grouping policy
28. Why should you partition data written by an AWS Glue job?
Partitioning splits a dataset into subdirectories based on one or more column values (commonly date, region, or category), and both Glue itself and query engines reading the catalog use that structure to skip irrelevant data entirely rather than scanning every file.
Without partitions, a query for last week's orders, or a Glue job filtering on a date range, still has to read every file in the table, which gets slower and more expensive as the dataset grows; with partitions, the same query prunes down to only the matching subdirectories before reading a single row.
The tradeoff is over-partitioning: choosing a partition key with very high cardinality (like a customer ID) creates enormous numbers of tiny files and partitions, which slows crawlers, bloats Data Catalog metadata, and can actually hurt performance instead of helping it - so partition keys are usually chosen for coarse, frequently filtered dimensions like date.
Take quiz
Avoid using the Data Catalog
Encrypt data automatically
Skip irrelevant files instead of scanning everything
Very high cardinality, creating too many tiny partitions
Values that never change over time
Too few distinct values
29. What is the purpose of pushdown predicates in AWS Glue?
A pushdown predicate is a partition-level filter you supply when reading a partitioned table, which Glue applies before listing or reading any actual files - it prunes at the partition-metadata level in the Data Catalog rather than filtering rows after they have already been read into memory.
This matters because a naive filter applied after create_dynamic_frame.from_catalog() still requires Glue to list and potentially scan every partition's files first; a pushdown predicate passed via the push_down_predicate argument skips that listing step entirely for partitions that do not match, cutting both runtime and S3 request cost.
datasource = glueContext.create_dynamic_frame.from_catalog( database="sales_db", table_name="orders", push_down_predicate="(year=='2026' and month=='08')" )
Pushdown predicates only work on the columns used as partition keys - they cannot filter on a non-partition column, since that value is not known until the file itself is read.
Take quiz
Partition metadata, before files are listed or read
Individual row values after loading
The Data Catalog's IAM policy
Computed or derived columns
Columns that are partition keys
Any column in the table
30. How does AWS Glue handle schema evolution in the Data Catalog?
When a crawler re-runs against a dataset whose shape has changed - a new column appended, a column dropped, or a type changed - it compares the new inferred schema against the existing catalog table and, depending on its configured schema-change policy, either updates the table's schema, adds the new columns while keeping old ones, marks the table for a partial update, or leaves it untouched and logs the difference for review.
For ETL jobs reading that same evolving table, a DynamicFrame naturally tolerates a column appearing only in some records (representing it as null elsewhere, or as a choice type if the types conflict), whereas a Spark DataFrame would need a schema-merge strategy or would fail outright on mismatched columns.
Because catalog updates by default apply globally across all consumers (Athena, Redshift Spectrum, other Glue jobs), a schema change from a crawler auto-update - like a column silently being reinterpreted as a different type - can quietly break a downstream query, which is why many teams set crawlers to log the change but not update the table, and apply schema changes through a reviewed, manual process instead.
Take quiz
Failing the job immediately
Representing it as null where absent, or as a choice type on conflict
Deleting those records
Auto-update the table on every crawl
Disable the crawler entirely
Log changes but do not update the table automatically
31. What is the difference between a Glue workflow and an AWS Step Functions state machine?
Both orchestrate multi-step pipelines, but they differ in scope and flexibility: a Glue workflow only orchestrates Glue-native resources (crawlers, jobs, triggers), while Step Functions can orchestrate almost any AWS service - Lambda, ECS, SageMaker, SNS, and Glue jobs among them - as steps in a single state machine.
| Glue workflow | Step Functions |
| Orchestrates only Glue crawlers, jobs, and triggers | Orchestrates any supported AWS service |
| Built-in visual graph inside the Glue console | Visual state machine in its own console |
| Simple conditional, scheduled, or on-demand triggers | Rich branching, retries, parallel states, error handling |
| No extra service charge beyond Glue usage | Billed separately per state transition |
Teams often pick Glue workflows for pipelines that are entirely Glue-based and simple, and reach for Step Functions once the pipeline needs to call non-Glue services, needs finer-grained error handling and retries, or is one piece of a larger cross-service application.
Take quiz
Only trigger Glue crawlers
Replace the Data Catalog
Orchestrate steps across many different AWS services
State transitions
Stored catalog objects
DPU-hours
32. When would you choose Glue Ray jobs over Spark jobs?
Choose Ray when the workload is fundamentally Python-native, iterative, or ML-oriented - things like distributed hyperparameter search, custom Python preprocessing with irregular parallelism, or feature engineering pipelines built around Python libraries (scikit-learn, XGBoost) that do not map cleanly onto Spark's DataFrame model.
Spark remains the better fit for classic large-scale ETL: SQL-style joins, aggregations, and columnar transformations over huge datasets, where Spark's query optimizer and mature Data Catalog integration outperform hand-written Python parallelism.
Ray jobs run on the Z.2X worker type rather than the G.x family, are priced under the same DPU-hour model, and are a relatively recent addition to Glue, so most existing production pipelines and community examples still assume Spark - Ray is the right call specifically when the task is Python or ML-shaped, not when it is a relational transformation that Spark already handles well.
Take quiz
Python-native ML preprocessing and iterative workloads
Large SQL-style joins and aggregations
Crawler scheduling
G.1X
Z.2X
G.025X
33. How do you use Glue interactive sessions for development?
Interactive sessions give you a live, on-demand Spark backend you can attach a Jupyter notebook or an IDE to, without waiting for a full job to be defined and started - useful for exploring a dataset or writing a script iteratively before turning it into a production job.
You start a session (via the Glue Studio notebook, SageMaker or Jupyter with the aws-glue-sessions kernel, or the API), attach your notebook to it, and run cells against real Glue-managed Spark compute; the session stays warm between cells so you are not paying the cold-start cost repeatedly, and it auto-terminates after a configurable idle timeout to avoid runaway billing.
Once the logic works, you export or save the notebook's cells as a Glue job script, so interactive sessions function as a development environment that hands off directly into the same production job type, rather than a separate throwaway tool.
Take quiz
Running unattended nightly production jobs
Iterative, exploratory development against live Spark compute
Storing table metadata
Can only run for exactly 5 minutes
Run forever until manually stopped
Auto-terminate after a configurable idle timeout
34. What is the difference between the relationalize and unnest transforms in AWS Glue?
Both flatten nested, semi-structured data (like JSON with nested objects and arrays) into tabular form, but they take different approaches and produce different shapes of output.
unnest flattens nested structs into top-level columns in the same DynamicFrame, using a dot- or underscore-joined naming convention for the new column names, and is a lighter-weight operation suited to moderately nested data without repeating arrays.
relationalize goes further: it fully normalizes deeply nested JSON - including arrays - into a set of separate DynamicFrames that mirror a relational schema, generating join-key-style columns so the pieces can be reassembled with joins, which is the right tool when the source has arrays of objects that need to become proper child tables rather than flattened columns.
flat = nested_dyf.unnest() tables = nested_dyf.relationalize( root_table_name="orders", staging_path="s3://bucket/tmp/" )
Take quiz
Produces multiple separate relational tables
Only works on Parquet files
Flattens nested structs into columns within the same frame
Arrays of nested objects that should become child tables
A single top-level string field
Only flat, non-nested fields
35. What is the difference between AWS Glue Data Quality rules on the Data Catalog versus inside an ETL job?
Data Quality rules can be attached in two places, and they serve different moments in the data lifecycle: catalog-level rules run against a table already registered in the Data Catalog, on a schedule or on demand, checking the data at rest; ETL-job-level rules run inline inside a Spark job's script, checking a DynamicFrame in flight before it is written anywhere.
Catalog-level checks are simpler to set up (no code - point-and-click rule recommendation against an existing table) and are good for ongoing monitoring of tables other pipelines or teams already populate, while job-level checks let you branch on the result programmatically - for example, routing failing records to a quarantine path with evaluate_data_quality_multiframe instead of just reporting a score after the fact.
A common pattern combines both: job-level checks gate what gets written into a curated table in the first place, and catalog-level checks continue monitoring that same table afterward to catch drift introduced by any other process that writes to it.
Take quiz
Programmatically route failing records, such as to quarantine, during the job
Only report a score after the fact
Skip writing DQDL rules entirely
Gating writes inside a single job
Ongoing monitoring of a table already in the catalog
Replacing the Schema Registry
36. Explain the execution flow of an AWS Glue crawler?
A crawler run moves through a fixed sequence of internal stages every time it executes, regardless of the data store it targets.
- Connect - the crawler uses its configured connection (if any) to reach the data store, launching elastic network interfaces into the specified VPC and subnet for JDBC sources.
- Sample and classify - it lists objects or tables, samples a subset of files, and runs them through the ordered list of classifiers until one returns a confident match for format and schema.
- Infer schema and group - it infers column names and types from the classified sample and groups files with compatible schemas into partitions of a single table, based on similarity and the crawler's grouping policy.
- Compare to existing catalog state - for a re-crawl, it diffs the newly inferred schema against what is already in the Data Catalog for that table.
- Write to the Data Catalog - it creates new tables and partitions, updates existing ones, or only logs the difference, according to the crawler's configured schema-change and deletion-behavior policies.
flowchart TD
A["Crawler starts"] --> B["Connect to data store"]
B --> C["Sample files"]
C --> D["Run classifiers in priority order"]
D --> E["Infer schema, group into partitions"]
E --> F{Table already exists?}
F -- No --> G["Create new table"]
F -- Yes --> H["Diff against existing schema"]
H --> I["Update, log, or skip per policy"]
G --> J["Crawl complete"]
I --> J
Take quiz
Run a Spark job
Connect to the configured data store
Write to the Data Catalog
Always overwriting it unconditionally
Asking the user interactively
Diffing the new schema against the existing catalog entry and applying its configured policy
37. Explain the internal working of AWS Glue job bookmarks?
Internally, a bookmark is a state object Glue persists per job (not per run), keyed by the job name, that records exactly which input was already consumed as of the job's last successful completion.
For an S3 source, that state is a set of file paths paired with size and last-modified timestamp; at the start of a run, Glue's bookmark-aware source APIs (with transformation_ctx set) compare the current file listing against that stored state and only hand the new or changed files to the job's DAG. For a JDBC source, the state is instead the maximum value seen so far in a column you designate, and each run issues a query bounded to values greater than that watermark.
Critically, the bookmark only advances after the job run reports success - a failed or cancelled run leaves the previous bookmark state untouched, so a fixed and re-run job safely reprocesses the same batch rather than skipping it, and each transformation_ctx string in the script is what lets Glue track bookmark state independently for multiple sources read within a single job.
Take quiz
Row-level checksums
IAM role ARNs
File path, size, and last-modified timestamp
The job run completes successfully
The crawler runs
The job script is saved
38. Why doesn't enabling job bookmarks always prevent reprocessing of data?
Bookmarks track what looked new to Glue, not what has already been handled by every downstream system, so several situations cause reprocessing even with bookmarks enabled.
- A file gets overwritten in place with a new last-modified time - the bookmark sees a "new" file and reprocesses it, even if you intended it as unchanged.
- The source is not one Glue can bookmark reliably - some connection types and certain JDBC configurations without a strictly increasing tracking column do not support incremental bookmarking at all.
- The
transformation_ctxon a read node is changed, removed, or duplicated across script edits - bookmark state is keyed to that context string, so altering it detaches the source from its previous history. - A prior run partially failed after writing some output but before reporting success - the bookmark did not advance, so the next run reprocesses that entire batch, including files whose output already landed.
Because of the last case, a downstream sink written by a bookmarked job should generally be idempotent - safe to receive the same batch twice - rather than assuming bookmarks alone guarantee exactly-once delivery.
Take quiz
A source file is overwritten in place with a new timestamp
The job never runs at all
The crawler is deleted
Read-only
Idempotent, so reprocessing the same batch is safe
Deleted after every run
39. How can you optimize the cost of AWS Glue ETL jobs?
Most Glue cost overruns come from a handful of repeatable patterns, and addressing them usually cuts a bill significantly without touching the job's actual logic.
- Right-size DPU allocation - the default of 10 workers is far more than most jobs need; profile actual usage and scale down, or turn on auto-scaling so Glue adjusts workers to the job's real Spark-stage parallelism instead of running at a fixed size the whole time.
- Use Flex execution for any batch job that can tolerate a delayed start, for roughly a 34% lower DPU-hour rate than Standard.
- Enable job bookmarks so incremental runs process only new data instead of the full dataset every time.
- Use columnar formats and partition pruning - switching CSV or JSON sources to Parquet, and filtering with pushdown predicates on partition keys, both cut the volume of data actually read.
- Watch the Data Catalog itself - beyond the first million free objects and requests, excessive partitions and table versions add a smaller but easy-to-forget recurring charge.
Together, right-sizing workers, adding Flex where possible, and enabling bookmarks are typically the highest-leverage changes, since they compound: fewer, cheaper DPU-hours applied to a smaller volume of re-read data.
Take quiz
Renaming job scripts
Right-sizing workers, Flex execution, and job bookmarks together
Increasing the crawler's schedule frequency
The console UI
IAM roles
The Data Catalog's stored objects and requests
40. How do you troubleshoot an AWS Glue job that fails with an out-of-memory error?
An out-of-memory failure in a Glue Spark job almost always means either too little memory per executor for the data being processed, or a shape of the job - skew, a huge broadcast join, excessive caching - that concentrates far more data on one executor than it can hold.
- Check the Spark UI and CloudWatch metrics for the failed run to see whether a small number of tasks handled a disproportionate share of the data (a sign of skew) versus every task failing uniformly (a sign of simple under-provisioning).
- If under-provisioned, move to a larger worker type (G.2X or G.4X instead of G.1X) rather than just adding more workers of the same size, since the issue is per-executor memory, not total cluster memory.
- If a broadcast join is the cause - Spark tries to broadcast a "small" table that turned out to be large - lower or disable the broadcast join threshold, or explicitly force a sort-merge join instead.
- If a small number of partition keys dominate the data, salt the join key or repartition on a composite key to spread that data across more executors.
- Check for unnecessary collect or cache calls pulling large amounts of data onto the driver or holding more in memory than needed for the rest of the job.
Reproducing the failure at a smaller scale in a Glue interactive session is often the fastest way to confirm which of these is the actual cause before changing production job settings.
Take quiz
A corrupted S3 bucket
An expired IAM role
Data skew
A broadcast join threshold that's too permissive
Job bookmarks
Too many crawlers running
41. Explain the lifecycle of an AWS Glue job run?
A single job run passes through a well-defined set of states from submission to completion, which the Glue console, API, and CloudWatch all expose identically.
flowchart LR
A[STARTING] --> B[RUNNING]
B --> C{Outcome}
C -- Success --> D[SUCCEEDED]
C -- Script or data error --> E[FAILED]
C -- User or API stop --> F[STOPPING] --> G[STOPPED]
C -- Time limit exceeded --> H[TIMEOUT]
In STARTING, Glue provisions the requested workers (or, for Flex, waits for spare capacity) and initializes the Spark or Python shell environment; in RUNNING, the actual script executes, emitting metrics and logs to CloudWatch throughout.
On success the job reaches SUCCEEDED and, if bookmarks are enabled, its bookmark state advances only at this point; on an unhandled exception it reaches FAILED with the stack trace available in the job's error logs; a manual stop or an API call moves it through STOPPING to STOPPED; and exceeding the job's configured timeout produces a distinct TIMEOUT state so it is distinguishable from a script-level failure when auditing run history.
Take quiz
Only after it reaches SUCCEEDED
When it is first created
When it enters RUNNING
FAILED
TIMEOUT
STOPPED
42. Which is better and why: G.2X or G.4X worker type for a memory-heavy join?
Neither is universally "better" - it depends on whether the bottleneck is memory per executor or total parallelism, and on cost per unit of useful work rather than sticker price per worker.
G.2X gives 8 vCPU and 32 GB per worker with one executor per worker; G.4X gives 16 vCPU and 64 GB per worker, also one executor per worker, at double the DPU (and cost) per worker. For a join that is failing or spilling to disk on G.2X purely because a single executor cannot hold its shuffle partition in memory, moving to G.4X directly doubles the memory available to each executor, which often resolves the failure outright - simply adding more G.2X workers would not help, since it increases total cluster capacity but not the memory ceiling of any individual executor handling a skewed partition.
If the join instead just needs more overall parallelism on evenly distributed data - not more memory per task - the same total DPU budget spent on more G.2X workers, rather than fewer G.4X ones, usually gives better throughput, since you get more, smaller executors working the problem in parallel for the same spend. So the right answer is G.4X specifically when the constraint is per-executor memory (skew, large shuffle partitions, wide joins), and G.2X when the constraint is overall parallelism on well-distributed data.
Take quiz
Total cluster parallelism on evenly distributed data
Per-executor memory, such as a skewed shuffle partition
Data Catalog request limits
Data is heavily skewed
The job is a Python shell job
Data is evenly distributed and more parallelism helps
43. How does AWS Glue Data Quality evaluate rules written in DQDL?
At evaluation time, Glue parses the DQDL ruleset into individual rule expressions - each naming a rule type (like Completeness, IsUnique, ColumnValues, or a custom SQL-based rule), a target column or expression, and a threshold or condition - and compiles them into checks the underlying Deequ engine runs as a set of Spark aggregations over the target dataset.
Deequ computes the relevant metric for each rule in largely a single pass over the data where possible - for example, null counts and distinct counts can share a scan - rather than running one full scan per rule, which keeps evaluation cost roughly proportional to data size rather than to rule count.
Each rule then evaluates to Pass or Fail against its threshold, an overall Data Quality score is computed as the proportion of rules that passed, and - critically for remediation - Glue can also return the specific rows that caused a rule to fail, not just the aggregate pass or fail, so you can route or inspect exactly the records responsible for a drop in score.
Take quiz
Apache Airflow
AWS Config
The open-source Deequ framework
The specific rows that caused a rule to fail
A new IAM policy
A rewritten DQDL ruleset
44. What happens when a Glue crawler encounters incompatible schemas across files in the same table?
The crawler does not fail the crawl outright; instead it applies its configured grouping and schema-comparison behavior to decide whether the incompatible files still belong to one table.
If the differences are compatible - a superset or subset of columns, or types that can be safely widened, such as int to long - the crawler merges them into one table with a combined schema, representing columns absent from some files as nullable. If the differences are genuinely incompatible - conflicting types for the same column name, or a fundamentally different file format mixed into the same prefix - the crawler's grouping policy typically splits the data into separate tables, or separate partitions with divergent schemas, which then causes query-time errors in engines that expect one schema per table, rather than silently coercing the data.
This is exactly the fragmentation issue interview candidates are expected to recognize and fix by tightening the source layout or classifiers - because leaving it as an accidental multi-table split, rather than an intentional one, means downstream Athena or Redshift Spectrum queries against "the table" quietly miss part of the data.
Take quiz
Merges them into one table with a combined, nullable-where-absent schema
Deletes the older files
Refuses to crawl the dataset
Automatic data deletion
One table silently missing part of the data
A Step Functions failure
45. Explain the internal working of DynamicFrame's resolveChoice with the "cast" action?
When Glue detects a column with a "choice" type - meaning different records reported different underlying types for the same field during schema inference - each affected record internally carries a small tagged union: the value plus which of the observed types it actually is.
Calling resolveChoice(specs=[("col", "cast:long")]) tells Glue to attempt, per record, to coerce that column's value into the target type (long here) regardless of which original type it was tagged as - a string like "42" is parsed into a long, an integer is widened, and so on - producing a single, uniformly typed column in the output DynamicFrame.
Records where the value genuinely cannot be cast to the target type - a string like "N/A" being cast to long, for instance - do not silently vanish; depending on the surrounding job's error-handling configuration, they either produce a null in that column, or are captured separately if the job has error-record handling turned on, so the resulting record count and any lost values are traceable rather than a mystery.
Take quiz
Rename the column
Coerce each record's value to a single target type regardless of its original tagged type
Delete every record with a choice type
A silently duplicated row
An automatic schema rollback
A null, or a separately captured error record depending on configuration
46. How can you optimize a Glue crawler's runtime and cost across many small S3 partitions?
Crawlers bill on the same DPU-hour model as jobs, with a 10-minute minimum per run, so a crawler that has to list and sample thousands of small partitions individually gets both slow and disproportionately expensive relative to the data volume it is actually cataloging.
- Use incremental crawls - configure the crawler to only crawl new S3 folders or partitions since its last run instead of re-scanning the entire dataset's existing partitions every time.
- Add explicit partition indexes on the catalog table, which speeds up how quickly Athena and Spark can resolve which partitions match a query's filters, independent of the crawl itself.
- Consolidate small files upstream - many tiny partitions are often a symptom of many tiny files; compacting them via a scheduled Glue job reduces both crawl time and the small-file problem that also slows down downstream jobs.
- Use exclude patterns to skip known-irrelevant prefixes (temp or staging folders, non-data files) so the crawler does not spend time sampling paths that will never become table data.
- Register new partitions directly via the Glue API or an ALTER TABLE ADD PARTITION statement for a well-known, regular partition scheme, which can replace routine crawler runs entirely once the schema itself has stabilized.
Incremental crawling combined with upstream file compaction usually gives the largest combined win, since it shrinks both the crawler's own bill and the cost of every downstream job or query touching the same data.
Take quiz
Increasing the crawler's worker count
Disabling classifiers
Enabling incremental crawling of only new partitions
Many small underlying files that should be compacted
An oversized Data Catalog free tier
Too few classifiers
47. Why should you avoid using coalesce(1) on large datasets in AWS Glue Spark jobs?
coalesce(1) merges all of a DataFrame or DynamicFrame's partitions down to exactly one, which forces every executor to funnel its data through a single task on a single executor for the final write - eliminating the parallelism that made the job fast in the first place.
On a large dataset, that single task has to hold and write far more data than any one executor's memory is really sized for, which commonly causes the exact out-of-memory or extremely slow-write symptoms teams are trying to avoid, and it also produces one enormous output file instead of many right-sized ones, which then makes downstream reads - Athena, Spark, Redshift Spectrum - less parallelizable too, since a single huge file cannot be split across as many reader tasks.
The usual fix is repartition(n) to a sensible number of output files, sized to a target file size like 128 to 256 MB rather than to an arbitrarily small count, or bucketing and partitioning by a business key if downstream queries filter on it - coalesce(1) is only appropriate for genuinely small result sets where one tidy output file matters more than write performance or downstream read parallelism.
Take quiz
Forces the entire dataset's final write through a single task and executor
Disables job bookmarks
Increases the number of parallel output files
Disabling partitioning entirely
repartition(n) to a sensible target file size
Always using coalesce(1) regardless of size
48. How do you troubleshoot slow AWS Glue jobs caused by data skew?
Skew means a small number of partition keys, or shuffle partitions, hold a disproportionate share of the data, so a handful of tasks run far longer than the rest while most executors sit idle waiting for them - visible in the Spark UI as a stage where the max task duration is many times the median.
- Confirm it is skew, not just under-provisioning, by checking the Spark UI's stage view for a small number of long-running tasks against a shuffle or join stage specifically, rather than uniformly slow tasks across the board.
- Salt the skewed key - append a small random suffix to the join or group-by key on both sides of a join (or during an aggregation) to spread one hot key's rows across multiple partitions, then aggregate the partial results back together.
- Use adaptive query execution - Spark's AQE, available on recent Glue versions, can automatically split skewed shuffle partitions at runtime; confirming it is enabled is often the cheapest fix before hand-salting a key.
- Broadcast the smaller side of a skewed join explicitly if one side is small enough, avoiding a shuffle-based join entirely for that step.
- Repartition on a composite key - the original key plus a secondary, more evenly distributed column - rather than the skewed key alone, when salting a join is not practical.
Because a skewed job can look identical to an under-provisioned one from the outside - both are just "slow" - checking the per-task duration distribution in the Spark UI before changing worker count or type is the step that prevents throwing money at the wrong problem.
Take quiz
Insufficient IAM permissions
Data skew
An expired job bookmark
Deleting the hot key's rows
Disabling the shuffle entirely
Spreading a hot key's rows across multiple partitions using a random suffix
49. What is the difference between AWS Glue's Spark-native fine-grained access control and Lake Formation credential vending?
Both enforce Lake Formation's table, column, and row-level permissions on data a Glue Spark job reads, but they differ in where and how that enforcement happens.
| Credential vending | Spark-native FGAC (Glue 5.0+) |
| Lake Formation hands out scoped, temporary S3 credentials per query | Permissions enforced inside the Spark engine during query planning |
| Works broadly across engines (Athena, Redshift Spectrum, Glue) | Specific to the Glue Spark runtime on supported versions |
| Filtering happens at the storage access layer | Filtering happens natively within the Spark execution plan |
| Broad support for read and write paths | Currently read-focused; some write and Iceberg-via-Lake-Formation limitations apply |
In practice, credential vending is the older, more universally supported mechanism, while Spark-native FGAC is a newer, tighter integration specific to Glue 5.0 Spark jobs that avoids some of the overhead of per-query credential exchange - but its current limitations, like unsupported data writes in some configurations, mean many production setups still rely on credential vending for anything beyond straightforward reads.
Take quiz
Only through S3 bucket policies
Only for Athena queries
Natively within Spark's own query execution
Has restrictions on certain data write paths
Can't be used with the Data Catalog
Doesn't support any read operations
50. Explain the tradeoffs of using AWS Glue Flex execution class for production ETL pipelines?
Flex trades startup-time guarantees for a lower per-DPU-hour rate, so the tradeoff is really about how much your pipeline's SLA depends on predictable start times versus how much it benefits from roughly 34% cheaper compute.
On the upside, for a batch job with a wide execution window - a nightly load that just needs to finish before business hours, for instance - Flex gives the identical Spark engine and identical results at a meaningfully lower cost, with zero code changes beyond the execution-class setting, and Glue only bills for capacity it actually acquires rather than what you requested if some workers get reclaimed mid-run.
On the downside, Flex is not supported for Python shell, streaming, or ML jobs at all, its start time is inherently variable since it depends on spare capacity being available, and a chain of dependent Flex jobs in a tightly timed workflow can compound that variability into a much less predictable end-to-end pipeline duration - which makes Flex a poor fit for anything feeding a hard downstream deadline, such as a dashboard refresh a business team is watching at a fixed time each morning, even though it is a near-free win for genuinely flexible batch workloads.