BigData / Apache Parquet Interview Questions
What is Schema Evolution in Parquet?
Schema evolution is the ability to change a Parquet dataset's schema over time without rewriting existing files. Parquet natively supports:
- Adding columns — new columns appear as
nullin older files when read together with newer files. - Renaming columns — supported via field IDs (used in formats like Iceberg on top of Parquet).
- Widening types — e.g., INT32 → INT64 is safe; narrowing is not.
In Apache Spark, set mergeSchema = true to merge schemas across multiple Parquet files automatically:
df = spark.read.option("mergeSchema", "true").parquet("s3://bucket/data/")
Schema evolution is critical for long-lived data lakes where upstream producers add new fields without co-ordinating with all downstream consumers.
More Related questions...