BigData / Apache Parquet Interview Questions
How does Parquet support nested data (structs, lists, maps)?
Parquet uses the Dremel encoding (Google's paper, 2010) to represent arbitrarily nested data in a flat columnar layout. Two extra per-value integers are stored alongside each column's data:
- Definition level — how many optional fields in the path are actually defined (non-null). Encodes null positions without storing nulls explicitly.
- Repetition level — which repeated field in the path started a new list element. Encodes list boundaries.
Example — a column orders.items.price where items is a repeated field:
Row 1: orders.items = [{price: 10}, {price: 20}] Row 2: orders.items = [] (empty list) Stored as: price column: [10, 20] repetition: [0, 1] (0=new top-level record, 1=repeated value) definition: [2, 2, 1] (includes row 2's empty-list marker)
This lets engines reconstruct nested structures perfectly while retaining columnar access for analytics on leaf fields.
More Related questions...