BigData / Apache Parquet Interview Questions
What is the difference between Parquet, ORC, and Avro?
All three are Apache-ecosystem formats but optimised for different workloads:
| Feature | Parquet | ORC | Avro |
|---|---|---|---|
| Layout | Columnar | Columnar | Row-based |
| Ecosystem fit | Spark, Presto, Hive, cloud lakes | Hive-native; Spark | Kafka, Flink, Hadoop MR |
| Compression | Excellent (ZSTD, Snappy) | Excellent (ZLIB, Snappy, ZSTD) | Good (Deflate, Snappy) |
| Schema evolution | Column adds; mergeSchema | Column adds; type promotions | Full backward/forward compatibility |
| Nested data | Strong (Dremel-style) | Good | Native (JSON-like nested) |
| Best for | Analytics, data lake | Hive OLAP, HBase integration | Streaming, messaging, OLTP-like writes |
In practice: most modern data lake stacks (Databricks, AWS lake formation, Google BigQuery) default to Parquet. ORC remains dominant in Hive-heavy shops.
More Related questions...