BigData / Apache Parquet Interview Questions
How do you read and write Parquet files in PySpark?
Use spark.read.parquet(path) to load files and DataFrameWriter.parquet(path) to write. Choose compression and partition columns to cut scan cost.
Spark provides first-class Parquet support via the DataFrameReader and DataFrameWriter APIs.
Read:
# Read a single file or directory of Parquet files df = spark.read.parquet("s3://my-bucket/events/") # With options df = (spark.read .option("mergeSchema", "true") .parquet("hdfs:///datalake/transactions/"))
Write:
# Overwrite with Snappy compression (default) df.write.mode("overwrite").parquet("s3://my-bucket/output/") # Partition by date and region, use ZSTD (df.write .partitionBy("date", "region") .option("compression", "zstd") .mode("append") .parquet("s3://my-bucket/partitioned/"))
Register as temp view for SQL:
df.createOrReplaceTempView("events") spark.sql("SELECT date, SUM(revenue) FROM events GROUP BY date").show()
More Related questions...