Prev Next

AI / Apache Paimon Interview questions

How do you troubleshoot excessive small files from streaming writes into a Paimon table?

Streaming writes commit frequently and in small batches, which naturally produces many small data files — left unmanaged, this degrades both read performance (more files to open and merge) and puts pressure on the underlying filesystem/object store.

  1. Check the $files system table. Query my_table$files to see the actual file count and size distribution per partition/bucket, confirming small files really are the issue before tuning anything.
  2. Verify compaction is running. Streaming writers can run compaction inline, but heavy-throughput pipelines are often better served by a dedicated, separately-scaled compaction job so writing isn't slowed down by compaction work.
  3. Tune commit frequency. If the upstream job commits (checkpoints) too often relative to data volume, each commit produces disproportionately small files; widening the checkpoint interval reduces file count at the cost of end-to-end latency.
  4. Review bucket count. Too many buckets for the actual data volume spreads writes thin, producing small files per bucket; Paimon recommends sizing buckets so each holds roughly 200MB–1GB, and reducing bucket count (or switching to Dynamic Bucket) if buckets are consistently under that.
  5. Consider Postpone Bucket mode for very high-throughput, low-latency ingestion, deferring bucketing and merging to a scheduled batch compaction instead of doing it inline on every commit.
Which system table would you query first to confirm a small-file problem?
Too many buckets for the actual data volume tends to:

More Related questions...

What is Apache Paimon? What is the purpose of Paimon's Catalog abstraction? What are the four types of metastores Paimon catalogs support? Define a primary key table in Paimon? What are the two main table types in Paimon? Describe a Paimon Snapshot? What are Manifest files used for in Paimon? What is a Bucket in Paimon? List the merge engines Paimon supports for primary key tables? What is the default merge engine in Paimon? What is the purpose of the changelog-producer table property? How do you use the sequence.field option in Paimon? What is Paimon's LSM tree used for? What are Sorted Runs in a Paimon LSM tree? Describe what Paimon's system tables are for? How do you use the $snapshots system table? What is the purpose of Tags in Paimon? What is an Append Table in Paimon? How do you apply schema evolution during CDC ingestion into Paimon? What are the consistency guarantees Paimon provides for writers? What is the purpose of the Audit Log system table? Describe the Read-optimized (ro) system table? Why does Paimon combine a lake format with an LSM-tree structure? How does the Deduplicate merge engine handle a DELETE record? What is the difference between the Partial Update and Aggregation merge engines? What is the difference between Fixed Bucket and Dynamic Bucket modes? When should you choose Postpone Bucket mode? What happens when Paimon's sorted runs contain overlapping primary key ranges? Why should changelog-producer be enabled only when necessary? What is the difference between the input and lookup changelog producers? How does the full-compaction changelog producer differ from lookup? Why do concurrent writers to the same partition only get snapshot isolation instead of full isolation? When should you choose the Hive catalog over the filesystem catalog? How can you optimize primary key lookups using the bucket-key option? What is the difference between the sequence.field and rowkind.field options? Explain the execution flow of a two-phase commit when a Paimon writer flushes data? Why is Cross Partitions Upsert more expensive than a normal bucketed upsert? How do you troubleshoot excessive small files from streaming writes into a Paimon table? What is the difference between the audit_log and binlog system tables? Explain the internal working of automatic tag creation with a watermark? When should you choose the REST catalog over the Hive catalog? What happens when you roll back a Paimon table to an earlier tag? How does Paimon achieve streaming-batch unification on the same table? Why does Paimon recommend keeping bucket data size between 200MB and 1GB? What is the difference between the First Row and Deduplicate merge engines? How can you optimize query performance using the read-optimized system table? What is the difference between Paimon's and Apache Iceberg's core design philosophy? Why did Apache Paimon originally start as part of the Flink project, and what is it called now? Which is better and why: lookup or full-compaction changelog producer for a 30-minute-latency pipeline? How do you troubleshoot duplicate rows appearing when using Dynamic Bucket mode with multiple concurrent write jobs?
Show more question and Answers...


Comments & Discussions