BigData / Apache Hudi Interview Questions
Describe the history and origin of Apache Hudi?
Hudi was developed at Uber in 2016 under the internal codename "Hoodie," built to solve a real scaling problem: Uber's largest Spark jobs were using over 1,000 executors to rewrite entire datasets just to absorb upstream inserts, updates, and deletes.
Uber open-sourced Hudi in 2017, then donated it to the Apache Software Foundation in January 2019, entering the Apache Incubator. It graduated to a full Apache Top-Level Project in May 2020.
Since then, the project has grown well beyond Uber's original use case, adding a multi-modal indexing subsystem, non-blocking concurrency control, and an LSM-tree-based timeline as part of the Hudi 1.x release line.
More Related questions...