// the find
apache/hudi
Upserts, Deletes And Incremental Processing on Big Data.
Apache Hudi is a table format with a write path and background table services layered over Parquet files in object storage, with Spark and Flink as the main writers. It suits teams running a data lake who need upserts, deletes and incremental reads, and who are prepared to run compaction and cleaning themselves.
- Record-level indexes, built on row-oriented formats and bloom filters and maintained by writes, let upserts and deletes locate affected file groups directly instead of scanning whole partitions.
- Change-data-capture queries return before and after images for each changed record between two commits, so downstream consumers can replay changes without a separate change log.
- Table services (cleaning, clustering, compaction, index building) run inline in the Spark or Flink writer or as standalone jobs, with configurable scheduling and built-in failure handling, instead of a set of maintenance scripts you write yourself.
- The build matrix covers Spark 3.3 through 4.2 and Flink 1.18 through 2.2, and catalog sync exists for Hive Metastore, AWS Glue, BigQuery and Apache XTable, so the project tracks current engine versions.
- The bundle you get depends on the Spark, Scala and Flink flags together. The Flink bundle cannot be built with the Scala 2.13 profile, so a wrong combination is easy to produce and the README table is long enough to show how.
- Java 11 or 17 is the only range the README lists, and the Spark 4 profiles require 17. Anyone on a newer JDK or a locked-down build image should check before building.
- The compression default is a footgun. ZSTD is the write default on Spark 3.5 and newer, but Spark 3.3 and 3.4 stay on GZIP because of an off-heap leak in the non-vectorized reader (PARQUET-2160). The README concedes that keeping GZIP does not protect you when another engine writes ZSTD files.
- The README is a feature list that links out to the quickstart and says nothing about operating costs, such as compaction overhead or how to choose between copy-on-write and merge-on-read tables.