// the find
apache/gravitino
World's most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.
Apache Gravitino is a Java metadata service that sits in front of existing sources such as Hive, MySQL, HDFS and S3 and exposes one model and API over them. It is aimed at data platform teams that run several query engines against several sources and would otherwise maintain a separate catalog and permission setup for each.
The Iceberg REST catalog and the Trino connector are first-party parts of the project, so engines that already speak those protocols can point at Gravitino directly. Access control, auditing and discovery are defined once against the metadata model, not per engine, which is the part that saves real work when permissions drift across sources. Apache governance gives it a neutral home, and the repo is under active development with a push two days ago.
The README says Windows is not supported for building, so contributors on that platform will be working in WSL or a Linux VM. The AI model and feature tracking is flagged WIP, and the model catalog topics oversell how far that part has got. The front page defers nearly everything to the docs site, and the geo-distribution and federation claims are asserted rather than shown, so test them on your own topology before depending on them. It is also a server you run, with its own metadata store and per-source connector configuration to maintain, not a library you import.