Apache Gravitino
Metadata Plane · Table & Technical Catalogues
Federated metadata catalogue for tables, files, streams and models, managed in place across sources and served to Spark, Trino and Flink through one namespace.
- Open source
Metadata Plane · Table & Technical Catalogues
Commercial distribution of Apache Gravitino, adding SCIM provisioning, role-based access control across federated catalogues, an administration UI, audit logging and a supported Helm chart.
Datastrato Enterprise is a commercial distribution of Apache Gravitino, and its documentation is unusually plain about what that means: the same engine, the same API, the same concepts. Metalakes, catalogues, schemas and tables keep their meanings, configuration keys and connector names are unchanged, and the documentation says "Gravitino" when it means the engine and "Datastrato Enterprise" when it means the distribution.
What it adds is the operational layer. Identity comes from SCIM provisioning against an existing directory, with local accounts as a fallback when the identity provider is unavailable, and role-based access control is enforced across every federated catalogue including requests arriving through the Iceberg REST endpoint.
Governance covers tags, policies and ownership across connected sources, with coverage reported per catalogue and per asset type, an audit log of metadata operations, and lineage collected from Spark through OpenLineage. An administration UI handles catalogue connection, access review and monitoring.
Operationally it is a supported Helm chart with health and readiness endpoints wired for orchestration and metrics for every service, plus support with a defined response commitment and a maintained release line.
Deployment is deliberately narrow: Kubernetes only, with no tarball, no Docker Compose and no bare-metal path, on the stated reasoning that one artefact means one thing to test, support and upgrade. The chart and its images come from one OCI registry, which also means one mirroring step for air-gapped clusters.
The company is explicit that nothing in the data path is proprietary, that metadata stays in open formats, and that moving between the distribution and upstream Gravitino in either direction does not require rewriting integrations. Version 1.3.0 tracks the upstream release of the same number.
The same headings are used for every table & technical catalogues entry, so two tools can be read side by side.
Price on request
No pricing is published. Trial licence keys are issued through the Datastrato trial page and production keys come from an account team, so the commercial terms are a conversation rather than a list price.
Datastrato is the company that created Gravitino and donated it to the Apache Software Foundation, where it is now a top-level project. It employs many of the project's contributors and maintains this distribution on top of the upstream engine. Its own documentation frames the relationship carefully, pointing readers to gravitino.apache.org for the upstream project and stating that where the two sets of documentation differ, the Enterprise one describes only what the distribution adds. It publishes no pricing.
Metadata Plane · Table & Technical Catalogues
Federated metadata catalogue for tables, files, streams and models, managed in place across sources and served to Spark, Trino and Flink through one namespace.
Metadata Plane · Table & Technical Catalogues
Open-source catalogue that implements the Apache Iceberg REST catalogue API.
Metadata Plane · Table & Technical Catalogues
Managed technical metadata catalogue for data on AWS, used by services such as Athena, EMR and Redshift.
Metadata Plane · Table & Technical Catalogues
Apache Hive's metadata service, still widely used as a table catalogue by Spark and other engines.
Metadata Plane · Table & Technical Catalogues
Governance catalogue for data and AI assets, available as open source and managed in Databricks.
Drafted with AI assistance and checked against the vendor’s own documentation.