Metadata Plane · Table & Technical Catalogues

Datastrato Enterprise

Commercial distribution of Apache Gravitino, adding SCIM provisioning, role-based access control across federated catalogues, an administration UI, audit logging and a supported Helm chart.

Overview

Datastrato Enterprise is a commercial distribution of Apache Gravitino, and its documentation is unusually plain about what that means: the same engine, the same API, the same concepts. Metalakes, catalogues, schemas and tables keep their meanings, configuration keys and connector names are unchanged, and the documentation says "Gravitino" when it means the engine and "Datastrato Enterprise" when it means the distribution.

What it adds is the operational layer. Identity comes from SCIM provisioning against an existing directory, with local accounts as a fallback when the identity provider is unavailable, and role-based access control is enforced across every federated catalogue including requests arriving through the Iceberg REST endpoint.

Governance covers tags, policies and ownership across connected sources, with coverage reported per catalogue and per asset type, an audit log of metadata operations, and lineage collected from Spark through OpenLineage. An administration UI handles catalogue connection, access review and monitoring.

Operationally it is a supported Helm chart with health and readiness endpoints wired for orchestration and metrics for every service, plus support with a defined response commitment and a maintained release line.

Deployment is deliberately narrow: Kubernetes only, with no tarball, no Docker Compose and no bare-metal path, on the stated reasoning that one artefact means one thing to test, support and upgrade. The chart and its images come from one OCI registry, which also means one mirroring step for air-gapped clusters.

The company is explicit that nothing in the data path is proprietary, that metadata stays in open formats, and that moving between the distribution and upstream Gravitino in either direction does not require rewriting integrations. Version 1.3.0 tracks the upstream release of the same number.

Key features and capabilities

The same headings are used for every table & technical catalogues entry, so two tools can be read side by side.

What it catalogues
  • Identical to Apache Gravitino by design: metalakes, catalogues, schemas, tables, filesets and models
  • Configuration keys, REST endpoints, client libraries and connector names unchanged from upstream
  • Tags, policies and ownership applied across every connected source
  • Coverage reported per catalogue and per asset type, so governance gaps are visible
Table formats
  • The upstream connector set, since the engine is Gravitino itself
  • Marketed emphasis on federating in place: Iceberg, Hive, Glue, relational sources, Kafka, files and models
  • Metadata held in open formats, with nothing proprietary in the data path
Engines that can use it
  • Documented engine coverage for Trino, Spark and Flink
  • Anything that speaks the Iceberg REST protocol
  • Access control applies to Iceberg REST requests, not only to the native API
Access control and auditing
  • Role-based access control enforced across every federated catalogue
  • SCIM provisioning from an existing identity provider, with local accounts as a fallback
  • Audit log of metadata operations
  • Lineage collected from Spark through OpenLineage
Interoperability
  • Same API as upstream Gravitino, so integrations carry across unchanged
  • Migration stated as possible in both directions without rewriting integrations
  • Upstream documentation is referenced directly for anything the distribution does not change
Table maintenance and operations
  • Supported Helm chart, with a maintained release line
  • Health and readiness endpoints wired for orchestration; readiness means the entity store is reachable
  • Metrics exposed for every service the server runs
  • Administration UI for catalogue connection, access review and monitoring
  • Support with a defined response commitment
How it runs
  • Kubernetes only: no tarball, no Docker Compose, no bare-metal path
  • Kubernetes 1.29 or later and Helm 3.8 or later, which is the floor for OCI registry support
  • Minimum for the Gravitino pod: 2 CPU cores and 4 GiB memory, plus a default storage class
  • Activated with a licence key; the chart and images ship from one OCI registry, so air-gapped mirroring is a single step
  • Bring your own relational database for production; the chart can deploy one for evaluation only
  • API and UI on port 8090; a local evaluation path runs the same chart against local Kubernetes

Pricing

Price on request

No pricing is published. Trial licence keys are issued through the Datastrato trial page and production keys come from an account team, so the commercial terms are a conversation rather than a list price.

Vendor pricing page →

Demos and videos

About Datastrato

Datastrato is the company that created Gravitino and donated it to the Apache Software Foundation, where it is now a top-level project. It employs many of the project's contributors and maintains this distribution on top of the upstream engine. Its own documentation frames the relationship carefully, pointing readers to gravitino.apache.org for the upstream project and stating that where the two sets of documentation differ, the Enterprise one describes only what the distribution adds. It publishes no pricing.

datastrato.ai

Other table & technical catalogues tools

Apache Gravitino

Metadata Plane · Table & Technical Catalogues

Federated metadata catalogue for tables, files, streams and models, managed in place across sources and served to Spark, Trino and Flink through one namespace.

  • Open source

Apache Polaris

Metadata Plane · Table & Technical Catalogues

Open-source catalogue that implements the Apache Iceberg REST catalogue API.

  • Open source

AWS Glue Data Catalog

Metadata Plane · Table & Technical Catalogues

Managed technical metadata catalogue for data on AWS, used by services such as Athena, EMR and Redshift.

  • Cloud service

Hive Metastore

Metadata Plane · Table & Technical Catalogues

Apache Hive's metadata service, still widely used as a table catalogue by Spark and other engines.

  • Open source

Unity Catalog

Metadata Plane · Table & Technical Catalogues

Governance catalogue for data and AI assets, available as open source and managed in Databricks.

  • Open core

Drafted with AI assistance and checked against the vendor’s own documentation.