Metadata Plane · Data Catalogues & Discovery

DataHub

Open-source metadata platform for discovery, lineage and governance, with a managed cloud offering.

Overview

DataHub is a metadata platform for finding, governing and monitoring data. It was built at LinkedIn and open-sourced in 2020, and is now developed by the company behind it, which sells a managed version alongside the open-source core.

Metadata arrives through connectors, run either from the user interface or the command line, and can also be pushed through APIs or Kafka. The vendor quotes different connector counts in different places: 80 or more production-grade connectors in the repository, and 150 or more systems in the integrations directory.

Discovery covers datasets, dashboards, ML models and raw files, with profiling statistics and usage information feeding search. Lineage comes in three ways: extracted automatically by connectors, edited by hand in the interface, or emitted through the API, and it runs to column level as well as table and pipeline level.

Governance features include a business glossary with term groups and inheritance, tags, ownership and PII tracking. Data contracts, which bundle schema, freshness, volume and column assertions, belong to the paid cloud product rather than the open-source core.

The split matters when choosing. The Apache 2.0 core is self-hosted on Kubernetes and needs Kafka, a relational database and Elasticsearch behind it. The cloud product adds an AI assistant, anomaly detection, compliance forms, access-request workflows and an uptime commitment. The vendor says more than 3,000 organisations run it, and names Netflix, Pinterest and Optum among its users.

Key features and capabilities

The same headings are used for every data catalogues & discovery entry, so two tools can be read side by side.

Metadata ingestion
  • Connector counts as published: 80+ production-grade connectors, 150+ systems in the directory
  • Ingestion from the user interface or the command line, using recipes
  • Push-based ingestion through APIs and Kafka
  • Captures schema, lineage, usage and profiling statistics
Search and discovery
  • Search across datasets, dashboards, ML models and raw files
  • Dataset profiling and data contracts surfaced in the interface
  • Cloud only: usage-aware search ranking and an AI assistant
  • Cloud only: AI-generated documentation and query-history mining
Lineage
  • Table-level, column-level and pipeline lineage
  • Extracted automatically by connectors that support it
  • Editable by hand in the interface, or emitted through the API
Glossary, classification and policy
  • Business glossary with term groups, inheritance and Git or API management
  • Tags, ownership and PII tracking
  • Cloud only: data contracts bundling schema, freshness, volume and column assertions
  • Cloud only: compliance forms, metadata tests and change proposals
Ownership and collaboration
  • Ownership assignment and documentation on assets
  • Slack, Google Drive, Notion and Confluence among the collaboration connectors
  • Cloud only: access-request workflows and change proposals
Integrations and APIs
  • Warehouses: Snowflake, BigQuery, Redshift, Databricks, Teradata, Microsoft Fabric
  • BI: Tableau, Looker, Power BI, Metabase, Superset
  • Orchestration and transformation: Airflow, Dagster, Prefect, dbt
  • APIs and SDKs, a Model Context Protocol integration, and Kafka for metadata events
How it runs
  • Self-hosted with Helm charts on Kubernetes
  • Needs Kafka, a relational database, Elasticsearch and a graph index
  • Managed cloud service with AWS PrivateLink; regions are not published

Pricing

Open source

The core is free under the Apache 2.0 licence. The managed cloud product publishes no prices: its pricing page is unavailable and every route leads to a demo, so treat it as quote-only. A trial without commitment is offered, and the vendor publishes a return-on-investment calculator instead of a price list.

Vendor pricing page →

Demos and videos

About DataHub (Acryl Data)

DataHub is developed by Acryl Data, Inc., which now trades under the DataHub name; the site still carries the Acryl copyright. It was co-founded by Swaroop Jagadish, previously at Airbnb, and Shirshanka Das, who built LinkedIn's metadata architecture and created the original project. The company was formed in 2021 with investment from LinkedIn, and is based in Palo Alto. It has raised $65m in total, including a $35m Series B in May 2025 led by Bessemer Venture Partners. The open-source project is Apache 2.0 but company-led rather than foundation-governed.

Palo Alto, California · datahub.com

Other data catalogues & discovery tools

Alation

Metadata Plane · Data Catalogues & Discovery

Enterprise data catalogue for search, discovery, stewardship and governance.

  • Commercial

Amundsen

Metadata Plane · Data Catalogues & Discovery

Data discovery tool built at Lyft, with search ranked by usage. Archived in September 2026 and no longer maintained; listed for reference and migration planning.

  • Open source

Atlan

Metadata Plane · Data Catalogues & Discovery

Active metadata platform for data cataloguing, discovery and governance.

  • Commercial

OpenMetadata

Metadata Plane · Data Catalogues & Discovery

Open-source metadata platform for discovery, lineage, data quality and collaboration.

  • Open core

Drafted with AI assistance and checked against the vendor’s own documentation.