Observability Plane · Data Observability

Bigeye

Data observability platform with automated monitoring and lineage-based root cause analysis.

Overview

Bigeye now calls itself an enterprise AI trust platform, with data observability one of six modules alongside metadata management, lineage, data sensitivity, governance and an AI guardian.

Its vocabulary is its own rather than borrowed. Checks are metrics, organised into five dimensions: pipeline reliability, uniqueness, completeness, distributions and validity.

The mechanism differs from Monte Carlo's in a way worth understanding. A metric is a statistic calculated over the data and tracked as a time series, so Bigeye monitors the data itself rather than warehouse metadata alone.

Autometrics are profiling-driven suggestions, so onboarding is guided rather than zero-configuration: Bigeye profiles your tables, proposes metrics, and you deploy them per schema, table or column.

Autothresholds are the learned part, and the documentation is unusually open about how they work, running statistical tests on the series, blind-testing forecasting techniques, then adjusting for your sensitivity setting.

Two numbers to plan around: the default training window is 21 days, and seasonality needs three or more cycles of a pattern before it can be inferred.

Feedback is explicit and load-bearing. Closing an issue as a bad alert makes the autothreshold treat those points as expected, widening the band, which is a cleaner loop than most competitors document.

Its real differentiator is legacy and on-premises reach, with over fifty connectors spanning analytical and transactional databases, BI and ETL, including Informatica, SAS, Talend, Teradata and IBM DB2.

That strength has a source worth knowing: Data Advantage Group and its MetaCenter product joined Bigeye in June 2023, and MetaCenter's metadata collection now powers its column-level lineage.

Deltas, for comparing one dataset against another, are unique here and genuinely useful for validating a migration. Pricing is per table, published only in the terms of service, with no pricing page at all.

Key features and capabilities

The same headings are used for every data observability entry, so two tools can be read side by side.

What it monitors
  • Pipeline reliability at table level, covering freshness, volume, row count, hours since latest value and read queries
  • Completeness at column level, covering nulls, empty strings and not-a-number values
  • Distributions covering minimum, maximum, average, variance, skew, kurtosis, median, percentile and sum
  • Uniqueness and validity, the latter including UUID, email, phone, postcode and financial identifiers
  • More than 70 metric types, plus custom SQL rules, join rules and deltas between datasets
How incidents are detected
  • Metrics are statistics calculated over the data, so it monitors the data rather than metadata alone
  • Autothresholds run statistical tests on the series, blind-test forecasting techniques, then model uncertainty
  • The default training window is 21 days, with history backfilled where a row creation time exists
  • Seasonality needs three or more cycles of a pattern before it can be inferred
  • Closing an issue as a bad alert retrains the threshold to treat those points as expected
Coverage and onboarding
  • More than 50 connectors, working across on-premises and cloud, spanning analytical, transactional, BI and ETL
  • Onboarding is guided rather than zero-configuration; you deploy metrics from schema, table or column pages
  • Autometrics are profiling-driven suggestions, such as nulls on all columns or email validation above a match rate
  • Monitoring adapts as new tables, columns and jobs appear over time
  • Lineage-enabled monitoring deploys coverage upstream from a chosen asset
Triage and root cause
  • Issues with a unified inbox, and collections that group metrics and target notifications
  • Every alert carries lineage-aware context showing where the issue started and what depends on it
  • Root-cause tracing across sources, with impacted tables, columns and schemas listed
  • Downstream issues can be muted or closed together
  • Cost anomaly detection flags spend changes the morning after something changes
Lineage and impact
  • Fully automated column-level lineage for modern and legacy sources, with no manual mapping
  • Table level by default, expanding individual nodes to reveal column-level detail on demand
  • Covers transactional databases, warehouses, lakes, ETL platforms, BI tools and legacy systems
  • Three display modes, showing metric status, data sensitivity classification, or lineage alone
  • BI coverage across Tableau, Power BI, Looker and Qlik Sense, with a lineage API and CSV export
Integrations
  • Warehouses and databases including Snowflake, BigQuery, Redshift, Postgres, MySQL and Oracle
  • Teradata and IBM DB2 are reachable only through the self-hosted agent
  • Alerting to Slack, email, webhooks, Teams, Jira, ServiceNow, Azure Boards and PagerDuty
  • A REST API with a Python SDK, a CLI and service accounts for non-human identities
  • Single sign-on across Okta, Microsoft Entra and OpenID Connect; dbt and orchestrator integrations are not published as a list
Where it runs and what it costs
  • Agentless, connecting directly with allowlisted addresses and read-only service accounts
  • Agent-based through Docker or Kubernetes for restricted environments, unlocking Teradata and DB2
  • An agent orchestrator manages self-hosted agents and runs jobs from the interface
  • Regions are not published, and self-hosting the whole platform is not offered
  • Priced per table, with a unit price and overage per table beyond the agreed quantity

Pricing

Price on requestQuote only; no pricing page exists

Bigeye publishes no prices at all; its pricing page does not exist, and no tiers, figures, free tier or trial length appear on the marketing site. What is published sits in the legal documents: the unit of pricing is per table, at a unit price, with overage charged per table beyond the agreed quantity. A startup package exists on a 30-day rolling term capped at a year, billed monthly with a 99% uptime commitment, but its fees and eligibility are specified only on an order form, so everything requires a sales conversation.

Vendor pricing page →

Demos and videos

About Bigeye

Bigeye was founded in 2019 by Kyle Kirwan, its chief executive, and Egor Gryaznov, who met at Uber working on data pipelines, and operates remote-first from South San Francisco as Toro Data Labs, Inc. It is private, having raised $73.5m in total including a $5m investment from USAA announced in October 2024, with investors Costanoa, Sequoia and Coatue plus angels including Datadog's chief executive. It acquired Data Advantage Group in June 2023, bringing twenty years of enterprise cataloguing and the MetaCenter product that now powers its column-level lineage. It reports over 500 customers.

Founded 2019 · South San Francisco, California · bigeye.com

Other data observability tools

Elementary

Observability Plane · Data Observability

dbt-native data observability, with an open-source package and a cloud platform.

  • Open core

Monte Carlo

Observability Plane · Data Observability

Data observability platform that monitors freshness, volume, schema and quality across the data stack. Now trading as Monte Carlo AI, with montecarlodata.com redirecting, and extended to monitoring AI agents.

  • Commercial

Sifflet

Observability Plane · Data Observability

Data observability platform that combines monitoring, lineage and a data catalogue.

  • Commercial

Drafted with AI assistance and checked against the vendor’s own documentation.