// the find
grai-io/grai-core
Grai is a self-hosted data lineage tool that builds a column-level graph across warehouses, databases, and dbt/BI tools, then lets you run lineage-aware tests in CI so a schema change in one system doesn't silently break a downstream dashboard. It's aimed at data engineers running dbt-centric warehouse stacks who want lineage tracking without buying into a hosted SaaS.
The GitHub Actions integration is the actual differentiator — most lineage tools stop at visualization, this one runs impact tests against your lineage graph as a CI check on PRs. Connector coverage is wide (Snowflake, BigQuery, Redshift, Postgres, MySQL, MSSQL, dbt, Fivetran, flat files) and each one gets its own CI workflow in the repo, so they're actually exercised rather than just declared supported. The metadata layer is factored into a separate grai-schemas package, so connectors talk to a standardized schema instead of each rolling its own format.
Last push was 8 months ago with only 317 stars and 20 forks post-YC-launch, which for a company also selling a hosted 'Grai Cloud' product raises the question of whether the open-source path is still getting engineering attention or has been deprioritized in favor of the paid version. Self-hosting means running Django + Postgres + a React frontend plus whichever connector packages you need — a lot of surface area for what's conceptually a metadata graph. The README ships default login credentials (null@grai.io / super_secret) with no visible warning to rotate them before exposing an instance, which is the kind of thing that ends up live on the internet. Looker support is still alpha and BigQuery/Snowflake connectors depend on those vendors' metadata APIs staying stable, which is a maintenance burden that tends to lag in slower-moving projects.