// the find
tuva-health/tuva-core
Main repo including core data model, data marts, data quality tests, and terminology sets.
Tuva Core is a dbt package that transforms raw healthcare claims and clinical data into a standardized analytics data model — encounters, cost, utilization, medications — across six warehouses (Snowflake, Databricks, BigQuery, Fabric, Redshift, DuckDB). It's built for healthcare data/analytics engineers doing population health, quality reporting, or claims analytics, not general dbt users.
The encounter-grain modeling is genuinely deep — 20+ distinct models for ED, SNF, hospice, ambulatory surgery, dialysis, home health, etc., the kind of domain work most in-house healthcare warehouses spend years building badly. Cross-database macros (safe_divide, try_to_cast_date, date_part) abstract vendor SQL differences across all six supported warehouses, and DuckDB support means you can run the full integration suite locally without touching a real warehouse. Data quality is dbt-native (YAML unit tests plus structural/logical checks) instead of bolted on, and the documented 1.0 contracts (nullable flag semantics, which 14 tables extension columns can flow through) cut down the usual ambiguity in shared data models. Terminology, value sets, and synthetic data are versioned independently from code and pulled from cloud storage, so you can pin a stable data snapshot without forking the package.
Installing it is not a five-minute dbt deps — you're wiring a root project with dozens of feature flags that must be native YAML booleans (env_var() strings are rejected), mapping source data into an Input Layer contract, and choosing a data-asset version correctly. SQL Server support has a real silent-failure trap: the default case-insensitive collation makes 'MALE' pass a logical DQ check written for 'male', and fixing it after the fact means rebuilding objects, not just changing a setting. Splitting Core from the feature packages (CMS HCC, quality measures, FHIR preprocessing, etc.) means you're manually tracking version compatibility across eight-plus separately released repos yourself, with no umbrella install. Production stability also leans on pinning immutable GitHub Releases rather than a stable Hub-registry path, and editing a 'released' data snapshot requires a named person's (Aaron's) explicit break-glass authorization — governance by individual, not by tooling.