// the find
Daniel15568/percentify
A pip-installable library of 17 one-call statistics and data-quality helpers that accept either pandas or Polars objects and return the same type. It is aimed at people doing exploratory analysis who want a quick answer to a specific question, not a full profiling suite.
- profiler() returns a ranked findings table with a suggested fix for each issue, and an errors subset that can be asserted on in CI. That turns a describe() dump into something a pipeline can act on.
- Type-preserving across backends: a Polars input comes back as Polars, so there are no conversion shims in user code. The tests/test_polars.py file suggests this is covered by tests, not just claimed.
- Each function names the underlying library it wraps (scipy, statsmodels, scikit-learn), so the convenience layer doesn't hide the math from anyone who wants to check it.
- The note that has_missing is computed before rounding is the right kind of detail. It stops a single missing row from showing as 0.00 and disappearing.
- The README doesn't say how the 0 to 100 health score is computed. The CI gate in the examples (health >= 80) is only as meaningful as that formula, and nothing in the docs explains what moves it.
- The install line lists only numpy and pandas, but the README says several functions wrap scipy, statsmodels and scikit-learn. Unless those are optional extras, the install instructions are incomplete. Check pyproject.toml before adopting.
- Polars is described as first-class, but every example in the README uses pandas. A short Polars example in the README would back that claim better than the tip box does.
- The 'Before' screenshot in the profiler section is hosted under a different GitHub account (dmitriy1ikobe) than the repo itself (Ad-meliorael). It is probably a leftover from a fork, and it makes the README look less maintained than the code appears to be.