// the find
jmcarpenter2/swifter
A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner
swifter is a pandas/modin extension that replaces .apply() with .swifter.apply(), picking the fastest execution path automatically — vectorized, pandas apply, or dask parallel — based on a quick benchmark on a sample of the data. It's for anyone who's tired of manually deciding when to parallelize a slow .apply() call.
The vectorize-first strategy is the right call: it tries to express the function as a vectorized pandas op before reaching for multiprocessing, which avoids the common mistake of parallelizing something that should've just been vectorized. It also covers groupby.apply, which most 'speed up pandas' libraries skip. The API is a genuine drop-in — swap .apply for .swifter.apply — so adoption cost is near zero, and modin interop is a nice bonus for anyone already on that stack.
The sampling mechanism that decides which backend to use is a real footgun: the README itself warns that functions with side effects on external variables will run twice (once during the sample, once for real) and silently corrupt state — that's the kind of bug that's invisible until it isn't. There's no guidance on the dataframe-size threshold where the overhead of sampling and backend dispatch actually pays off, so small-to-medium frames can end up slower than a plain .apply(). Last push was March 2024, so it's maintained but not actively evolving, and debugging through a dask worker when something goes wrong is a worse experience than the plain pandas error you started with.