// the find
Trinkle23897/Fast-Poisson-Image-Editing
A fast poisson image editing implementation that can utilize multi-core CPU or GPU to handle a high-resolution image input.
A Python package (`fpie`) that solves the Poisson image editing problem with Jacobi iteration, offering eight backends from NumPy to CUDA, HIP, MPI and Taichi behind one CLI and one Python API. It is for people who need to blend high-resolution images or video frames and want to choose their hardware path, not for anyone expecting a polished end-user editor.
- The backend matrix is the main draw. The same solver interface runs on NumPy, Numba, GCC, OpenMP, CUDA, HIP, MPI and Taichi, so you can benchmark your own hardware against the others instead of relying on one headline speedup.
- Two solver layouts are implemented side by side. EquSolver relabels mask pixels into a compact index, while GridSolver keeps the 2D layout and saves four index reads per pixel when the mask covers the whole image. The README says GridSolver only beats EquSolver once `--grid-x` and `--grid-y` are tuned, and it states that caveat openly.
- The gradient modes (`src`, `avg`, `max`) are exposed as flags with before-and-after examples on the same inputs, so the PIE paper's variants can be compared directly rather than taken on faith.
- Benchmarks and the written course report are linked from the docs, and the project discloses its 15-618 origin, so the performance numbers can be checked against the methodology behind them.
- Jacobi is a slow-converging choice. The examples run 5,000 to 10,000 iterations, each one a sweep over the whole mask. Conjugate gradient or a multigrid preconditioner would reach the same residual in far fewer sweeps, so most of the parallel work speeds up a weak solver rather than cutting the work itself.
- The test suite pulls its eight input sets from other people's GitHub repos at run time (`tests/data.py`), and the README's results table is hosted the same way. Link rot or an offline machine breaks the example workflow, and nothing in the repo pins those inputs.
- Installation is a dependency matrix, not a `pip install`. NumPy, Numba and Taichi install with pip, but GCC, OpenMP, CUDA, HIP and MPI need a compiler toolchain, and macOS OpenMP needs `gcc-11` substituted for clang. Most people will hit a build failure before they reach the parallel backends the project is named for.