// the find
jonaswinkler/paperless-ng
A supercharged version of paperless: scan, index and archive all your physical documents
paperless-ng is a self-hosted pipeline that takes whatever your scanner dumps into a folder, OCRs it, indexes it for full-text search, and auto-tags it using matching rules it learns from your existing documents. It's for people who want their paper archive searchable without hand-labeling every page - but this specific repo stopped shipping in Feb 2023 and was handed off to the community fork paperless-ngx, so it's mostly of historical interest now.
The matching engine for tags, correspondents, and document types improves as your corpus grows, so the tagging workload drops over time instead of staying flat. Search goes beyond basic full-text: relevance ranking, highlighted matches, and a 'more like this' feature, which most self-hosted search setups don't bother with. Format support extends past PDFs/images to Office documents via Apache Tika, and the email ingestion rules let you move, flag, or delete source mail per account after consumption, not just dump everything in one inbox-wide rule.
This repo is archived in favor of paperless-ngx - anyone starting fresh should go there instead, which makes evaluating this fork mostly academic. Documents are stored unencrypted on disk by design, and the README is explicit that you should never run this on an untrusted host, so protecting sensitive scans (tax records, IDs) is entirely on your filesystem/OS, not the app. Resource usage is heavier than the use case implies - roughly 300MB RAM just to run the web server, more while consuming documents - for what's conceptually OCR plus an index. The REST API diverged from the original paperless project's, so anything built against the old API is a rewrite, not a drop-in upgrade.