// the find
scambier/obsidian-text-extractor
A (companion) plugin to facilitate the extraction of text from images (OCR) and PDFs.
An Obsidian companion plugin that runs OCR on images and extracts text from PDFs, .docx and .xlsx files, so other plugins such as Omnisearch can search their contents. Extraction happens on demand: the first call does the work and stores the result as a small JSON file in the plugin folder. It suits desktop Obsidian users who keep a lot of scanned material in their vaults.
The API surface is small: extractText, canFileBeExtracted and isInCache. A calling plugin can check whether a file is extractable and whether a result is already cached before paying for an expensive extraction. Caching per file as JSON in the plugin directory means ordinary vault sync can carry the cache between devices. The extraction library lives in its own lib/ folder, built with Rollup, and the Obsidian wrapper is built separately with esbuild. The PDF and Office work runs in web workers, which keeps the UI thread free. Coverage goes beyond images and PDFs to docx and xlsx, and there is a separate macOS OCR module alongside Tesseract.
The author says the project is unmaintained, so any bug you hit, including the PDF failures, stays open unless you fork it. PDF extraction is the weak point: the README admits it often fails, pointing to issues #7 and #21, and PDFs are a core use case. Nothing works on mobile. The cache only helps if a desktop already processed the file, and on a phone an unprocessed file returns an empty string, which a caller may mistake for a document with no text. OCR language data downloads on demand, and the README says the plugin needs an internet connection to work at all, which rules out a fully offline vault.