Installation¶
Prerequisites¶
The package requires Python 3.10 or newer.
Install¶
pip install pdfexorcist
With uv, as a command-line tool or as a project dependency:
uv tool install pdfexorcist
uv add pdfexorcist
This installs four of the five default engines: pdfplumber, PyMuPDF, PDFium and camelot. The fifth, pdftotext, comes from poppler (see System tools).
Check what is installed¶
pdfexorcist engines
┏━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Engine ┃ Installed ┃ Used ┃ Reads ┃ What it is ┃
┡━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ pdftotext │ yes │ default │ text layer │ poppler's pdftotext │
│ pdfplumber │ yes │ default │ text layer │ pdfminer word boxes │
│ pymupdf │ yes │ default │ text layer │ MuPDF word boxes │
│ camelot │ yes │ default │ text layer │ camelot table finder │
│ pdfium │ yes │ default │ text layer │ PDFium, Chrome's PDF engine │
│ tabula │ no │ --engines │ text layer │ tabula-java (needs Java) │
│ tesseract │ yes │ --ocr │ pixels (OCR) │ Tesseract OCR, CPU │
...
To install the missing engines (copy and paste):
tabula: pip install "pdfexorcist[tabula]"
5 default engines ready: text PDFs need at least 3. 5 OCR engines ready: scans and photos
need at least 2.
The last lines print the exact command for each missing engine on your system. pdfexorcist doctor is the same command.
Optional engines¶
Each extra is named after the engine it adds; all adds every engine. See
Engines for what each one reads.
Extra |
Engine |
Platform |
|---|---|---|
|
Chandra OCR 2 model |
GPU: CUDA or Apple Silicon |
|
PP-OCR |
Linux, macOS, Windows; CPU |
|
Apple Vision |
macOS |
|
GLM-OCR model via MLX |
Apple Silicon Macs |
|
tabula-java |
needs Java |
pip install "pdfexorcist[all]"
uv tool install "pdfexorcist[paddleocr,ocrmac]"
The names ocr, paddle, mac and mlx also work.
chandra, glmocr and paddleocr download their models on first use.
System tools¶
Three engines run a program that pip cannot install.
Engine |
Program |
macOS |
Debian, Ubuntu |
Windows |
|---|---|---|---|---|
|
poppler |
|
|
|
|
Tesseract |
|
|
|
|
Java |
|
|
|
An engine whose program is missing is skipped with a warning. Without pdftotext, four default engines remain, and three of them must still agree.
PaddleOCR and OpenCV¶
PaddleOCR needs opencv-contrib-python as the only OpenCV wheel in the environment. camelot and mlx-vlm pull in opencv-python or opencv-python-headless. The wheels share one cv2 folder, and PaddleOCR then crashes with a segmentation fault. After installing, keep only one:
pip uninstall -y opencv-python opencv-python-headless
pip install --force-reinstall --no-deps opencv-contrib-python==4.10.0.84
From source¶
git clone https://github.com/saketlab/pdfexorcist
cd pdfexorcist
pip install -e ".[paddleocr]"