Installation

Prerequisites

The package requires Python 3.10 or newer.

Install

pip install pdfexorcist

With uv, as a command-line tool or as a project dependency:

uv tool install pdfexorcist
uv add pdfexorcist

This installs four of the five default engines: pdfplumber, PyMuPDF, PDFium and camelot. The fifth, pdftotext, comes from poppler (see System tools).

Check what is installed

pdfexorcist engines
┏━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Engine     ┃ Installed ┃ Used      ┃ Reads        ┃ What it is                         ┃
┡━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ pdftotext  │ yes       │ default   │ text layer   │ poppler's pdftotext                │
│ pdfplumber │ yes       │ default   │ text layer   │ pdfminer word boxes                │
│ pymupdf    │ yes       │ default   │ text layer   │ MuPDF word boxes                   │
│ camelot    │ yes       │ default   │ text layer   │ camelot table finder               │
│ pdfium     │ yes       │ default   │ text layer   │ PDFium, Chrome's PDF engine        │
│ tabula     │ no        │ --engines │ text layer   │ tabula-java (needs Java)           │
│ tesseract  │ yes       │ --ocr     │ pixels (OCR) │ Tesseract OCR, CPU                 │
...
To install the missing engines (copy and paste):
  tabula: pip install "pdfexorcist[tabula]"

5 default engines ready: text PDFs need at least 3. 5 OCR engines ready: scans and photos
need at least 2.

The last lines print the exact command for each missing engine on your system. pdfexorcist doctor is the same command.

Optional engines

Each extra is named after the engine it adds; all adds every engine. See Engines for what each one reads.

Extra

Engine

Platform

chandra

Chandra OCR 2 model

GPU: CUDA or Apple Silicon

paddleocr

PP-OCR

Linux, macOS, Windows; CPU

ocrmac

Apple Vision

macOS

glmocr

GLM-OCR model via MLX

Apple Silicon Macs

tabula

tabula-java

needs Java

pip install "pdfexorcist[all]"
uv tool install "pdfexorcist[paddleocr,ocrmac]"

The names ocr, paddle, mac and mlx also work.

chandra, glmocr and paddleocr download their models on first use.

System tools

Three engines run a program that pip cannot install.

Engine

Program

macOS

Debian, Ubuntu

Windows

pdftotext

poppler

brew install poppler

sudo apt install poppler-utils

conda install -c conda-forge poppler

tesseract

Tesseract

brew install tesseract

sudo apt install tesseract-ocr

winget install UB-Mannheim.TesseractOCR

tabula

Java

brew install openjdk

sudo apt install default-jre

winget install Microsoft.OpenJDK.21

An engine whose program is missing is skipped with a warning. Without pdftotext, four default engines remain, and three of them must still agree.

PaddleOCR and OpenCV

PaddleOCR needs opencv-contrib-python as the only OpenCV wheel in the environment. camelot and mlx-vlm pull in opencv-python or opencv-python-headless. The wheels share one cv2 folder, and PaddleOCR then crashes with a segmentation fault. After installing, keep only one:

pip uninstall -y opencv-python opencv-python-headless
pip install --force-reinstall --no-deps opencv-contrib-python==4.10.0.84

From source

git clone https://github.com/saketlab/pdfexorcist
cd pdfexorcist
pip install -e ".[paddleocr]"