Engines and pages

pdfexorcist.register(name: str, default: bool = True, in_process: bool = False) → Callable[[Callable[[Path], Iterator[tuple[int, list[list[tuple[float, str]]]]]]], Callable[[Path], Iterator[tuple[int, list[list[tuple[float, str]]]]]]]

Add an extractor under name.

Parameters:
  • name – The engine’s name in methods and –engines.

  • default – Run it when no engines are named.

  • in_process – Never read in a worker process (a model held on the GPU).

pdfexorcist.EXTRACTORS: dict[str, Callable[[Path], Iterator[Page]]]

Every registered engine by name: the built-in ones and any added with register(). An engine takes a PDF path and yields (page_no, lines).

pdfexorcist.pages_matching(pdf: Path, start: str | None = None, stop: str | None = None, match: str | None = None, within: Sequence[int] | None = None) → list[int]

Return pages selected by case-insensitive text patterns.

Parameters:
  • pdf – PDF whose text is searched.

  • start – Pattern opening the page run.

  • stop – Pattern closing that run.

  • match – Pattern filtering selected pages.

  • within – Candidate 1-based page numbers.

pdfexorcist.parse_pages(spec: str, n_pages: int | None = None) → list[int]

Return unique, ordered pages from comma-separated ranges.

pdfexorcist.fit_page_boxes(pdf: Path, out_dir: Path) → Path

Copy of pdf whose page boxes cover all their text, or pdf itself.

Poppler and MuPDF clip text drawn outside a page box; pdfminer and PDFium read it, so without this the engines would read different tables.