Engines and pages¶
- pdfexorcist.register(name: str, default: bool = True, in_process: bool = False) Callable[[Callable[[Path], Iterator[tuple[int, list[list[tuple[float, str]]]]]]], Callable[[Path], Iterator[tuple[int, list[list[tuple[float, str]]]]]]]¶
Add an extractor under name.
- Parameters:
name – The engine’s name in methods and –engines.
default – Run it when no engines are named.
in_process – Never read in a worker process (a model held on the GPU).
- pdfexorcist.EXTRACTORS: dict[str, Callable[[Path], Iterator[Page]]]¶
Every registered engine by name: the built-in ones and any added with
register(). An engine takes a PDF path and yields(page_no, lines).
- pdfexorcist.pages_matching(pdf: Path, start: str | None = None, stop: str | None = None, match: str | None = None, within: Sequence[int] | None = None) list[int]¶
Return pages selected by case-insensitive text patterns.
- Parameters:
pdf – PDF whose text is searched.
start – Pattern opening the page run.
stop – Pattern closing that run.
match – Pattern filtering selected pages.
within – Candidate 1-based page numbers.
- pdfexorcist.parse_pages(spec: str, n_pages: int | None = None) list[int]¶
Return unique, ordered pages from comma-separated ranges.
- pdfexorcist.fit_page_boxes(pdf: Path, out_dir: Path) Path¶
Copy of pdf whose page boxes cover all their text, or pdf itself.
Poppler and MuPDF clip text drawn outside a page box; pdfminer and PDFium read it, so without this the engines would read different tables.