Extracting

pdfexorcist.extract(pdf: Path | str, methods: list[str] | None = None, parse: Callable[[Iterable[tuple[int, list[list[tuple[float, str]]]]]], Iterator[dict]] = parse_rows, key: Sequence[str] = KEY, min_agree: int | None = None, checks: Sequence[Callable[[DataFrame], Series]] = (), min_unopposed: int | None = None, families: dict[str, list[str]] | None = None, normalize: Callable[[str], str | None] | None = None, pages: Sequence[int] | None = None, on_engine: Callable[[str, str, object], None] | None = None, fit_boxes: bool = True, return_readings: Literal[False] = False, jobs: int = 1) → DataFrame
pdfexorcist.extract(pdf: Path | str, methods: list[str] | None = None, parse: Callable[[Iterable[tuple[int, list[list[tuple[float, str]]]]]], Iterator[dict]] = parse_rows, key: Sequence[str] = KEY, min_agree: int | None = None, checks: Sequence[Callable[[DataFrame], Series]] = (), min_unopposed: int | None = None, families: dict[str, list[str]] | None = None, normalize: Callable[[str], str | None] | None = None, pages: Sequence[int] | None = None, on_engine: Callable[[str, str, object], None] | None = None, fit_boxes: bool = True, *, return_readings: Literal[True], jobs: int = 1) → tuple[DataFrame, Readings]

Read pdf with each engine, parse, vote and check.

Parameters:
  • pdf – A PDF, an image (.png, .jpg, …; name OCR engines in methods), or an http(s) URL, downloaded into $PDFEXORCIST_CACHE.

  • methods – Engines to run; default the text-layer engines.

  • parse – Turns (page_no, lines) into dicts carrying key and “value”; parse_by_columns places values by x.

  • key – Columns that name a cell.

  • min_agree – Engines that must agree; default a strict majority, at least 3. Missing or crashing engines do not lower it.

  • checks – Run on verified cells; they fill the failed column.

  • min_unopposed – Also accept a value this many engines read when none read another.

  • families – Engines that are not independent vote once per family. A text layer made by ocrmypdf makes the text engines one family.

  • normalize – Rewrites each value before the vote; None drops it.

  • pages – 1-based pages to read; result pages are renumbered 1..len(pages).

  • on_engine – Called as (name, event, detail) with “start”, “done” (readings) or “skipped” (reason).

  • fit_boxes – Grow page boxes over text drawn outside them; False when that text belongs to another page.

  • return_readings – Also return Readings: each engine’s readings, lines and boxes, and the PDF it read.

  • jobs – Worker processes; each built-in engine reads up to jobs chunks of pages. Chandra, GLM-OCR and your own engines read in this process. Same result as 1.

Returns:

One row per key with its value, status and failed checks; with return_readings, (result, Readings).

pdfexorcist.vote.vote(readings: DataFrame, key: Sequence[str] = KEY, min_agree: int | None = None, min_unopposed: int | None = None, families: dict[str, list[str]] | None = None) → DataFrame

One row per key from one row per (method, key…, value) reading.

A value wins when at least min_agree methods read it and no other value ties it. A method that reads two values for one cell loses its vote there; extra columns take their most common reading.

Parameters:
  • readings – One row per reading, with a method column.

  • key – Columns that name a cell.

  • min_agree – Methods that must agree; default a strict majority, at least 3.

  • min_unopposed – Also accept a value this many methods read when none read another.

  • families – Methods that are not independent; each family settles on its own majority, then votes once (min_agree then defaults to a majority of families, at least 2).