Extracting¶
- pdfexorcist.extract(pdf: Path | str, methods: list[str] | None = None, parse: Callable[[Iterable[tuple[int, list[list[tuple[float, str]]]]]], Iterator[dict]] = parse_rows, key: Sequence[str] = KEY, min_agree: int | None = None, checks: Sequence[Callable[[DataFrame], Series]] = (), min_unopposed: int | None = None, families: dict[str, list[str]] | None = None, normalize: Callable[[str], str | None] | None = None, pages: Sequence[int] | None = None, on_engine: Callable[[str, str, object], None] | None = None, fit_boxes: bool = True, return_readings: Literal[False] = False, jobs: int = 1) DataFrame¶
- pdfexorcist.extract(pdf: Path | str, methods: list[str] | None = None, parse: Callable[[Iterable[tuple[int, list[list[tuple[float, str]]]]]], Iterator[dict]] = parse_rows, key: Sequence[str] = KEY, min_agree: int | None = None, checks: Sequence[Callable[[DataFrame], Series]] = (), min_unopposed: int | None = None, families: dict[str, list[str]] | None = None, normalize: Callable[[str], str | None] | None = None, pages: Sequence[int] | None = None, on_engine: Callable[[str, str, object], None] | None = None, fit_boxes: bool = True, *, return_readings: Literal[True], jobs: int = 1) tuple[DataFrame, Readings]
Read pdf with each engine, parse, vote and check.
- Parameters:
pdf – A PDF, an image (.png, .jpg, …; name OCR engines in methods), or an http(s) URL, downloaded into $PDFEXORCIST_CACHE.
methods – Engines to run; default the text-layer engines.
parse – Turns (page_no, lines) into dicts carrying key and “value”; parse_by_columns places values by x.
key – Columns that name a cell.
min_agree – Engines that must agree; default a strict majority, at least 3. Missing or crashing engines do not lower it.
checks – Run on verified cells; they fill the failed column.
min_unopposed – Also accept a value this many engines read when none read another.
families – Engines that are not independent vote once per family. A text layer made by ocrmypdf makes the text engines one family.
normalize – Rewrites each value before the vote; None drops it.
pages – 1-based pages to read; result pages are renumbered 1..len(pages).
on_engine – Called as (name, event, detail) with “start”, “done” (readings) or “skipped” (reason).
fit_boxes – Grow page boxes over text drawn outside them; False when that text belongs to another page.
return_readings – Also return Readings: each engine’s readings, lines and boxes, and the PDF it read.
jobs – Worker processes; each built-in engine reads up to jobs chunks of pages. Chandra, GLM-OCR and your own engines read in this process. Same result as 1.
- Returns:
One row per key with its value, status and failed checks; with return_readings, (result, Readings).
- pdfexorcist.vote.vote(readings: DataFrame, key: Sequence[str] = KEY, min_agree: int | None = None, min_unopposed: int | None = None, families: dict[str, list[str]] | None = None) DataFrame¶
One row per key from one row per (method, key…, value) reading.
A value wins when at least min_agree methods read it and no other value ties it. A method that reads two values for one cell loses its vote there; extra columns take their most common reading.
- Parameters:
readings – One row per reading, with a method column.
key – Columns that name a cell.
min_agree – Methods that must agree; default a strict majority, at least 3.
min_unopposed – Also accept a value this many methods read when none read another.
families – Methods that are not independent; each family settles on its own majority, then votes once (min_agree then defaults to a majority of families, at least 2).