Parsers

pdfexorcist.parse_rows(pages: Iterable[tuple[int, list[list[tuple[float, str]]]]], min_values: int = 1) → Iterator[dict]

Default parser: one dict per value, keyed by KEY, plus the label text.

A trailing number in a label (“Population 2019”) is read as a value; pass a document parser to extract() when that matters.

pdfexorcist.parse_by_columns(pages: Iterable[tuple[int, list[list[tuple[float, str]]]]], min_values: int = 2, tol: float = 0.45) → Iterator[dict]

Like parse_rows, but a value’s column comes from its x.

Column centres are the median x of the i-th value over rows with the page’s most common value count; each value goes to the nearest centre within tol x the spacing, so a value an engine missed shifts nothing. Tokens in one column are joined (“1” “437” -> “1437”).

pdfexorcist.group_words(words: list[tuple], gap: float = 8.0) → list[list[tuple[float, str]]]

(x0, x1, y_bottom, text[, y_top]) words to lines of (x_centre, text) BoxCells.

A word joins a line when their vertical extents overlap by half the shorter one; a word without y_top is 1 pt tall; a word drawn twice (fake bold) counts once. Label words within gap pt join one cell; values never join.

pdfexorcist.KEY: list[str] = ["page", "row", "col"]

The default parser’s key: a cell is named by its page, its row label and its column number.