Parsers¶
- pdfexorcist.parse_rows(pages: Iterable[tuple[int, list[list[tuple[float, str]]]]], min_values: int = 1) Iterator[dict]¶
Default parser: one dict per value, keyed by KEY, plus the label text.
A trailing number in a label (“Population 2019”) is read as a value; pass a document parser to extract() when that matters.
- pdfexorcist.parse_by_columns(pages: Iterable[tuple[int, list[list[tuple[float, str]]]]], min_values: int = 2, tol: float = 0.45) Iterator[dict]¶
Like parse_rows, but a value’s column comes from its x.
Column centres are the median x of the i-th value over rows with the page’s most common value count; each value goes to the nearest centre within tol x the spacing, so a value an engine missed shifts nothing. Tokens in one column are joined (“1” “437” -> “1437”).
- pdfexorcist.group_words(words: list[tuple], gap: float = 8.0) list[list[tuple[float, str]]]¶
(x0, x1, y_bottom, text[, y_top]) words to lines of (x_centre, text) BoxCells.
A word joins a line when their vertical extents overlap by half the shorter one; a word without y_top is 1 pt tall; a word drawn twice (fake bold) counts once. Label words within gap pt join one cell; values never join.
- pdfexorcist.KEY: list[str] = ["page", "row", "col"]¶
The default parser’s key: a cell is named by its page, its row label and its column number.