Other recipe settings¶
The gallery recipes leave many settings at their defaults. Each recipe here changes a few of them and runs on a gallery fixture. The generic ones are in examples/settings/; the others sit next to the recipe they vary, so they can use its Python. tests/test_recipe_docs_examples.py runs each one. Every setting is listed in the format reference.
Rows or columns, and a total column¶
Table 1A.4 of Crime in India 2023 prints, for each State/UT, road-accident deaths three times: in total, by Hit and Run, and by Other Accidents. Each has Incidence, Victims and Rate, so total incidence equals Hit and Run incidence plus Other incidence. Dadra and Nagar Haveli and Daman and Diu wraps its name around its row, which leaves the serial number 31 at the start of the line.
examples/settings/cii_rows.toml reads the page with the rows parser and checks that total column:
name = "Crime in India 1A.4, rows parser"
sample = "../../tests/fixtures/cii/cii_2023_1A.4_p47.pdf"
[parser]
type = "rows"
min_values = 3
[[checks]]
column = "col"
total = 0 # the total and its parts, instead of a rule
parts = [3, 6]
op = "="
by = ["page", "row"] # one check per line
name = "Total I = Hit and Run I + Other I"
[output]
layout = "table"
format = "csv"
pdfexorcist extract tests/fixtures/cii/cii_2023_1A.4_p47.pdf --recipe examples/settings/cii_rows.toml -o cii_rows.csv
Agreed 352 97.5% of cells
Unresolved 9 engines disagreed: left blank, not guessed
Failed checks 3 agreed values that break a rule
x Total I = Hit and Run I + Other I: 3 cells
Wrote cii_rows.csv (table layout, 40 rows).
Cells to review, with the reason and every engine's reading: cii_rows.review.csv
Some agreed values break a rule (see the failed column), so the exit code is 1.
page,label,0,1,2,3,4,5,6,7,8,9
1,1 Andhra Pradesh,7488,8036,14.1,549,575,1.0,6939,7461,13.0,
...
1,,31,77,91,6.0,36,36,2.8,41,55,3.2
The rows parser numbers values in reading order, so on the unlabelled line 31 becomes column 0 and every value moves one column right. 31 is not 6.0 + 2.8, so those cells fail the check.
examples/settings/cii_columns.toml reads the same page with the columns parser, which puts each value in the column it sits under. Its parts that differ:
[parser]
type = "columns"
min_values = 3
tolerance = 0.6 # default 0.45; see below
[[checks]]
column = "col"
total = 1 # the serial number is now column 0
parts = [4, 7]
by = ["page", "row"]
missing = "fail" # a row with a total but without a part fails
name = "Total I = Hit and Run I + Other I"
[output]
layout = "table"
format = "xlsx"
pdfexorcist extract tests/fixtures/cii/cii_2023_1A.4_p47.pdf --recipe examples/settings/cii_columns.toml -o cii_columns.xlsx
pdftotext 387 readings 0.0s
pdfplumber 387 readings 0.1s
pymupdf 387 readings 0.0s
camelot 387 readings 0.3s
pdfium 387 readings 0.0s
Agreed 387 97.5% of cells
Unresolved 10 engines disagreed: left blank, not guessed
Checks 1 rule all passed
Wrote cii_columns.xlsx (table layout, 40 rows).
Cells to review, with the reason and every engine's reading: the review sheet.
The table sheet, as CSV:
page,label,0,1,2,3,4,5,6,7,8,9
1,Andhra Pradesh,1,7488,8036,14.1,549,575,1.0,6939,7461,13.0
...
1,,31,77,91,6.0,36,36,2.8,41,55,3.2
1,D&N Daman Haveli & Diu and,,,,,,,,,,
The serial numbers have a column of their own, and the unlabelled row’s values stay in theirs. Its name, read apart from its values, is a row of its own; only PyMuPDF reads it as a line with values, so its cells are unresolved and on the review sheet.
tolerance is how far, in column spacings, a value may sit from its column before it is dropped. pdftotext places text by character column, which is coarser than the other engines’ points: at the default 0.45 it drops a few values, at 0.6 none.
A rounded total¶
Column 3 of Crime in India 2021 Table 1A.1 is the mid-year population in lakhs, rounded to 0.1. The States and UTs add up to 13672.0, and the page prints 13671.8 for All India. examples/settings/cii_population.toml allows a relative gap and writes one row per cell:
name = "Crime in India 1A.1, population"
sample = "../../tests/fixtures/cii/cii_2021_1A.1_p43.pdf"
[pages]
select = "1"
[parser]
type = "rows"
min_values = 3
[[checks]]
column = "label"
rule = "TOTAL ALL INDIA = rest"
by = ["page", "col"]
where = "col == 3 and label != 'TOTAL STATE(S)' and label != 'TOTAL UT(S)'"
tolerance = 0.0001 # 0.01% of the total: 1.4 lakh here
[output]
layout = "cells"
format = "json"
pdfexorcist extract tests/fixtures/cii/cii_2021_1A.1_p43.pdf --recipe examples/settings/cii_population.toml -o population.json
Agreed 234 100.0% of cells
Unresolved 0
Checks 1 rule all passed
Wrote population.json (cells layout).
One record:
{
"page": 1,
"row": "totalallindia#1",
"col": 3,
"label": "TOTAL ALL INDIA",
"value": "13671.8",
"status": "verified",
"failed": "",
"n_agree": 5,
"votes": "13671.8×5",
"sources": "camelot|pdfium|pdfplumber|pdftotext|pymupdf",
"dissent": ""
}
Without tolerance the rule fails on the 0.2 gap. abs_tolerance allows a fixed gap instead, as in the population projections.
Pages by number¶
examples/crs_state_tables/table4.toml is the CRS recipe with [pages] select = "2": the fixture holds Table 1 on page 1 and Table 4 (still births) on page 2.
[pages]
select = "2"
pdfexorcist extract tests/fixtures/crs/crs_2023_t1_t4.pdf --recipe examples/crs_state_tables/table4.toml -o table4.csv
Reading page 2 of tests/fixtures/crs/crs_2023_t1_t4.pdf with 5 engines (a majority must agree),
recipe examples/crs_state_tables/table4.toml
...
Agreed 324 97.3% of cells
Unresolved 9 engines disagreed: left blank, not guessed
Checks 2 rules all passed
page,label,Rural_Male,Rural_Female,Rural_Person,Urban_Male,Urban_Female,Urban_Person,Total_Male,Total_Female,Total_Person
2,India,,,,,,,,,
2,Andhra Pradesh,435,390,825,589,583,1172,1024,973,1997
The output keeps the page’s number in the original file. The unresolved cells are the India row, as in the full run.
Page boxes¶
examples/mccd/unfitted.toml is the MCCD Table 4 recipe with [pages] fit_boxes = false, run on the 2011 fixture: a two-page spread drawn half outside its page box. By default pdfexorcist widens each page box to cover its text before the engines read it. Without that:
[pages]
fit_boxes = false
pdfexorcist extract tests/fixtures/mccd/mccd_2011.pdf --recipe examples/mccd/unfitted.toml -o unfitted.csv
pdftotext 1,782 readings 0.0s
pdfplumber 1,782 readings 0.2s
pymupdf 1,782 readings 0.0s
camelot 3,564 readings 11.6s
pdfium 1,782 readings 0.1s
Agreed 1,782 50.0% of cells
Unresolved 1,782 engines disagreed: left blank, not guessed
Failed checks 768 agreed values that break a rule
x state 'All States (Total)' >= the rest, within each page, code, pocc, sex: 768 cells
Note: 1,044 agreed cells were read with different values in different places (e.g. two pages); left
blank in the table and listed for review.
Four engines read only the half inside the box; camelot reads all of it. Half the cells have no majority, and on page 2 the clipped reading starts at another column, so its values are named after the wrong States and the All States check fails. With the default, every cell agrees. Set fit_boxes = false only when the text outside a box belongs to another page, as in the NFHS-5 India Report (NFHS-5 fact sheets).
Every engine must agree¶
examples/nfhs6/unanimous.toml is the NFHS-6 State recipe with these [engines] settings:
[engines]
use = ["pdftotext", "pdfplumber", "pymupdf", "camelot", "pdfium"]
min_agree = 5
min_unopposed = 3
min_agree = 5 alone, given here on the command line with -k 5:
pdfexorcist extract tests/fixtures/nfhs6_state/nfhs6_chandigarh_p145-147.pdf --recipe examples/nfhs6/state.toml -k 5 -o chandigarh.csv
Reading tests/fixtures/nfhs6_state/nfhs6_chandigarh_p145-147.pdf with 5 engines (5 must agree),
recipe examples/nfhs6/state.toml
pdftotext 404 readings 0.0s
pdfplumber 404 readings 0.4s
pymupdf 404 readings 0.0s
camelot 384 readings 0.8s
pdfium 276 readings 0.0s
Agreed 260 64.4% of cells
Unresolved 144 engines disagreed: left blank, not guessed
Checks 2 rules all passed
pdfium and camelot read nothing for some values, so those cells cannot reach 5. min_unopposed = 3 also accepts a value that 3 engines read when no engine read another:
pdfexorcist extract tests/fixtures/nfhs6_state/nfhs6_chandigarh_p145-147.pdf --recipe examples/nfhs6/unanimous.toml -o chandigarh.csv --force
Agreed 404 100.0% of cells
Unresolved 0
Checks 2 rules all passed
A value two engines read differently still needs all five. use names the engines, so one that is not installed is an error instead of a smaller vote.
Engines that vote once¶
pdfplumber and camelot both read the text layer with pdfminer, so a pdfminer misreading would count twice. examples/nfhs5/families.toml is the NFHS-5 State recipe with:
[engines]
families = { pdfminer = ["pdfplumber", "camelot"] }
The two settle on one value first, and that value counts as one vote. Five engines make four voters, and 3 must agree.
pdfexorcist extract tests/fixtures/nfhs5_state/nfhs5_chandigarh_p3-6.pdf --recipe examples/nfhs5/families.toml -o chandigarh.csv
pdftotext 476 readings 0.0s
pdfplumber 524 readings 0.4s
pymupdf 524 readings 0.0s
camelot 472 readings 1.0s
pdfium 68 readings 0.1s
Agreed 477 91.0% of cells
Unresolved 47 engines disagreed: left blank, not guessed
Checks 2 rules all passed
Most cells lost against the plain run were read only by pdfplumber, camelot and PyMuPDF: two independent readings, not three. On row 126, pdfplumber and camelot read different values, so the family casts no vote. Families matter most for a PDF whose text layer came from OCR, where every text engine repeats one reading. See Engines.
Cleaning values¶
examples/nfhs5/numbers.toml is the NFHS-5 State recipe with numbers cleaned before the vote and Parquet output:
[normalize]
remove_commas = true # "5,586" -> "5586"
function = "clean.py:drop_na" # "na" -> no reading, so a blank cell
[output]
format = "parquet"
clean.py, next to the recipe:
def drop_na(value: str) -> str | None:
"""Return the value to vote on, or None to drop the reading."""
return None if value == "na" else value
pdfexorcist extract tests/fixtures/nfhs5_state/nfhs5_chandigarh_p3-6.pdf --recipe examples/nfhs5/numbers.toml -o chandigarh.parquet
Agreed 490 99.4% of cells
Unresolved 3 engines disagreed: left blank, not guessed
Checks 2 rules all passed
Wrote chandigarh.parquet (table layout, 131 rows).
Cells to review, with the reason and every engine's reading: chandigarh.review.parquet
Read back with pandas (pd.read_parquet("chandigarh.parquet")), rows 3, 6 and 47:
geo indicator_no level nfhs5_urban nfhs5_rural nfhs5_total nfhs4_total
Chandigarh 3 state 918 868 917 934
Chandigarh 6 state 93.6 * 93.6 NaN
Chandigarh 47 state 5586 * 5546 2357
The na readings are gone; brackets and * are kept. The values stay text in the Parquet file.
Printed numbers that need cleaning¶
tests/fixtures/synthetic/cleaning.pdf is a small synthetic table (make_fixtures.py next to it builds it). It prints numbers the way some reports do: raised-dot decimals (65·3), Unicode minus signs (−4·0), a footnote mark (13·4*) and a small-sample value in brackets ((7·0)). examples/settings/cleaning.toml cleans them before the vote and checks the table after it:
[normalize]
raised_dots = true # "65·3" -> "65.3"
minus_signs = true # "−4.0" -> "-4.0"
remove_footnote_marks = true # "13.4*" -> "13.4"
[parser]
type = "rows"
[postprocess]
function = "numbers.py:add_number" # adds `number`: "(7.0)" -> 7.0
[[checks]]
column = "col"
total = 2
parts = [0, 1]
by = ["page", "row"]
value = "number"
name = "Total = A + B"
[[checks]]
column = "label"
rule = "All regions = rest"
by = ["page", "col"]
value = "number"
numbers.py, next to the recipe:
def add_number(cells: pd.DataFrame, pdf: object) -> pd.DataFrame:
text = cells["value"].astype(str).str.strip("()")
return cells.assign(number=pd.to_numeric(text, errors="coerce"))
pdfexorcist extract tests/fixtures/synthetic/cleaning.pdf --recipe examples/settings/cleaning.toml -o cleaning.csv
Agreed 12 100.0% of cells
Unresolved 0
Checks 2 rules all passed
page,label,0,1,2
1,North,65.3,-4.0,61.3
1,South,12.5,(7.0),19.5
1,East,500.0,13.4,513.4
1,All regions,577.8,16.4,594.2
The checks read number (value = "number"): (7.0) is not a number, so a check on value would leave the South row unchecked. The cell itself keeps its brackets.
OCR for an image¶
examples/settings/cleaning_ocr.toml reads the same table from an image, cleaning.png. An image has no text layer, so ocr = true adds every installed OCR engine; a majority of them, at least 2, must agree:
[engines]
ocr = true
[parser]
type = "columns" # OCR engines place words by x; values go to the nearest column
pdfexorcist extract tests/fixtures/synthetic/cleaning.png --recipe examples/settings/cleaning_ocr.toml -o cleaning.csv
Reading tests/fixtures/synthetic/cleaning.png with 5 engines (3 must agree), recipe
examples/settings/cleaning_ocr.toml
tesseract 0 readings 0.2s
chandra 12 readings 0.0s
paddleocr 12 readings 6.0s
ocrmac 8 readings 26.2s
glmocr 12 readings 3.0s
tesseract read no values on these pages.
Agreed 12 85.7% of cells
Unresolved 2 engines disagreed: left blank, not guessed
Checks 1 rule all passed
Every cell of the table is agreed and equals the PDF’s. The unresolved cells are a row only one engine read (East 13.4.). The engines depend on what is installed (pdfexorcist engines); the lake photo names its OCR engines with use instead.