CRS State Tables 1-4

The annual Civil Registration System (CRS) report prints four State/UT tables of one shape: registered births, deaths, infant deaths and still births. Each row is a State/UT, then Male, Female and Person for Rural, Urban and Total.

The files are in examples/crs_state_tables/: recipe.toml and crs_state_table.py. The fixture holds Table 1 and Table 4 of the 2023 report, in tests/fixtures/crs/.

The recipe

name = "CRS State Tables 1-4"

[pages]
match = 'TABLE\s*-?\s*[1-4]\s*:\s*NUMBER OF .{0,30}REGISTERED BY SEX'

[engines]
ocr = false  # a text layer: the five default engines, 3 must agree

[parser]
type = "custom"
function = "crs_state_table.py:parse"
key = ["page", "state", "col"]

# ">=" because some States count Others/Not stated only in Person or Total
[[checks]]
column = "sex"
rule = "Person >= Male + Female"
by = ["page", "state", "area"]

[[checks]]
column = "area"
rule = "Total >= Rural + Urban"
by = ["page", "state", "sex"]

[output]
layout = "table"
rows = ["page", "label"]  # one row per table and State/UT
columns = "col"           # Rural_Male ... Total_Person
format = "csv"

match keeps only the four table pages of a full report. The parser reads rows as the generic parser does, then names each value’s area and sex, so the checks can use them.

parse() in crs_state_table.py, in outline (pseudocode):

def parse(pages):
    for page_no, lines in pages:
        seen = set()
        for label, values in lines split into text and values:
            skip a line without exactly 9 values
            no label: take the text lines just above and below, if they have no values
            a fake-bold label ("IIIInnnnddddiiiiaaaa") is read once ("India")
            state = the label in lower-case letters; skip it if seen on this page
            for col, value in zip(Rural_Male ... Total_Person, values):
                yield {"page", "state", "col", "label", "area", "sex", "value"}

Run it

pdfexorcist extract tests/fixtures/crs/crs_2023_t1_t4.pdf --recipe examples/crs_state_tables/recipe.toml -o crs_2023.csv
  pdftotext   666 readings  0.0s
  pdfplumber  666 readings  0.1s
  pymupdf     648 readings  0.0s
  camelot     648 readings  0.8s
  pdfium      648 readings  0.0s

Agreed          648  97.3% of cells
Unresolved       18  engines disagreed: left blank, not guessed
Checks      2 rules  all passed
page,label,Rural_Male,Rural_Female,Rural_Person,Urban_Male,Urban_Female,Urban_Person,Total_Male,Total_Female,Total_Person
1,India,,,,,,,,,
1,Andhra Pradesh,153272,137315,290587,244351,227155,471506,397623,364470,762093
1,Arunachal Pradesh,12565,12367,24932,9168,9297,18465,21733,21664,43397
...
1,Dadra and Nagar Haveli and Daman and Diu,1562,1353,2916,5149,4878,10032,6711,6231,12948
1,Lakshadweep,423,384,807,-,-,-,423,384,807

pdfexorcist show tests/fixtures/crs/crs_2023_t1_t4.pdf --recipe examples/crs_state_tables/recipe.toml -o page.png draws what the recipe reads:

CRS Table 1 page with its agreed values boxed and each row named by State

What is unusual

  • The India row is printed in fake bold: each glyph is drawn four times, a little offset. The engines read it differently (pdfplumber reads 5518639 as 5555555511118888666633339999), so its cells are unresolved and listed in the review file.

  • “Dadra and Nagar Haveli and Daman and Diu” wraps around its row, half above and half below. A row without a label takes the text lines on either side.

  • camelot returns some rows a second time. A State is printed once a page, so the parser keeps its first row.

Every agreed value equals crsindia’s PDF-verified data (tests/test_crs_state_table.py).