NFHS-5 fact sheets¶
India’s National Family Health Survey 2019-21 (NFHS-5) publishes “Key Indicators” fact sheets from IIPS. Two recipes read them with one parser:
Example |
Reads |
|---|---|
|
The India and State/UT sheet: 131 indicators over four pages, NFHS-5 Urban, Rural and Total, and NFHS-4 Total. |
|
A State’s compendium: 104 indicators per district over three pages, NFHS-5 Total and NFHS-4 Total, or NFHS-5 Total alone where NFHS-4 has no comparable estimate. |
The parser is the NFHS-6 one adapted (see NFHS-6 fact sheets): rows by their printed indicator number, a wrapped label’s values on a line of their own, and a row yields values only when it holds exactly the page’s number of columns. The NFHS-5 sheets add:
Placeholders
(45.6)(small sample),*(suppressed) andna(not available), and numbers such as1,037(sex ratio) and3,385(Rs.). All are kept as printed.A value printed above its row’s baseline, which engines read before the row’s label (Kangra’s row 16).
Text outside the page box, never printed, right of the table (Raigarh’s page 3). The parser drops it.
A contents page with numbered rows (“1. North Goa 7”), and titles with an en dash or wrapped over two lines.
The fixtures are pages of the IIPS PDFs in tests/fixtures/nfhs5_state/ and tests/fixtures/nfhs5_district/.
Run it¶
pdfexorcist extract tests/fixtures/nfhs5_state/nfhs5_chandigarh_p3-6.pdf --recipe examples/nfhs5/state.toml -o chandigarh.csv
Reading tests/fixtures/nfhs5_state/nfhs5_chandigarh_p3-6.pdf with 5 engines (a majority must agree),
recipe examples/nfhs5/state.toml
pdftotext 476 readings 0.0s
pdfplumber 524 readings 0.5s
pymupdf 524 readings 0.0s
camelot 472 readings 1.1s
pdfium 68 readings 0.1s
Agreed 521 99.4% of cells
Unresolved 3 engines disagreed: left blank, not guessed
Checks 2 rules all passed
Wrote chandigarh.csv (table layout, 131 rows).
Cells to review, with the reason and every engine's reading: chandigarh.review.csv
geo,indicator_no,level,nfhs5_urban,nfhs5_rural,nfhs5_total,nfhs4_total
Chandigarh,1,state,86.8,(69.2),86.7,83.7
Chandigarh,6,state,93.6,*,93.6,na
Chandigarh,47,state,"5,586",*,"5,546","2,357"
Chandigarh,68,state,,*,,
Row 68 stays blank: camelot reads it with row 40’s numbers, and only two engines read it as printed. A district sheet, with the district recipe:
pdfexorcist extract tests/fixtures/nfhs5_district/nfhs5_kangra_p31-33.pdf --recipe examples/nfhs5/district.toml -o kangra.csv
Agreed 208 100.0% of cells
Unresolved 0
Checks 2 rules all passed
Wrote kangra.csv (table layout, 104 rows).
geo,indicator_no,level,nfhs5_total,nfhs4_total
"Kangra, Himachal Pradesh",3,district,"1,051","1,107"
"Kangra, Himachal Pradesh",16,district,1.5,2.4
"Kangra, Himachal Pradesh",49,district,(83.9),(68.6)
pdfexorcist show tests/fixtures/nfhs5_state/nfhs5_india_p3-6.pdf --recipe examples/nfhs5/state.toml -o page.png draws what the recipe reads:

pdfexorcist show tests/fixtures/nfhs5_district/nfhs5_kangra_p31-33.pdf --recipe examples/nfhs5/district.toml -o page.png draws what the recipe reads:

The recipe¶
name = "NFHS-5 State fact sheet"
[pages]
match = '[-–]\s*Key\s+Indicators' # "Bihar - Key Indicators", not the contents page
[engines]
ocr = false # a text layer: the five default engines, 3 must agree
[parser]
type = "custom"
function = "nfhs5_factsheet.py:parse"
key = ["page", "indicator_no", "column"]
[[checks]]
function = "nfhs5_factsheet.py:checks"
[output]
layout = "table"
rows = ["geo", "indicator_no"] # one row per geography and indicator
columns = "column" # nfhs5_urban, nfhs5_rural, nfhs5_total, nfhs4_total
format = "csv"
The district recipe differs only in its name and sample.
The key is the page, the printed indicator number and the column. The geography (
geo, from the page title) is an extra column.The checks allow no percentage over 100 (sex ratios, rates per 1,000, TFR and Rs. are not percentages) and require each State Total to lie between its Urban and Rural. TFR and the three child mortality rates are left out of the second rule: they are not ratios of sums, so their Total need not lie between.
Two district pages (Mahisagar and Wardha) skip the number 10 and print 11-32 for indicators 10-31, so “32.” appears on two pages. The vote keeps both, since its key holds the page; the table has one row for them, which it leaves blank and lists for review.
Checked against an independent extraction¶
tests/test_nfhs5_factsheet.py compares the vote with values that two extractions agree on: the vote, and pratapvardhan/NFHS-5’s CSVs. Each cell where they differ was read on the rendered page. On the full set (37 State sheets and 36 compendiums from nfhsiips.in):
State sheets |
District sheets |
|
|---|---|---|
Values agreed by the vote |
19,385 of 19,388 |
132,910 of 132,910 (705 districts) |
Compared with the CSV |
19,385 |
62,710 (its 341 districts) |
Equal |
19,258 |
62,635 |
CSV wrong |
0 |
10 (a missed value; 9 bracket notes on |
Page misprinted |
0 |
65 (the two misnumbered pages) |
CSV from another edition |
127 (literacy, December 2020) |
0 |
The parser handles each case this check caught it on: one-column pages, the raised value, text outside the page box, title variants and the contents page. The case-by-case table, sources and evidence crops are in tests/fixtures/nfhs5_state/VERIFICATION.md.
The NFHS-5 India Report (FR375) prints no fact sheets; its tables by State corroborate the 2021 literacy values. Each of its pages carries the facing page’s text outside its box, so read it with [pages] fit_boxes = false (or extract(..., fit_boxes=False)): growing the box would put two pages’ tables side by side. Its tables then need their own parser.
Sources¶
IIPS publishes the sheets on nfhsiips.in. The fixtures are pages of:
The independent extraction is pratapvardhan/NFHS-5 at commit 93c67fe: NFHS-5-States.csv and NFHS-5-Districts.csv.