NTA NEET (UG) toppers

The National Testing Agency’s list of NEET (UG) 2026 re-examination toppers opens with “List of Top 138 candidates scoring equal to or more than 690 marks” on pages 1-6. Each row has a Sr. No., Application No., candidate name, gender, category, percentile, NEET rank and State. The pages are 300 dpi scans with no text layer, so only OCR engines can read them. Pages 7-15 hold shorter lists (female, male, EWS, SC … toppers) and a marks table.

The files are in examples/nta_toppers/: recipe.toml and nta_toppers.py. The fixture is pages 1 and 2, in tests/fixtures/nta/.

The recipe

name = "NEET (UG) 2026 top 138"

[engines]
use = ["paddleocr", "ocrmac", "glmocr"]

[parser]
type = "custom"
function = "nta_toppers.py:parse"
key = ["sr_no", "field"]

[[checks]]
function = "nta_toppers.py:rank_rises"

[[checks]]
function = "nta_toppers.py:percentile_falls"

[output]
layout = "table"
rows = ["sr_no"]
columns = "field"
format = "csv"

Each value is keyed by the printed Sr. No. and its column. Two of the three engines must agree.

Run it

pdfexorcist extract https://cdnbbsr.s3waas.gov.in/s37bc1ec1d9c3426357e69acd5bf320061/uploads/2026/07/20260716180970800.pdf \
  --recipe examples/nta_toppers/recipe.toml --pages 1-6 -o neet_top138.csv
  paddleocr  966 readings  254.9s
  ocrmac     928 readings  7.9s
  glmocr     845 readings  39.2s

Agreed          965  99.9% of cells
Unresolved        1  engines disagreed: left blank, not guessed
Checks      2 rules  all passed

A few columns of neet_top138.csv:

sr_no,rank,percentile,category,state
1,1,99.9999,General,PUNJAB
5,5,99.99965,OBC-NCL (Central List),MAHARASHTRA
25,25,99.9978999,Gen-EWS,GUJARAT
44,44,99.9976999,General,CHANDIGARH (UT)
...
138,138,99.9930996,General,DELHI

The unresolved cell is the Application No. of Sr. No. 89: only PaddleOCR reads it. Apple Vision reads that row’s percentile as 99.995499šę, so the parser skips the row, and GLM-OCR leaves out the column.

pdfexorcist show tests/fixtures/nta/neet_2026_top138_p1-2.pdf --recipe examples/nta_toppers/recipe.toml -o page.png draws what the recipe reads:

Scanned NEET toppers list: each row boxed and named by its Sr. No.

The parser and the checks

parse() in nta_toppers.py, in outline (pseudocode):

def parse(pages):
    for each line:
        a "List of ..." title other than the top candidates -> stop reading
        a row's line: a Sr. No. first, then a percentile and the rank after it
            place its tokens by what they are: the 12-digit Application No.,
            Male or Female, the percentile, the rank; the name lies between
            the Application No. and the gender, the category between the
            gender and the percentile, the State after the rank
            add the wrapped words held from the lines above, each to the
            column its x falls in
            yield {"sr_no", "field", "value"} for each column
        upper-case words, or the category's words -> hold as wrapped text
        anything else (letterhead, header, page number) -> drop what is held

rank_rises fails both ranks of two consecutive rows when the later rank is not higher. percentile_falls does the same when the later percentile is higher.

What is unusual

  • Every cell is bottom-aligned, so a wrapped name, “OBC-NCL (Central List)” or a two-word State sits on the lines above its row’s numbers. PaddleOCR and Apple Vision return those as lines of their own; GLM-OCR returns each row as one line.

  • The category is normalised from however it wrapped or was spaced (“OBC- NCL (Central List)”, “OBC-NCL (Central List)”).

  • Apple Vision reads some decimal points as commas (99,9930996); the parser reads them as points.

  • GLM-OCR leaves the Application No. column out of some pages and sometimes adds cells after the State (KERALA | Male | General). The parser stops the State at the first word that is not in capitals.

  • Tesseract is left out: it reads the table rules as text and runs rows together.

tests/test_nta_toppers.py compares every agreed value with a hand transcription of pages 1 and 2.