Pulling 7,760 numbers out of a printed PDF
Eleven pages of tables laid out three-up for print, turned into one spreadsheet — and then checked against the structure the document guarantees, so the result can be signed off without anyone hand-counting a single row.
Result
- 1,940
- rows extracted
- 7,760
- values captured
- 0
- gaps found
- 0.47s
- processing time
SourceIRS Publication 1040 (2025) — Tax and Earned Income Credit TablesPages 3-13. Public document — download it and check any row on this page against the printed original.
Proof
Checked against the document's own worked example
Printed in the PDF
“They find the $25,300-25,350 taxable income line, read down the married filing jointly column, and the amount shown is $2,562.”
Row in my output
income_at_least 25,300
income_but_less_than 25,350
married_filing_jointly 2,562
Expected 2,562. Matches exactly.
The problem
Why this is not a copy and paste job
- 01
Three tables share every line
The page prints three column blocks side by side, so one line of text holds three unrelated table rows. Read top to bottom, the numbers are nonsense.
- 02
Headers repeat on every page
Eleven pages of column headings sit in the middle of the data and have to be excluded without excluding real rows that look similar.
- 03
Band markers look like data
A lone figure sits above each group of rows as a heading. A parser that trusts numbers picks these up as records.
- 04
A worked example on page 2
The document explains itself with a sample table whose numbers are indistinguishable from real ones. Anything that reads that page ingests fake rows.
Verification
How I know nothing is missing
- 1,940
- Rows expected
- 1,940
- Rows found
- 0
- Duplicate bands
- 0
- Monotonicity breaks
From the income range and the band width
Every band accounted for
No row captured twice
Tax never falls as income rises
Coverage runs unbroken from $3,000 to $100,000. Thirty-three numeric lines were read and produced no row — those are the band markers, correctly ignored.
Sample
The extracted table
14 rows
| $3,000 | $3,050 | 303 | 303 | 303 | 303 | 3 |
| $4,000 | $4,050 | 403 | 403 | 403 | 403 | 3 |
| $5,000 | $5,050 | 503 | 503 | 503 | 503 | 3 |
| $6,000 | $6,050 | 603 | 603 | 603 | 603 | 3 |
| $7,000 | $7,050 | 703 | 703 | 703 | 703 | 3 |
| $8,000 | $8,050 | 803 | 803 | 803 | 803 | 3 |
| $9,000 | $9,050 | 903 | 903 | 903 | 903 | 3 |
| $10,000 | $10,050 | 1,003 | 1,003 | 1,003 | 1,003 | 3 |
| $11,000 | $11,050 | 1,103 | 1,103 | 1,103 | 1,103 | 3 |
| $12,000 | $12,050 | 1,205 | 1,203 | 1,205 | 1,203 | 4 |
| $13,000 | $13,050 | 1,325 | 1,303 | 1,325 | 1,303 | 4 |
| $14,000 | $14,050 | 1,445 | 1,403 | 1,445 | 1,403 | 4 |
| $15,000 | $15,050 | 1,565 | 1,503 | 1,565 | 1,503 | 4 |
| $16,000 | $16,050 | 1,685 | 1,603 | 1,685 | 1,603 | 4 |
Fourteen of 1,940 rows. The full file is below.
Judgement
The section I did not extract
$0 - $2,999
variable band widths ($5/$10/$25) and a broken PDF text layer that splits single rows across multiple lines; also shares the page with a worked 'Sample Table' example whose numbers mimic real rows.
Recommendation: manual entry, ~120 rows, or a second pass using positional extraction rather than text flow. Reporting this is worth more than a silent 95% — a client who is not told about the gap finds it themselves, later, in front of someone else.
Deliverables
The files a client receives
Have a file like this?
Send a sample and what you need it to look like. You will get a quote and a turnaround, not a discovery call.
