All work
Live data · public source

Pulling 7,760 numbers out of a printed PDF

Eleven pages of tables laid out three-up for print, turned into one spreadsheet — and then checked against the structure the document guarantees, so the result can be signed off without anyone hand-counting a single row.

Result

1,940
rows extracted
7,760
values captured
0
gaps found
0.47s
processing time

SourceIRS Publication 1040 (2025) — Tax and Earned Income Credit TablesPages 3-13. Public document — download it and check any row on this page against the printed original.

Proof

Checked against the document's own worked example

IRS worked example, printed on page 2 of the same PDF
Match

Printed in the PDF

They find the $25,300-25,350 taxable income line, read down the married filing jointly column, and the amount shown is $2,562.

Row in my output

income_at_least 25,300
income_but_less_than 25,350
married_filing_jointly 2,562

Expected 2,562. Matches exactly.

The problem

Why this is not a copy and paste job

The document was laid out to be read on paper, not parsed. Four things in it will quietly corrupt a naive extraction.
  1. 01

    Three tables share every line

    The page prints three column blocks side by side, so one line of text holds three unrelated table rows. Read top to bottom, the numbers are nonsense.

  2. 02

    Headers repeat on every page

    Eleven pages of column headings sit in the middle of the data and have to be excluded without excluding real rows that look similar.

  3. 03

    Band markers look like data

    A lone figure sits above each group of rows as a heading. A parser that trusts numbers picks these up as records.

  4. 04

    A worked example on page 2

    The document explains itself with a sample table whose numbers are indistinguishable from real ones. Anything that reads that page ingests fake rows.

Verification

How I know nothing is missing

The document guarantees a structure: income bands step by exactly $50 with no gaps, and tax owed never falls as income rises. Both are machine-checkable, which turns “looks right” into a pass or a fail.
1,940
Rows expected

From the income range and the band width

1,940
Rows found

Every band accounted for

0
Duplicate bands

No row captured twice

0
Monotonicity breaks

Tax never falls as income rises

Coverage runs unbroken from $3,000 to $100,000. Thirty-three numeric lines were read and produced no row — those are the band markers, correctly ignored.

Sample

The extracted table

Every thousandth band, from the delivered CSV. Sort any column, or filter by an income figure.

14 rows

$3,000$3,0503033033033033
$4,000$4,0504034034034033
$5,000$5,0505035035035033
$6,000$6,0506036036036033
$7,000$7,0507037037037033
$8,000$8,0508038038038033
$9,000$9,0509039039039033
$10,000$10,0501,0031,0031,0031,0033
$11,000$11,0501,1031,1031,1031,1033
$12,000$12,0501,2051,2031,2051,2034
$13,000$13,0501,3251,3031,3251,3034
$14,000$14,0501,4451,4031,4451,4034
$15,000$15,0501,5651,5031,5651,5034
$16,000$16,0501,6851,6031,6851,6034

Fourteen of 1,940 rows. The full file is below.

Judgement

The section I did not extract

Excluded

$0 - $2,999

variable band widths ($5/$10/$25) and a broken PDF text layer that splits single rows across multiple lines; also shares the page with a worked 'Sample Table' example whose numbers mimic real rows.

Recommendation: manual entry, ~120 rows, or a second pass using positional extraction rather than text flow. Reporting this is worth more than a silent 95% — a client who is not told about the gap finds it themselves, later, in front of someone else.

Deliverables

The files a client receives

Have a file like this?

Send a sample and what you need it to look like. You will get a quote and a turnaround, not a discovery call.