Table extraction from PDFs and images: how it works and which method to use

For operations and engineering teams who need the rows and columns out of PDFs, scans and photos: how each method works, which tool fits which job, and how to test and check the output before it reaches your systems.

Illustration of an invoice line-item table extracted from a PDF into a structured table

Key takeaways

  • Table extraction finds a table on a page, works out its rows, columns and cells, and returns the contents as structured data such as CSV, Excel or JSON.
  • It happens in two stages: table detection (where is the table?) and table structure recognition (which cell does each value belong to?). Scans and images need OCR first.
  • The hard cases are merged cells, borderless layouts, tables that span pages, poor scans and a different layout from every sender.
  • Methods run from rules and templates to machine learning, deep learning and vision-language models. Free libraries such as Camelot, pdfplumber and PyMuPDF work well on native PDFs you control.
  • Extraction isn't finished until the table is validated: rows that add up to the total, balances that reconcile, and uncertain cells sent to a person.
On this page
  1. What is table extraction?
  2. How table extraction works
  3. Why tables are hard to extract
  4. Table extraction techniques
  5. How to extract a table from a PDF
  6. How to extract a table from an image
  7. How to validate extracted table data
  8. How to choose table extraction software
  9. How Docsumo extracts tables
  10. The bottom line
  11. Frequently asked questions

Table extraction is the process of finding a table in a PDF, scan or image, working out its rows, columns and cells, and returning the contents as structured data such as CSV, Excel or JSON. It runs in two stages, table detection and table structure recognition, with OCR first when the page is an image. This guide covers how it works, the methods, the free tools, and how to test and check the output.

What is table extraction?#

Invoices, bank statements, rent rolls and claim forms keep their most useful data in tables: line items, transactions, units, charges. A PDF only stores instructions such as "put this text here, draw this line there"; it has no idea what a row or a cell is. Table extraction rebuilds that structure, so tabular data locked in a PDF comes out as rows and columns with each value in the right cell.

An invoice table detected on the page (table detection), split into rows and columns (table structure recognition), and output as a clean table of description, price, quantity and amount (table extraction)

Table detection finds where each table sits on the page. Table structure recognition then finds the rows, columns and cells inside it, including headers and merged cells. What you start from changes the first step:

  • Native PDFs

    Made by software, with a text layer, so tools read the characters and their positions directly.
  • Scanned PDFs

    A picture of a page with no text layer. OCR reads the characters first, and its errors carry into the table.
  • Images and photos

    Screenshots, phone photos and faxes: a scan, plus skew, shadows and uneven light to correct.

How table extraction works#

Whatever the method, the pipeline is the same: get the text, find the table, rebuild its grid, then clean and check the result.

  • Native PDFs
  • Scanned PDFs
  • Images and photos
Table AI
  1. 01OCR if needed
  2. 02Detect the table
  3. 03Recognize rows and columns
  4. 04Map headers and clean rows
  5. 05Validate
CSV, Excel, JSON or your system
From a page to a usable table

Tools differ most in the last two stages: dropping subtotals and repeated headers, joining tables across pages, mapping columns to your fields and checking totals.

Why tables are hard to extract#

Five things break simple extractors:

  • Merged cells and nested headers

    A header that spans sub-columns, or a small table inside a cell, throws simple parsers off by a column or more.
  • Tables that span pages

    Page-by-page tools cut the rows on page 2 off from the header on page 1.
  • Borderless layouts

    Spacing replaces ruling lines, so tight columns run together.
  • Scan quality

    Blur, skew, rotated pages and low resolution cause OCR misreads that every later step inherits.
  • A different layout from every sender

    One vendor's invoice has 5 columns, the next has 8, and templates break as senders multiply.

Table extraction techniques#

There are four broad ways to do it, each built on the limits of the last:

  • Rules and templates

    Ruling lines, spacing and header words such as Qty and Amount mark out the table; every new layout needs new rules.
  • Machine learning

    A classifier learns which regions are tables from features an engineer designs, such as edges, text density and alignment.
  • Deep learning

    Networks learn the features themselves: CascadeTabNet (2020) finds tables and their cells in one model, and Microsoft's Table Transformer learned from nearly one million tables in scientific articles.
  • LLMs and vision-language models

    They read a table in context and can infer a missing header, but cost more and can return values that aren't on the page.
CriteriaRule-basedMachine learningDeep learningLLM and vision
Training dataNoneLabeled examplesLarge labeled setsNone to start; examples help
New layoutsNew rules each timeModerateGood, if trained on similar documentsGood
Main riskBreaks on layout changesFeature upkeepWeak on unfamiliar document typesValues that aren't on the page
Cost to runVery lowLowModerateHighest

In production they're often combined, with confidence scores and validation rules deciding which cells a person checks.

How to extract a table from a PDF#

For a one-off, Excel's Data > Get Data > From File > From PDF lists the tables in a PDF that has a text layer. In code, free Python libraries read that text layer well; scanned PDFs need OCR first, except where noted.

LibraryBest forScanned PDFs
CamelotRuled tables (lattice) and whitespace-separated tables (stream) in native PDFsOnly with its optional ML and OCR add-on
tabula-pySimple, well-structured tables; a Python wrapper for tabula-java (needs Java)No
pdfplumberDetailed access to characters, lines and tables through find_tables and extract_tablesNo; it has no OCR
PyMuPDFPage.find_tables(), with line- and text-based strategies and export to pandasNot on its own
img2tableTables in images and PDFs, including borderless tablesYes, with an OCR engine you choose

With a library, you build the multi-page joining, header mapping, validation and review yourself. For a worked pdfplumber example, see how to extract data from PDF.

How to extract a table from an image#

A photo or screenshot has no text layer, so OCR comes first. Pick the route that fits how often you do it:

  1. One table, nowExcel's Data > From Picture turns an image or a screenshot into cells and asks you to review them before inserting: that's where to catch misreads.
  2. A script you controlThe Python library img2table finds tables in images and PDFs, borderless ones included, with an OCR engine such as Tesseract, PaddleOCR or EasyOCR.
  3. Many images, every dayA document AI platform corrects skew, reads the table, scores each cell and sends the uncertain ones to a person.

Whatever the route, straighten the image, crop to the table and use the highest resolution you have.

How to validate extracted table data#

Extraction is half the job; validation is what makes the data safe to use.

  • Rows add up to the totalLine items that sum to 10,050 against an extracted total of 10,500 send the record to review before it reaches the ERP.
  • Balances reconcileOn a bank statement, the opening balance plus credits minus debits equals the closing balance.
  • Documents agreeInvoice lines match the purchase order; rent roll totals match the operating statement.
  • Every page is therePage counts and running totals confirm that a multi-page table isn't missing rows.
  • Uncertain cells go to a personHigh-confidence rows go through; the rest go to review.

How to choose table extraction software#

What are you extracting?

For line items and transactions at volume
Docsumo (our product) or another document AI platform: invoice line items, bank statement transactions, rent rolls for CRE underwriting, loss runs, bills of lading. Rows are checked, and uncertain cells go to a reviewer before the data reaches your systems through an API.

Before you commit, test the tool on your own files:

  1. Use your worst filesMerged headers, borderless tables, tables across three or more pages, rotated pages and poor scans. Clean samples make every tool look good.
  2. Score cells, not documentsCount right values and missed cells: a tool can find every table and still drop rows.
  3. Score the structureA right value in the wrong column is still wrong. The TEDS score compares the extracted table's structure and content with the original.
  4. Count rows on long tablesExtra rows mean repeated headers were kept; missing rows mean a page wasn't joined.
  5. Count the fixesHow many cells a person had to correct is the number that matters in production. More in how to measure OCR accuracy.

How Docsumo extracts tables#

Docsumo's Document AI extracts every line item and transaction row from documents such as invoices and bank statements, scans and handwritten text included. Tables that run across pages are joined into one table, with the headers mapped, and values the model is unsure about go to a reviewer, who can click each one to see its source line on the page. The rows reach your systems through the API and webhooks or an Excel export. It's built for business documents at volume, so for a single table from a research paper, a free library or Excel will do.

  • 99%field-level accuracy across 250+ document types
  • 95%+of documents processed straight through, without manual review

The bottom line#

Table extraction is detection plus structure recognition, with OCR first for scans and images. Free libraries handle native PDFs you control; at volume, confidence scores, validation and review make the output safe to post. Test on your messiest documents before you choose.

Book a demo with a few of your own documents, or start a free trial.

Frequently asked questions#

What is table extraction?

Table extraction finds the tables in a PDF, scan or image and turns them into structured rows and columns for a spreadsheet, database or business system. OCR alone returns the characters; table extraction also works out which row and column each value belongs to.

How do I extract a table from a PDF?

If the PDF has a text layer, Excel's Data > Get Data > From File > From PDF lists its tables, and Python libraries such as Camelot and pdfplumber do the same in code. A scanned PDF needs OCR first, and many PDFs from many senders call for a document AI platform with an API and a review queue.

How do I extract a table from an image or a scanned PDF?

Run OCR first, because an image has no text layer. Excel's Data from Picture handles a one-off; in code, img2table pairs an OCR engine with table detection. Send low-confidence cells to review, because scan quality drives accuracy.

How do I extract tables that span multiple pages?

Use a tool that joins tables across pages: it spots the continuation (the same columns, a repeated header or a row cut off at the page break), carries the header forward and drops the repeated header rows. To test one, count the rows: extra rows mean repeated headers were kept, and missing rows mean a page wasn't joined.

What is the best OCR for tables?

For business documents at volume, a document AI platform such as Docsumo (our product) adds checks and a review queue. For scans and photos you script yourself, cloud APIs such as Amazon Textract, Azure Document Intelligence and Google Document AI return tables as JSON. Native PDFs need no OCR at all: Camelot or pdfplumber reads the text layer.

What is a data extraction table?

In research, a data extraction table is the form a systematic review team fills in for each study it includes, recording the same details from every one, such as study design, participants, interventions and outcomes. It's a different thing from table extraction, which pulls tables out of documents.

See Docsumo read your own documents

Bring a few real samples. We'll show the fields extracted, the checks that ran and what a reviewer would see.