Document intelligence: what it is, how it works, and API vs platform

For teams comparing document AI options, from cloud APIs to full platforms. See what document intelligence adds to OCR, where it breaks, and how to choose.

Document intelligence guide cover: a magnifying glass over a document picking out name, date and total fields

Key takeaways

  • Document intelligence is AI that turns PDFs, scans and photos into structured data, including text, key-value pairs, tables and the document type, instead of bare characters.
  • It's a category of software, not one product. Microsoft's Azure Document Intelligence is one service in it, and document AI and document understanding are other names for the same field.
  • Over plain OCR it adds layout analysis, key-value extraction, table reconstruction and classification.
  • A cloud API extracts. A document intelligence platform adds the workflow around it, such as checks across documents and a review queue. Choose by who owns the exceptions, engineering or operations.
  • Plan for layout drift, miscalibrated confidence, cross-document mismatches and hard handwriting before go-live.
On this page
  1. What is document intelligence?
  2. What is document understanding?
  3. How document intelligence differs from basic OCR
  4. Core capabilities, and where each one struggles
  5. The document intelligence workflow
  6. Where document intelligence fails
  7. API versus platform: when each makes sense
  8. Before you pick a tool
  9. Frequently asked questions

Document intelligence is AI that turns unstructured files, such as PDFs, scans and photos, into structured data, meaning the text, the key-value pairs, the tables and the type of document. It builds on OCR, which only returns characters, by working out how those characters relate, such as which label a value belongs to or which row a price sits in. It's a category of software rather than one product. Microsoft sells a cloud service called Azure Document Intelligence, and Google, Amazon and platforms such as Docsumo, our product, sell the same capability.

We start with the definition and its two sister terms, then show what it adds to OCR, where it breaks, and how to choose between a cloud API and a platform.

What is document intelligence?#

Picture an invoice with "Total Due: $4,120.50" at the bottom. OCR sees a line of text. Document intelligence knows the figure is the amount owed and returns it as the invoice total, ready for your accounting system, with no one typing it in.

Vendors describe the core the same way. Microsoft says its service extracts key-value pairs, text and tables and returns them as structured JSON; Google says Document AI turns unstructured document data into structured data. That's the extraction layer. A business process needs more around it. Files get sorted and split, values get checked against other documents, and unsure fields go to a person before the data moves on. That's where an API and a platform differ, covered below.

What is document understanding?#

Document understanding is what researchers call the same problem. The goal is software that reads a page's text, layout and visual cues together, the way a person does, rather than only the characters OCR returns, so it can classify the document, pull out its fields or answer questions about it. Microsoft researchers define document AI, or document intelligence, as techniques for "reading, understanding, and analyzing business documents", and models such as LayoutLM are tested on form understanding, receipt understanding and document classification. Some vendors use it as a product name. UiPath's Document Understanding extracts data from forms and invoices, and Google calls Document AI a document processing and understanding platform.

For a buyer, the names point to one category.

NameWhere you'll see it
Document intelligenceMicrosoft's Azure Document Intelligence, and platforms that cover the whole workflow
Document AIGoogle Cloud's Document AI, research papers, and the category as a whole: see what is document AI
Document understandingMachine-learning research, and UiPath's Document Understanding

For tools, see AI document understanding tools compared.

How document intelligence differs from basic OCR#

Feed OCR a receipt and you get a string of characters. Document intelligence adds four layers on top.

  • Layout analysis

    Finds headers, paragraphs, tables and reading order. More in layout detection.
  • Key-value pairs

    Links a label to its value, such as "Invoice Number" to "INV-2024-0892". More in key-value pair extraction.
  • Table reconstruction

    Keeps each price on the row of its item and quantity. More in table extraction.
  • Classification

    Works out whether a file is a W-2, a bank statement or a purchase order before extraction starts. More in document classification.
One invoice three ways: unstructured OCR text, numbered header, table and totals regions, and fields with a checked total
OCR returns the characters, layout analysis finds the header, table and totals in reading order, and extraction names each value and checks that line items plus tax equal the total.

OCR output usually needs custom parsing or a person before it's usable. Document intelligence output can go straight to your systems once it passes validation.

Core capabilities, and where each one struggles#

Most products share the same capabilities. They differ in how well each one holds up on hard documents.

CapabilityWorks well onStruggles with
Printed textClean scans, standard fontsLow-resolution images, decorative fonts
HandwritingBlock letters, consistent stylesCursive, poor legibility, mixed languages
TablesBordered tables, simple layoutsBorderless tables, merged cells, tables across pages
Key-value pairsLabels in consistent positionsImplied labels, variable layouts
ClassificationDistinct document typesHybrid documents, new formats

Most vendors offer prebuilt models for common documents, such as invoices, receipts, IDs and tax forms, and custom models trained on your own layouts. Handwriting goes through intelligent character recognition, and accuracy depends on legibility.

The document intelligence workflow#

Extraction is one stage. The stages around it decide whether the output is safe to use without a person checking every page.

  • Email attachments
  • Scans and uploads
  • API submissions
Document intelligence
  1. 01Classify and split
  2. 02Extract with confidence scores
  3. 03Validate across documents
  4. 04Review exceptions
Your systems, through an API and webhooks
How a document moves from intake to your systems

Each extracted value carries a confidence score. A 0.95 on an invoice total means the model is fairly sure, not certain. Validation then checks the values against rules and other documents. Does the total equal the line items? Does the PO number match an open purchase order? Values below their threshold and failed checks go to a person, whose corrections can improve the model. The full workflow is in automated document processing.

Where document intelligence fails#

Four problems come up again and again. Plan for them before go-live.

  • Layout driftA vendor redesigns its invoice and a model trained on the old layout pulls the wrong fields. Track accuracy by sender so the change shows up in days.
  • Miscalibrated confidenceA model can report 95% confidence on a wrong value, because confidence measures certainty, not correctness. Set thresholds per field on your own documents.
  • Cross-document mismatchesThe invoice says $10,000, the PO $9,500, and the receiving record shows 95 units at $100. Single-document extraction can't catch it; you need checks across the set.
  • Handwriting edge casesDoctors' notes, rushed signatures and forms filled in outside the boxes often need a person, whatever the model.

Aim for predictable automation rather than total automation. You want to know in advance which documents will flow through untouched and which will need a person.

API versus platform: when each makes sense#

Microsoft, Google and Amazon sell document intelligence as cloud APIs. You send a document and get structured data back. Document intelligence platforms add the layers above extraction. The choice usually comes down to who owns the problem.

Who owns the problem?

For operations, replacing data entry
A document intelligence platform such as Docsumo, our product: auto-classification and splitting (Business plan), cross-document validation and case management (Enterprise plan), a review queue for low-confidence fields, and delivery through its API and webhooks.
For engineering, building a feature
A cloud API such as Azure Document Intelligence, Google Document AI or Amazon Textract. You get JSON back; the validation rules, review workflow and integrations around it are mostly yours to build.

A platform pays off once exceptions outgrow a shared inbox, or when several document types feed one decision, like a loan file or a claim. Docsumo runs in the cloud only; if documents can't leave your own data center, a cloud API that also ships containers, as Azure's does, is the closer fit.

Our take. Most teams start by comparing extraction accuracy, and most vendors now do well on clean pages. The better question is who deals with the pages that aren't clean. If your engineers own those exceptions, an API is a fine start. If an operations team answers for the numbers, buy the review queue and the checks along with the model.

  • 99%field-level accuracy across 250+ document types
  • 95%+of documents processed straight through, without manual review

Before you pick a tool#

Whatever a vendor calls it, the job is to turn documents into data you can trust. Run a week of your own messiest files through each option and count how many come through untouched. That number, and who handles the rest, should decide it.

Book a demo with a few of your own documents, or start a free trial. Plans are on the pricing page.

Frequently asked questions#

What is document intelligence?

Document intelligence is AI that extracts text, key-value pairs, tables and structure from documents such as PDFs, scans and photos, and returns them as structured data. It builds on OCR by working out how the characters on a page relate to each other.

Is document intelligence the same as Azure Document Intelligence?

No. Document intelligence is the category, meaning any software that turns documents into structured data. Azure Document Intelligence in Foundry Tools, formerly Azure Form Recognizer, is Microsoft's cloud service in that category. Google Document AI, Amazon Textract and document intelligence platforms such as Docsumo are others. For a side-by-side, see Azure Document Intelligence alternatives.

Is there a Google Document Intelligence?

Google's product in the category is called Document AI. Different name, same category. Google sells Document AI, Microsoft sells Azure Document Intelligence, and researchers use both names, plus document understanding. More in what is document AI.

How much does document intelligence cost?

Cloud APIs charge per page. Microsoft's list prices for Azure Document Intelligence in US East, as of September 2026, run from $1.50 per 1,000 pages for reading text to $30 for custom extraction, with 500 free pages a month. Platforms usually charge by plan and volume. Docsumo offers a 14-day free trial on up to 1,000 pages; see pricing.

What's the difference between document intelligence and IDP?

The terms overlap. Document intelligence usually means the extraction and understanding layer, often sold as a cloud API. Intelligent document processing (IDP) usually means the whole workflow, including validation, review and delivery to your systems. See what is intelligent document processing.

Is document intelligence the same as OCR?

No. OCR turns an image of text into characters. Document intelligence also finds the layout, links each value to its label, rebuilds tables and identifies the document type. See IDP vs OCR.

See Docsumo read your own documents

Bring a few real samples. We'll show the fields extracted, the checks that ran and what a reviewer would see.