Document intelligence: what it is, how it works, and API vs platform
For teams comparing document AI options, from cloud APIs to full platforms. See what document intelligence adds to OCR, where it breaks, and how to choose.

Key takeaways
- Document intelligence is AI that turns PDFs, scans and photos into structured data, including text, key-value pairs, tables and the document type, instead of bare characters.
- It's a category of software, not one product. Microsoft's Azure Document Intelligence is one service in it, and document AI and document understanding are other names for the same field.
- Over plain OCR it adds layout analysis, key-value extraction, table reconstruction and classification.
- A cloud API extracts. A document intelligence platform adds the workflow around it, such as checks across documents and a review queue. Choose by who owns the exceptions, engineering or operations.
- Plan for layout drift, miscalibrated confidence, cross-document mismatches and hard handwriting before go-live.
On this page
- What is document intelligence?
- What is document understanding?
- How document intelligence differs from basic OCR
- Core capabilities, and where each one struggles
- The document intelligence workflow
- Where document intelligence fails
- API versus platform: when each makes sense
- Before you pick a tool
- Frequently asked questions
Document intelligence is AI that turns unstructured files, such as PDFs, scans and photos, into structured data, meaning the text, the key-value pairs, the tables and the type of document. It builds on OCR, which only returns characters, by working out how those characters relate, such as which label a value belongs to or which row a price sits in. It's a category of software rather than one product. Microsoft sells a cloud service called Azure Document Intelligence, and Google, Amazon and platforms such as Docsumo, our product, sell the same capability.
We start with the definition and its two sister terms, then show what it adds to OCR, where it breaks, and how to choose between a cloud API and a platform.
What is document intelligence?#
Picture an invoice with "Total Due: $4,120.50" at the bottom. OCR sees a line of text. Document intelligence knows the figure is the amount owed and returns it as the invoice total, ready for your accounting system, with no one typing it in.
Vendors describe the core the same way. Microsoft says its service extracts key-value pairs, text and tables and returns them as structured JSON; Google says Document AI turns unstructured document data into structured data. That's the extraction layer. A business process needs more around it. Files get sorted and split, values get checked against other documents, and unsure fields go to a person before the data moves on. That's where an API and a platform differ, covered below.
What is document understanding?#
Document understanding is what researchers call the same problem. The goal is software that reads a page's text, layout and visual cues together, the way a person does, rather than only the characters OCR returns, so it can classify the document, pull out its fields or answer questions about it. Microsoft researchers define document AI, or document intelligence, as techniques for "reading, understanding, and analyzing business documents", and models such as LayoutLM are tested on form understanding, receipt understanding and document classification. Some vendors use it as a product name. UiPath's Document Understanding extracts data from forms and invoices, and Google calls Document AI a document processing and understanding platform.
For a buyer, the names point to one category.
| Name | Where you'll see it |
|---|---|
| Document intelligence | Microsoft's Azure Document Intelligence, and platforms that cover the whole workflow |
| Document AI | Google Cloud's Document AI, research papers, and the category as a whole: see what is document AI |
| Document understanding | Machine-learning research, and UiPath's Document Understanding |
For tools, see AI document understanding tools compared.
How document intelligence differs from basic OCR#
Feed OCR a receipt and you get a string of characters. Document intelligence adds four layers on top.
Layout analysis
Finds headers, paragraphs, tables and reading order. More in layout detection.Key-value pairs
Links a label to its value, such as "Invoice Number" to "INV-2024-0892". More in key-value pair extraction.Table reconstruction
Keeps each price on the row of its item and quantity. More in table extraction.Classification
Works out whether a file is a W-2, a bank statement or a purchase order before extraction starts. More in document classification.

OCR output usually needs custom parsing or a person before it's usable. Document intelligence output can go straight to your systems once it passes validation.
Core capabilities, and where each one struggles#
Most products share the same capabilities. They differ in how well each one holds up on hard documents.
| Capability | Works well on | Struggles with |
|---|---|---|
| Printed text | Clean scans, standard fonts | Low-resolution images, decorative fonts |
| Handwriting | Block letters, consistent styles | Cursive, poor legibility, mixed languages |
| Tables | Bordered tables, simple layouts | Borderless tables, merged cells, tables across pages |
| Key-value pairs | Labels in consistent positions | Implied labels, variable layouts |
| Classification | Distinct document types | Hybrid documents, new formats |
Most vendors offer prebuilt models for common documents, such as invoices, receipts, IDs and tax forms, and custom models trained on your own layouts. Handwriting goes through intelligent character recognition, and accuracy depends on legibility.
The document intelligence workflow#
Extraction is one stage. The stages around it decide whether the output is safe to use without a person checking every page.
- Email attachments
- Scans and uploads
- API submissions
- 01Classify and split
- 02Extract with confidence scores
- 03Validate across documents
- 04Review exceptions
Each extracted value carries a confidence score. A 0.95 on an invoice total means the model is fairly sure, not certain. Validation then checks the values against rules and other documents. Does the total equal the line items? Does the PO number match an open purchase order? Values below their threshold and failed checks go to a person, whose corrections can improve the model. The full workflow is in automated document processing.
Where document intelligence fails#
Four problems come up again and again. Plan for them before go-live.
- Layout driftA vendor redesigns its invoice and a model trained on the old layout pulls the wrong fields. Track accuracy by sender so the change shows up in days.
- Miscalibrated confidenceA model can report 95% confidence on a wrong value, because confidence measures certainty, not correctness. Set thresholds per field on your own documents.
- Cross-document mismatchesThe invoice says $10,000, the PO $9,500, and the receiving record shows 95 units at $100. Single-document extraction can't catch it; you need checks across the set.
- Handwriting edge casesDoctors' notes, rushed signatures and forms filled in outside the boxes often need a person, whatever the model.
Aim for predictable automation rather than total automation. You want to know in advance which documents will flow through untouched and which will need a person.
API versus platform: when each makes sense#
Microsoft, Google and Amazon sell document intelligence as cloud APIs. You send a document and get structured data back. Document intelligence platforms add the layers above extraction. The choice usually comes down to who owns the problem.
Who owns the problem?
A platform pays off once exceptions outgrow a shared inbox, or when several document types feed one decision, like a loan file or a claim. Docsumo runs in the cloud only; if documents can't leave your own data center, a cloud API that also ships containers, as Azure's does, is the closer fit.
Our take. Most teams start by comparing extraction accuracy, and most vendors now do well on clean pages. The better question is who deals with the pages that aren't clean. If your engineers own those exceptions, an API is a fine start. If an operations team answers for the numbers, buy the review queue and the checks along with the model.
- 99%field-level accuracy across 250+ document types
- 95%+of documents processed straight through, without manual review
Before you pick a tool#
Whatever a vendor calls it, the job is to turn documents into data you can trust. Run a week of your own messiest files through each option and count how many come through untouched. That number, and who handles the rest, should decide it.
Book a demo with a few of your own documents, or start a free trial. Plans are on the pricing page.
Frequently asked questions#
What is document intelligence?
Document intelligence is AI that extracts text, key-value pairs, tables and structure from documents such as PDFs, scans and photos, and returns them as structured data. It builds on OCR by working out how the characters on a page relate to each other.
Is document intelligence the same as Azure Document Intelligence?
No. Document intelligence is the category, meaning any software that turns documents into structured data. Azure Document Intelligence in Foundry Tools, formerly Azure Form Recognizer, is Microsoft's cloud service in that category. Google Document AI, Amazon Textract and document intelligence platforms such as Docsumo are others. For a side-by-side, see Azure Document Intelligence alternatives.
Is there a Google Document Intelligence?
Google's product in the category is called Document AI. Different name, same category. Google sells Document AI, Microsoft sells Azure Document Intelligence, and researchers use both names, plus document understanding. More in what is document AI.
How much does document intelligence cost?
Cloud APIs charge per page. Microsoft's list prices for Azure Document Intelligence in US East, as of September 2026, run from $1.50 per 1,000 pages for reading text to $30 for custom extraction, with 500 free pages a month. Platforms usually charge by plan and volume. Docsumo offers a 14-day free trial on up to 1,000 pages; see pricing.
What's the difference between document intelligence and IDP?
The terms overlap. Document intelligence usually means the extraction and understanding layer, often sold as a cloud API. Intelligent document processing (IDP) usually means the whole workflow, including validation, review and delivery to your systems. See what is intelligent document processing.
Is document intelligence the same as OCR?
No. OCR turns an image of text into characters. Document intelligence also finds the layout, links each value to its label, rebuilds tables and identifies the document type. See IDP vs OCR.
Sources
- Microsoft Learn: What is Azure Document Intelligence in Foundry Tools?
- Microsoft Learn: Azure Document Intelligence in Foundry Tools (formerly Azure AI Form Recognizer) FAQ
- Microsoft Azure: Azure Document Intelligence pricing
- Google Cloud: Document AI overview
- AWS: Amazon Textract features
- Cui, Xu, Lv and Wei (Microsoft Research): Document AI: Benchmarks, Models and Applications
- Xu et al. (Microsoft Research): LayoutLM: Pre-training of Text and Layout for Document Image Understanding
- UiPath: IXP and Document Understanding
First published . Last updated .