Skip to content
AI & Automation

AI Document Extraction: How to Extract Structured Data From Documents

How to extract structured data from documents with AI: OCR and layout parsing, templates versus ML versus language and vision models, schema design, structured outputs, tables, confidence, validation and evaluation.

Quick answer

AI document extraction turns documents into structured data in four steps: get the text and layout (native PDF text, OCR or a vision model), extract fields against an explicit schema (with templates, trained models or language and vision models using structured outputs), validate every field (formats, totals, cross-checks, nulls instead of guesses) and route low-confidence or failed fields to human review. Modern language and vision models handle varied layouts well, but they can invent values, so validation and field-level evaluation on your own documents are essential.

Where This Fits

This guide covers the extraction technique. The end-to-end business pipeline (ingestion, review, integration, storage) is in intelligent document processing, an applied example in AI invoice processing, and how extraction fits into automations in AI workflow automation.

Step 1: Get Text and Layout

Native digital PDFs usually contain text you can extract directly, though reading order and tables may need layout analysis. Scans and photos need OCR, ideally layout-aware so tables, columns and key-value pairs are preserved. Vision-capable models can read page images directly, which helps with complex layouts, stamps and handwriting, at a higher cost per page. Keep page and position references so reviewers can see where each value came from.

Step 2: Choose an Extraction Method

MethodHow it worksBest for
Templates and rulesFixed positions or anchors per layoutFew, stable layouts at high volume
Trained ML extractionModels trained on labelled examplesCommon document types with labelled data
Language model on textSchema prompt over OCR or PDF textVaried layouts, text-heavy documents
Vision-language modelSchema prompt over page imagesComplex layouts, tables, stamps, handwriting
HybridDifferent methods per document typeMixed real-world inputs

Step 3: Design the Schema

The schema is the contract between the document and your systems. Use clear names and types, formats (dates, currency codes), enums for categorical values, and arrays for line items. Make missing values explicitly nullable and instruct the model to return null rather than guess. Include a source quote or page reference per field where your provider supports it, to make review and validation easier.

Example: extraction schema for a delivery note (illustrative)
{
  "type": "object",
  "properties": {
    "delivery_note_number": { "type": "string" },
    "delivery_date": { "type": ["string", "null"], "format": "date" },
    "supplier_name": { "type": "string" },
    "po_number": { "type": ["string", "null"] },
    "lines": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "sku": { "type": ["string", "null"] },
          "description": { "type": "string" },
          "quantity": { "type": "number" },
          "unit": { "type": "string", "enum": ["each", "box", "pallet", "kg", "other"] }
        },
        "required": ["sku", "description", "quantity", "unit"],
        "additionalProperties": false
      }
    }
  },
  "required": ["delivery_note_number", "delivery_date", "supplier_name", "po_number", "lines"],
  "additionalProperties": false
}

Step 4: Use Structured Outputs

Major model providers can constrain output to a JSON schema (OpenAI's structured outputs and Anthropic's structured outputs, for example), which eliminates parsing failures. It does not guarantee correct values, so the next step matters more.

Extracting data from documents your systems cannot read?

ZSpace Labs builds extraction pipelines with schemas, validation and review screens, tuned on your own documents.

Start a Project

Step 5: Validate Every Field

  • Format checks: dates, IDs, currency codes, check digits
  • Arithmetic: line items sum to totals; tax matches rates
  • Cross-checks against system records: PO exists, supplier matches
  • Presence checks: required fields not null without a reason
  • Source checks: extracted value appears in the document text
  • Confidence routing: send uncertain fields to review
Structured output guarantees shape; validation checks the values.

Evaluating Extraction

Label a representative sample of documents, including poor scans and unusual layouts. Measure accuracy per field (not just per document), separately for header fields and line items, and track how often values are invented versus missed. Compare methods and models on this set before choosing, and rerun it when prompts, models or document sources change.

Costs

Costs come from OCR or document AI services per page, model tokens (images can be token-heavy), review time and infrastructure. Reduce them by routing simple, stable documents to cheaper methods, sending only relevant pages to models, using smaller models where accuracy holds and batching offline work.

Advantages and Limitations

AI extraction handles layout variety that defeats templates and dramatically reduces manual keying. It can invent plausible values, struggles with poor scans and long multi-page tables, and costs more per page than simple OCR rules. Validation, review and evaluation turn it from impressive to dependable.

Tables and Multi-Page Documents

Line items and tables cause most extraction errors. Detect tables with layout analysis or vision models, extract them as arrays with a defined row schema, handle tables that continue across pages by carrying headers forward, and reconcile row totals with document totals. For long documents, classify pages first and send only relevant pages to extraction, which improves accuracy and cuts cost.

Tools and Services

CategoryUse for
Cloud document AI servicesOCR, layout, prebuilt models for common documents
Open-source OCR and layout toolsSelf-hosted parsing and data control
Language and vision model APIs with structured outputsFlexible schema extraction across layouts
IDP platformsEnd-to-end pipelines with review UIs
Validation libraries and rules enginesField checks, cross-checks and business rules

Worked Example

An illustrative scenario, not a client case: an insurer extracts data from repair estimates in dozens of formats. A vision-language model with a strict schema returns line items and totals; validation checks arithmetic and that each part number appears in the OCR text; mismatches go to adjusters with the relevant region highlighted. Field-level accuracy is tracked weekly by repair shop to spot problem formats.

Common Mistakes

  • No nullable fields, so models guess
  • Measuring per-document rather than per-field accuracy
  • Trusting structured output without value validation
  • Ignoring multi-page tables
  • Sending whole documents when one page holds the data

Need reliable data from messy documents?

Talk to ZSpace Labs about AI document extraction and integration into your systems.

Start a Project

Conclusion

Good document extraction combines the right text source, a precise schema, structured outputs, rigorous validation and review for uncertain fields, measured on your own documents. Related: intelligent document processing and AI invoice processing.

FAQ

Common questions

Turning the content of documents such as PDFs, scans and images into structured fields and tables that software can use, using OCR, layout analysis and machine learning or language and vision models.

Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.