# Tool: Document OCR to Structured JSON (`scrapers_lat/document-ocr-json-tool`) Actor

Read any document image and return the full OCR text plus clean structured JSON fields using premium AI. Paid Apify plans only.

- **URL**: https://apify.com/scrapers\_lat/document-ocr-json-tool.md
- **Developed by:** [Scrapers Lat](https://apify.com/scrapers_lat) (community)
- **Categories:** AI, Developer tools, Business
- **Stats:** 2 total users, 1 monthly users, 92.6% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $96.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![Tool: Document OCR to Structured JSON](https://scrapers.lat/banners/document-ocr-json-tool.png)](https://console.apify.com/actors/document-ocr-json-tool)

## Tool: Document OCR to Structured JSON

> Turn any document image into clean data. Read the full text of a scan or photo and pull out the exact fields you need as structured JSON.

**📥 [Input](https://apify.com/scrapers_lat/document-ocr-json-tool/input-schema) · 📤 [Output](https://apify.com/scrapers_lat/document-ocr-json-tool/output-schema) · 💰 [Pricing](https://apify.com/scrapers_lat/document-ocr-json-tool/pricing) · ▶️ [Examples](https://apify.com/scrapers_lat/document-ocr-json-tool/examples)**

![Apify](https://img.shields.io/badge/Platform-Apify-1CE1CE?logo=apify\&logoColor=white)
![Paid plan](https://img.shields.io/badge/Requires-Paid%20Apify%20plan-DE7356)
![Output](https://img.shields.io/badge/Output-JSON%20%7C%20CSV%20%7C%20Excel-orange)

<table><tr>
<td align="center"><strong>Full OCR text</strong><br>read in reading order</td>
<td align="center"><strong>Your fields</strong><br>key/value JSON</td>
<td align="center"><strong>JSON / CSV / Excel</strong><br>output formats</td>
</tr></table>

<br>

### What it does

Give it a list of image URLs, each a photo or scan of a document. For every image it reads the text and returns a clean, structured record:

- **extractedText**: the full text found in the document, kept in reading order
- **fields**: the exact key/value data you asked for, or the important fields auto-detected for you
- **documentType**: a short label such as invoice, receipt, letter, form, ID or contract
- **imageUrl** and **observedAt** for every record

Point it at invoices, receipts, letters, forms, IDs, contracts or plain pages of text. Feed the results into a database, a spreadsheet or a workflow. Export to JSON, CSV or Excel, or pull it through the Apify API.

### Paid plans only

This tool runs on the latest, most reliable premium AI models. Because each run has a real model cost, it is available to paid Apify accounts only. A run started from a free account stops immediately with an upgrade message. Upgrade your Apify subscription to a paid plan and run it again.

### Who is it for

| Use case | Who benefits |
|---|---|
| Back office | Teams digitizing paper documents at scale |
| Operations | Ops turning scans into database rows |
| Developers | Builders who want OCR plus structured fields in one call |
| Research | Analysts pulling numbers out of document images |

### Input

- **imageUrls**: a list of image URLs, each a document to read.
- **fields**: optional list of specific fields to extract. Leave empty to auto-detect.
- **instructions**: optional plain-language instructions when a fixed field list does not fit.
- **maxItems**: optional cap on how many images to read this run.

### Notes

- Only text that is actually visible in the image is reported. Clear, well-lit images read best.
- Read documents you have the right to process, in line with applicable laws.

### Example use cases

Ready-to-run example tasks, each preconfigured for a common scenario. Open one and press run, or use it as a template:

- [OCR an invoice into structured JSON fields](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-invoice-to-json): Read an invoice image and extract vendor, invoice number, line items and totals into clean structured JSON fields.
- [OCR a scanned text page to JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-scanned-text-page): Extract all readable text from a scanned document page and return it as structured JSON for search and indexing.
- [OCR a store receipt into JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-store-receipt): Turn a photographed store receipt into JSON with merchant, date, items, tax and total for expense tracking and audits.
- [OCR an ID card to structured data](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-id-card): Read an identity card image and extract name, ID number, date of birth and expiry into structured JSON for KYC flows.
- [OCR a business card into contact JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-business-card): Scan a business card image and pull name, company, title, phone and email into clean JSON for your CRM imports.
- [OCR a tax form into JSON fields](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-tax-form): Extract labeled fields from a tax form image and return them as structured JSON for bookkeeping and data entry.
- [OCR a bank statement into JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-bank-statement): Read a bank statement image and extract account holder, period, transactions and balances into structured JSON.
- [OCR a purchase order to JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-purchase-order): Convert a purchase order image into JSON with PO number, supplier, items and totals for procurement workflows.
- [OCR a handwritten note into text JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-handwritten-note): Read a handwritten note image and return the transcribed text as clean JSON for digitizing forms, records and field notes.
- [OCR a shipping label into JSON](https://apify.com/scrapers_lat/document-ocr-json-tool/examples/ocr-ocr-shipping-label): Extract sender, recipient, tracking number and address from a shipping label image into structured JSON records.

### Export, API and AI agents (x402 + MCP)

Export the scraped data to **JSON, CSV or Excel**, pull it as a **dataset** through the Apify **API**, or wire it into your app with **no code**. This web scraper and data extractor also works for bulk data extraction and scheduled runs.

For AI agents: this Actor is available on **x402**, Apify's agentic payment standard built with Coinbase. An AI agent can discover, pay for and run it on its own with a funded wallet and a single HTTP request: no account, no subscription, no API key and no human in the loop. It also runs as an **MCP** tool inside Claude, Cursor and other AI clients out of the box. Learn more about [x402 agentic payments on Apify](https://docs.apify.com/platform/integrations/x402).

Built by [scrapers.lat](https://scrapers.lat).

# Actor input Schema

## `maxDocuments` (type: `integer`):

Maximum number of documents (images) to read this run. Optional.

## `imageUrls` (type: `array`):

A list of image URLs to read. Each should be a photo or scan of a document (invoice, receipt, letter, form, ID, contract, page of text).

## `fields` (type: `array`):

Optional. Specific fields to pull out of each document, for example name, date, total, address. Leave empty to auto-detect the important fields.

## `instructions` (type: `string`):

Optional. Plain-language instructions for what to extract when a fixed field list does not fit.

## Actor input object example

```json
{
  "imageUrls": [
    "https://tesseract.projectnaptha.com/img/eng_bw.png",
    "https://templates.invoicehome.com/invoice-template-us-neat-750px.png"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "imageUrls": [
        "https://tesseract.projectnaptha.com/img/eng_bw.png",
        "https://templates.invoicehome.com/invoice-template-us-neat-750px.png"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapers_lat/document-ocr-json-tool").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "imageUrls": [
        "https://tesseract.projectnaptha.com/img/eng_bw.png",
        "https://templates.invoicehome.com/invoice-template-us-neat-750px.png",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("scrapers_lat/document-ocr-json-tool").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "imageUrls": [
    "https://tesseract.projectnaptha.com/img/eng_bw.png",
    "https://templates.invoicehome.com/invoice-template-us-neat-750px.png"
  ]
}' |
apify call scrapers_lat/document-ocr-json-tool --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapers_lat/document-ocr-json-tool",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/mpLKZt7Ky5XrYOo64/builds/qUpgHW4j5eDMBsRRg/openapi.json
