# OCR – PDF to Word (Arabic and all other languages) (`mshshakir/ocr-arabic`) Actor

Convert Arabic PDFs to Word with Google Cloud Vision OCR. Optimized for manuscripts and books, it highlights low-confidence words in red for easy review. Get a clean, editable .docx file ready for publishing. Works for other languages too—fast, accurate, and reliable.

- **URL**: https://apify.com/mshshakir/ocr-arabic.md
- **Developed by:** [Mufaddal Shakir](https://apify.com/mshshakir) (community)
- **Categories:** AI, Integrations, Automation
- **Stats:** 3 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $30.00 / 1,000 page processeds

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Arabic OCR – PDF to Word

Convert Arabic PDF files into fully editable Word documents using Google Cloud Vision OCR — with automatic highlighting of low-confidence words for easy human review.

***

### ⬇️ How to Download Your Output

After the Actor finishes running:

1. Open the completed run
2. Click the **Storage** tab (top of the run page)
3. Click **Key-value store**
4. Find the key named **OUTPUT**
5. Click the **Download** button next to it

Your `.docx` file will download immediately. Open it in Microsoft Word or Google Docs.

> **Note:** The "Dataset" tab only shows billing records (one row per page processed). Your actual Word document is always in the Key-value store under the OUTPUT key.

***

### What This Actor Does

1. **Downloads** your PDF from a public URL
2. **Runs** Google Cloud Vision OCR (industry-leading Arabic accuracy)
3. **Compiles** all pages into a single `.docx` Word file
4. **Highlights** uncertain words in red so you can review them quickly
5. **Outputs** the DOCX file to the key-value store for download

***

### Why Use This Instead of Generic OCR?

- ✅ **Arabic-first** — built specifically for Arabic, not an afterthought
- ✅ **Manuscript support** — works on classical Arabic script, not just modern print
- ✅ **Confidence highlighting** — low-confidence words appear in red in the output document
- ✅ **Custom page numbering** — set the real-world starting page number of your PDF

***

### Input Parameters

| Parameter | Required | Default | Description |
|-----------|----------|---------|-------------|
| `pdfUrl` | ✅ Yes | — | Direct public URL to your PDF |
| `fileLabel` | No | `document` | Short name for the output file (no spaces) |
| `startingPageNumber` | No | `1` | Real-world page number of the first page |
| `confidenceThreshold` | No | `0.9` | Words below this confidence are highlighted red (0.0–1.0) |
| `gcsBucketName` | No | `kutub-scanning` | GCS bucket (leave as default) |

***

### Supported PDF URLs

Your PDF must be publicly accessible. Supported sources:

- **Direct URL** — any public `https://` link ending in `.pdf`
- **Google Drive** — share the file as "Anyone with the link", paste the share URL
- **Dropbox** — paste the share link (the Actor converts it automatically)

***

### Pricing

| Pages | Estimated Cost |
|-------|----------------|
| 10 | ~$0.80 |
| 50 | ~$2.00 |
| 100 | ~$3.50 |
| 500 | ~$15.50 |

Pricing = $0.50 Actor start fee + $0.03 per page processed.

***

### Example Use Cases

- Digitizing Islamic manuscript collections
- Converting scanned Arabic books for academic research
- Archiving Arabic legal or historical documents
- Bulk digitization for libraries and institutions

***

### Notes

- Your PDF must be **publicly accessible** via URL
- Very large PDFs (500+ pages) may take several minutes — this is normal
- Password-protected PDFs are not supported
- For best results, use high-resolution scans (300 DPI or above)
- The red-highlighted words in the output indicate lower OCR confidence — review these manually

# Actor input Schema

## `pdfUrl` (type: `string`):

Direct URL to the PDF file you want to process. Must be publicly accessible.

## `fileLabel` (type: `string`):

A short name for this file (no spaces). Used to name the output DOCX.

## `startingPageNumber` (type: `integer`):

The real-world page number of the first page in your PDF (e.g. if your PDF starts at page 45 of a book, enter 45).

## `confidenceThreshold` (type: `number`):

Words with OCR confidence below this value will be highlighted in red in the output DOCX. Range: 0.0 to 1.0.

## `gcsBucketName` (type: `string`):

Your Google Cloud Storage bucket name. Leave as default if you are using the managed service.

## Actor input object example

```json
{
  "pdfUrl": "https://example.com/arabic-document.pdf",
  "fileLabel": "kitab-al-ilm",
  "startingPageNumber": 1,
  "confidenceThreshold": 0.85,
  "gcsBucketName": "kutub-scanning"
}
```

# Actor output Schema

## `docxFile` (type: `string`):

Download your Arabic OCR output as an editable .docx file. Low-confidence words are highlighted in red.

## `pages` (type: `string`):

Dataset records showing each page processed, used for billing.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("mshshakir/ocr-arabic").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("mshshakir/ocr-arabic").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call mshshakir/ocr-arabic --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=mshshakir/ocr-arabic",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fp72gPzWrRfD5p5qu/builds/mELQl0XayqVgLrGvj/openapi.json
