# Website Content Crawler — Text, Titles & Metadata (`hipersoft/website-content-crawler`) Actor

Extract clean, readable content from any list of websites: page title, meta description, headings, main body text, word count and link/image counts. Optional same-domain crawl. Bulk-ready, no browser, no login. Great for LLM/RAG ingestion, content audits and research.

- **URL**: https://apify.com/hipersoft/website-content-crawler.md
- **Developed by:** [hiper soft](https://apify.com/hipersoft) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 4 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.0024 / page scraped

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Content Crawler — Text, Titles & Metadata

Turn any list of websites into clean, structured content. For each page the Actor
extracts the **title, meta description, headings, main body text (boilerplate
stripped), word count, language and link/image counts** — over plain HTTP, no
browser and no login. Optionally **crawl internal links** to pull a whole section.

Ideal for **LLM/RAG ingestion, content audits, competitive research, and building
text datasets** from many sites at once.

### What you get per page

| Field | Notes |
|---|---|
| `title`, `description`, `lang` | Page title, meta/OG description, `<html lang>`. |
| `text` | Readable body text with scripts/nav/header/footer stripped; prefers `<main>`/`<article>`. |
| `wordCount` | Word count over the full extracted text. |
| `headings` | `h1`/`h2` outline (level + text). |
| `internalLinkCount`, `externalLinkCount`, `imageCount` | Structure signals. |
| `links` | Arrays of internal/external links (toggle off to slim the output). |
| `canonicalUrl`, `domain`, `depth` | Canonical link, host, and crawl depth. |

### Input

```json
{
  "urls": ["https://apify.com/blog", "example.com"],
  "crawl": true,
  "maxPagesPerDomain": 20,
  "maxDepth": 2
}
```

- **urls** — pages/sites to extract, one per line. (Or use **startUrls**.)
- **crawl** — follow same-domain links (BFS). Off = just the given URLs.
- **maxPagesPerDomain** / **maxDepth** — crawl limits.
- **sameDomainOnly** — keep the crawl on the start domain (default on).
- **includeLinks** — include the internal/external link arrays (counts always included).
- **textMaxChars** — truncate stored body text (word count uses the full text).
- **maxConcurrency** / **maxItems** — parallelism and a global page cap.
- **proxyConfiguration** — optional; enable for reliable access at scale.

### Output (one row per page)

```json
{
  "input": "https://apify.com/blog",
  "url": "https://apify.com/blog",
  "domain": "apify.com",
  "depth": 0,
  "title": "Apify Blog",
  "description": "News and tutorials…",
  "lang": "en",
  "headings": [{ "level": 1, "text": "Apify Blog" }],
  "text": "…clean readable content…",
  "wordCount": 812,
  "internalLinkCount": 47,
  "externalLinkCount": 6,
  "imageCount": 12,
  "error": null
}
```

### FAQ

**Do I need an API key or login?**
No. The Actor works over plain HTTP with no browser and no login — just paste your list of URLs or domains and run.

**How many pages can I get per run?**
Control it with `maxPagesPerDomain`, `maxDepth` and a global `maxItems` cap. With crawl on, it follows same-domain links breadth-first, so one run can extract from a whole site section or many sites at once.

**Is crawling website content legal?**
The Actor collects only publicly available page content — the same pages any visitor can open. Public data is generally collectible, but you are responsible for respecting each site's terms and applicable law.

**Does it collect emails or contact info?**
No. It extracts page content — title, description, headings, body text, word count and link/image counts — not contact details. For emails, phones and socials, use the [Website Contact Scraper](https://apify.com/hipersoft/website-contact-scraper).

**What's the output format?**
One structured JSON row per page with clean, boilerplate-stripped text and metadata — ready for LLM/RAG ingestion, content audits or text datasets. Export as JSON, CSV or Excel.

### Related Actors

Pair content extraction with these web-tool companions:

- [Website Contact Scraper](https://apify.com/hipersoft/website-contact-scraper) — extract emails, phones and social profiles from the same sites.
- [Google Maps Scraper](https://apify.com/hipersoft/google-maps-scraper) — get local business listings by keyword and location.
- [Google Maps Email Extractor](https://apify.com/hipersoft/google-maps-email-extractor) — turn Maps results into contactable leads with emails and socials.

### Notes

Every site is hard-capped in time so one slow page can never stall a bulk run.
Only publicly available page content is collected; respect each site's terms.

# Actor input Schema

## `urls` (type: `array`):

Websites/pages to extract content from (e.g. "apify.com", "https://example.com/blog"). One per line.

## `startUrls` (type: `array`):

Alternative to "URLs": a request list of {"url": "..."} objects.

## `crawl` (type: `boolean`):

When on, follow same-domain links from each start URL (breadth-first) up to the page/depth limits below. When off, only the given URLs are fetched (one row each).

## `maxPagesPerDomain` (type: `integer`):

When crawling, the maximum pages to fetch per start URL.

## `maxDepth` (type: `integer`):

When crawling, how many link-hops deep to go from the start URL.

## `sameDomainOnly` (type: `boolean`):

Only follow links on the start URL's domain while crawling.

## `includeLinks` (type: `boolean`):

Include the arrays of internal/external links in each record (counts are always included).

## `textMaxChars` (type: `integer`):

Truncate extracted body text to this many characters (word count is computed on the full text).

## `maxConcurrency` (type: `integer`):

How many sites to process in parallel.

## `maxItems` (type: `integer`):

Cap total pages across all sites (0 = no cap).

## `proxyConfiguration` (type: `object`):

Optional. Enable a proxy if a target blocks datacenter IPs.

## Actor input object example

```json
{
  "urls": [
    "https://apify.com"
  ],
  "crawl": false,
  "maxPagesPerDomain": 1,
  "maxDepth": 1,
  "sameDomainOnly": true,
  "includeLinks": true,
  "textMaxChars": 40000,
  "maxConcurrency": 10,
  "maxItems": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("hipersoft/website-content-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("hipersoft/website-content-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com"
  ]
}' |
apify call hipersoft/website-content-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=hipersoft/website-content-crawler",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/mDJgCoWPdtY3wIZrx/builds/Rmf4MoKP7TW9oJxFA/openapi.json
