# Internet Archive & Wayback Machine Scraper (`mangudai/internet-archive-scraper`) Actor

Search the Internet Archive's 40M+ items, pull full item metadata and file lists, and query the Wayback Machine for URL snapshots. Books, audio, video, software, and archived pages on official archive.org APIs. No API key.

- **URL**: https://apify.com/mangudai/internet-archive-scraper.md
- **Developed by:** [Mangudäi](https://apify.com/mangudai) (community)
- **Categories:** Developer tools, SEO tools, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Internet Archive and Wayback Machine scraper

Search the Internet Archive's 40M+ items, pull the full record and file list for any item, and look up historical snapshots of any URL in the Wayback Machine. One actor, four jobs, all running on archive.org's official public APIs. No API key, no login, nothing that rate-limits into breakage.

### Four modes

Pick one in the **What to do** field:

- **Search items**: query the catalog of books, audio, video, software, images, web archives, and data. Filter by media type, sort by relevance, downloads, date, or title.
- **Item metadata + files**: give it archive.org identifiers and get the complete metadata plus every file, with format, size, and download base.
- **Wayback snapshots (all)**: list every capture of a URL from the Wayback Machine, with timestamp, status, MIME type, and the direct snapshot link. Filter by date range.
- **Wayback closest snapshot**: find the single capture nearest a date you name, or the most recent one.

### Search example

```json
{
  "mode": "search",
  "searchQueries": ["machine learning", "public domain films"],
  "mediaType": "texts",
  "sortBy": "downloads desc",
  "maxResults": 50
}
```

Each item comes back with identifier, title, creator, description, media type, collection, date, year, language, subject tags, download count, item size, file count, formats, and ready-made links to the details page, the download page, and the thumbnail.

### Wayback example

```json
{
  "mode": "waybackCaptures",
  "urls": ["example.com", "nasa.gov"],
  "waybackFrom": "2010",
  "waybackTo": "2015",
  "maxResults": 100
}
```

You get one row per snapshot: the capture timestamp, an ISO date, the original URL, the HTTP status, the content type, and a `waybackUrl` that opens the archived page.

### Who uses it

Researchers and librarians pulling public-domain sources, journalists and OSINT analysts checking what a page said on a given date, digital preservation teams cataloguing archived sites and old software, and anyone building a dataset from public archive metadata without clicking through archive.org by hand.

### How to find an item identifier

Open any item on archive.org. The identifier is the last part of the URL: for `archive.org/details/nasa_apollo11`, the identifier is `nasa_apollo11`. Put that in the **Item identifiers** field for metadata mode.

### Notes

Everything here reads public data through documented archive.org endpoints: the advanced search API, the item metadata API, the Wayback CDX API, and the availability API. It stores no credentials and pulls nothing private. Broad search terms return huge result sets, so narrow with a media type or a more specific query when you can.

# Actor input Schema

## `mode` (type: `string`):

Pick the job. Search finds items; item metadata pulls the full record and file list; the two Wayback modes look up historical snapshots of a URL.

## `searchQueries` (type: `array`):

One search per line. Used only in Search mode.

## `mediaType` (type: `string`):

Narrow a search to one media type.

## `sortBy` (type: `string`):

Order of search results.

## `maxResults` (type: `integer`):

Caps items per search query, or snapshots per URL in Wayback captures mode.

## `identifiers` (type: `array`):

Archive.org item identifiers, one per line (the last part of an archive.org/details/ URL). Used only in Item metadata mode.

## `urls` (type: `array`):

Web addresses to look up in the Wayback Machine, one per line. Used by both Wayback modes.

## `waybackFrom` (type: `string`):

Only snapshots on or after this date. Wayback captures mode.

## `waybackTo` (type: `string`):

Only snapshots on or before this date. Wayback captures mode.

## `waybackTimestamp` (type: `string`):

Find the snapshot nearest this date. Wayback closest snapshot mode. Leave empty for the most recent.

## Actor input object example

```json
{
  "mode": "search",
  "searchQueries": [
    "nasa apollo 11"
  ],
  "mediaType": "all",
  "sortBy": "relevance",
  "maxResults": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "nasa apollo 11"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mangudai/internet-archive-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["nasa apollo 11"] }

# Run the Actor and wait for it to finish
run = client.actor("mangudai/internet-archive-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "nasa apollo 11"
  ]
}' |
apify call mangudai/internet-archive-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=mangudai/internet-archive-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/VVgfpRuawATeYHUYp/builds/YMABmhiY1lWVzifx0/openapi.json
