# Wikimedia Commons Scraper — Free Images & Media (No Key) (`ninhothedev/wikimedia-commons-scraper`) Actor

$0.5/1K 🔥 Fast Wikimedia Commons scraper! Free images & media by search — url, author, license & description. No key. JSON, CSV, Excel or API in seconds. Pull thousands of free-to-use media for datasets & design ⚡

- **URL**: https://apify.com/ninhothedev/wikimedia-commons-scraper.md
- **Developed by:** [ninhothedev](https://apify.com/ninhothedev) (community)
- **Categories:** Developer tools, Videos
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Wikimedia Commons Scraper — Free Images & Media (No API Key)

Scrape **free, reusable images and media from [Wikimedia Commons](https://commons.wikimedia.org/)** at scale. Give the Actor a list of search terms and it returns clean, structured records for every matching file — direct file URL, author/artist, license, usage terms, description, categories, dimensions and more. No API key, no login, datacenter-proxy friendly.

Wikimedia Commons hosts **100M+ freely licensed media files** (public domain and Creative Commons). This Actor turns that library into a ready-to-use dataset for image sourcing, machine-learning pipelines, content and design workflows, and research.

### Features

- **No API key required** — uses the public MediaWiki API.
- **Datacenter-proxy friendly** — clean JSON endpoint, no browser needed.
- **Rich metadata** — direct URL, artist, license, usage terms, description, credit, categories, dimensions, MIME type, uploader.
- **HTML cleaned** — author/description/credit fields are stripped of HTML for direct use.
- **Multi-query** — pass many search terms in one run.
- **Cheap & fast** — lightweight HTTP requests, pay only for what you collect.

### Input

| Field | Type | Description |
|-------|------|-------------|
| `mode` | select | Scraping mode. `search` runs your text queries against the Commons file namespace. |
| `queries` | array | One or more search terms (e.g. `["sunset", "golden gate bridge"]`). |
| `maxItems` | integer | Max total media files to collect (default 100, max 1000). |

#### Example input

```json
{
  "mode": "search",
  "queries": ["sunset", "golden gate bridge"],
  "maxItems": 100
}
```

### Output

Each item is one media file:

```json
{
  "page_id": 312884,
  "title": "File:Sunset 2007-1.jpg",
  "url": "https://upload.wikimedia.org/wikipedia/commons/5/58/Sunset_2007-1.jpg",
  "description_url": "https://commons.wikimedia.org/wiki/File:Sunset_2007-1.jpg",
  "width": 2789,
  "height": 1980,
  "size": 1327087,
  "mime": "image/jpeg",
  "uploader": "Alvesgaspar",
  "artist": "Alvesgaspar",
  "license": "CC BY-SA 3.0",
  "usage_terms": "Creative Commons Attribution-Share Alike 3.0",
  "description": "Sunset at Porto Covo, Portugal.",
  "credit": "Own work",
  "date": "2007",
  "categories": ["Sunsets", "Porto Covo"],
  "source": "wikimedia_commons",
  "scraped_at": "2026-07-20T00:00:00+00:00"
}
```

### Pricing

Priced to be cheap: roughly **~$0.5 per 1,000 media files**. You only pay for the items you collect. Costs scale linearly with `maxItems`.

### Use cases

- **Free image sourcing** — find openly licensed images for blogs, apps and products.
- **ML / AI datasets** — build labelled image datasets with license provenance baked in.
- **Content & design** — pull illustrations, photos and diagrams with proper attribution.
- **Research** — study licensing, categories and media metadata across topics.

### Licensing note

Files on Wikimedia Commons are freely licensed, but many require **attribution** or share-alike terms. Always respect each file's `license` / `usage_terms` and credit the `artist` when required.

### Related Actors

- [Openverse Media Scraper](https://apify.com/ninhothedev/openverse-media-scraper)
- [Met Museum Scraper](https://apify.com/ninhothedev/met-museum-scraper)
- [Art Institute of Chicago Scraper](https://apify.com/ninhothedev/art-institute-chicago-scraper)
- [Wikipedia Scraper](https://apify.com/ninhothedev/wikipedia-scraper)

### Keywords

wikimedia commons scraper, free images api, creative commons images, public domain images, image dataset, media scraper, stock photos free, openverse alternative, ml training images, no api key image scraper

***

Questions or a custom media source? Reach out on the Actor's Apify page.

# Actor input Schema

## `mode` (type: `string`):

Scraping mode. Currently 'search' runs your text queries against the Wikimedia Commons file namespace and returns matching media files.

## `queries` (type: `array`):

One or more search terms. Each query is run against Wikimedia Commons and returns matching free images/media files.

## `maxItems` (type: `integer`):

Maximum total number of media files to collect across all queries.

## Actor input object example

```json
{
  "mode": "search",
  "queries": [
    "sunset",
    "golden gate bridge"
  ],
  "maxItems": 100
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "sunset",
        "golden gate bridge"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ninhothedev/wikimedia-commons-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "sunset",
        "golden gate bridge",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("ninhothedev/wikimedia-commons-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "sunset",
    "golden gate bridge"
  ]
}' |
apify call ninhothedev/wikimedia-commons-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ninhothedev/wikimedia-commons-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/53ciSGCVObg7PhELM/builds/aDWdSCc8acZxkYCta/openapi.json
