# Mojeek Image Scraper (`codingfrontend/mojeek-image-scraper`) Actor

Scrapes image search results from Mojeek (mojeek.com). Extracts image URL, source page, title, license, creator, and dimensions. Supports infinite scroll pagination.

- **URL**: https://apify.com/codingfrontend/mojeek-image-scraper.md
- **Developed by:** [Coding Frontned](https://apify.com/codingfrontend) (community)
- **Categories:** Developer tools, Automation, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.99 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Mojeek Image Scraper

Extract image search results from [Mojeek](https://www.mojeek.com) — a privacy-first, independent search engine powered by its own crawler. Mojeek images are sourced via [Openverse](https://openverse.org) (Creative Commons-licensed media from Flickr, Wikipedia, museums, and more).

### Features

- **Image results** — thumbnail URL, full-size image URL, source page, title, creator, license
- **Infinite scroll** — automatically scrolls to load more images up to `maxItems`
- **Pagination** — follows explicit next-page links for large result sets
- **License details** — extracts CC license type, license URL, and attribution info
- **Creator attribution** — captures photographer/creator name and profile URL
- **Provider info** — tracks which image provider (e.g. Openverse) supplied the image
- **Deduplication** — cross-page deduplication prevents duplicate images
- **Proxy support** — works with Apify residential/datacenter proxies

### Input Parameters

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `query` | string | *(required)* | Image search query |
| `maxItems` | integer | `100` | Maximum number of images to extract |
| `license` | string | `""` | License filter: `""` = all, `"cc"` = Creative Commons, `"commercial"` = commercial use |
| `proxyConfiguration` | object | — | Apify proxy config (recommended: residential) |

#### Example INPUT.json

```json
{
    "query": "nature photography",
    "maxItems": 100,
    "license": "cc",
    "proxyConfiguration": { "useApifyProxy": true }
}
```

### Output Fields

| Field | Type | Description |
|-------|------|-------------|
| `position` | integer | Rank in results (1-based) |
| `title` | string | Image title (when available) |
| `imageUrl` | string | Thumbnail image URL (via Mojeek proxy) |
| `largeImageUrl` | string | Full-size image URL (via Mojeek proxy) |
| `sourcePage` | string | Source page hosting the original image |
| `sourceDomain` | string | Domain of the source page |
| `license` | string | License type (e.g. `BY-SA`, `BY-NC-ND`) |
| `licenseUrl` | string | URL to the Creative Commons license |
| `reportUrl` | string | URL to report the image on Openverse |
| `creator` | string | Image creator/photographer name |
| `creatorProfileUrl` | string | Link to creator's profile |
| `provider` | string | Image provider (e.g. `openverse`) |
| `searchQuery` | string | Query that produced this result |
| `scrapedAt` | string | ISO 8601 scrape timestamp |

#### Example Output

```json
{
    "position": 1,
    "title": "Cats at sunset",
    "imageUrl": "https://www.mojeek.com/image?img=https://api.openverse.org/v1/images/abc123",
    "largeImageUrl": "https://www.mojeek.com/image?img=https://api.openverse.org/v1/images/abc123-large",
    "sourcePage": "https://www.flickr.com/photos/user123/photo456",
    "sourceDomain": "flickr.com",
    "license": "BY-SA",
    "licenseUrl": "https://creativecommons.org/licenses/by-sa/2.0/",
    "reportUrl": "https://openverse.org/image/abc123",
    "creator": "Jane Doe",
    "creatorProfileUrl": "https://www.flickr.com/photos/user123",
    "provider": "openverse",
    "searchQuery": "cats photography",
    "scrapedAt": "2025-05-01T12:00:00.000Z"
}
```

### How Pagination Works

Mojeek Images uses **infinite scroll** to load more results:

- Initial page loads ~24 images
- Each scroll loads ~24 more
- The scraper scrolls automatically until `maxItems` is reached
- Also follows explicit pagination links (`s=25`, `s=49`, etc.) for larger requests

### Dataset Views

The dataset provides two views:

1. **Image Results Overview** — key fields including thumbnail, source, license, creator
2. **Images By License** — focused view with full attribution: license URL, creator profile, provider

### Image Sources

Mojeek's image search is powered by [Openverse](https://openverse.org), which aggregates openly licensed media from:

- **Flickr** — user-contributed photography
- **Wikimedia Commons** — free media files
- **Museums & cultural institutions** — digitized collections
- **Open media databases** — openly licensed stock images

### Notes

- All images are Creative Commons licensed or openly licensed
- `license` values follow CC naming: `BY`, `BY-SA`, `BY-NC`, `BY-NC-ND`, `BY-NC-SA`, `BY-ND`
- Use a residential proxy for better reliability at scale
- Images are proxied via Mojeek's `/image?img=` endpoint for display

# Actor input Schema

## `query` (type: `string`):

The image search query.

## `maxItems` (type: `integer`):

Maximum number of images to scrape.

## `license` (type: `string`):

Filter images by license type. Leave empty for all images.

## `proxyConfiguration` (type: `object`):

Proxy settings for the scraper. Residential proxies recommended.

## Actor input object example

```json
{
  "query": "nature photography",
  "maxItems": 100,
  "license": "",
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "nature photography",
    "maxItems": 100,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("codingfrontend/mojeek-image-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "nature photography",
    "maxItems": 100,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("codingfrontend/mojeek-image-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "nature photography",
  "maxItems": 100,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call codingfrontend/mojeek-image-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=codingfrontend/mojeek-image-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6a0kGEyNugvKASWpU/builds/GMNxtiPuNF2cnmkFt/openapi.json
