# Museum Artworks — Met + Cleveland Unified Search (`scrupulous_waterbird_m4w/museum-artworks`) Actor

Search artwork records from the Met and Cleveland Museum of Art in one unified JSON schema. Returns artist, title, date, medium, dimensions, images, and license per record. CC0 / public-domain by default. Open APIs, no auth, no captcha.

- **URL**: https://apify.com/scrupulous\_waterbird\_m4w/museum-artworks.md
- **Developed by:** [Mori](https://apify.com/scrupulous_waterbird_m4w) (community)
- **Categories:** Other, Developer tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Museum Artworks — Met + Cleveland Unified Search (Apify Actor)

Search artwork records from the **Metropolitan Museum of Art** and the **Cleveland Museum of Art** in a unified JSON schema. Returns one record per artwork with artist, title, date range, medium, dimensions, images, credit line, source URL, license posture, and free-form tags.

### Source coverage

| Source | Auth | Captcha | Endpoint | Records | License / ToS |
|---|---|---|---|---|---|
| Metropolitan Museum of Art | none | none | `https://collectionapi.metmuseum.org/public/collection/v1/` | 417,474+ | Open Access under CC0 (`metmuseum.github.io`); robot-friendly public API; no documented rate limit, see "Polite use" below |
| Cleveland Museum of Art | none | none | `https://openaccess-api.clevelandart.org/api/artworks` | 68,758+ | Open Access API; images distributed under `share_license_status` (`CC0` / `CC-BY-SA` etc. exposed per record) |

Both endpoints return JSON, no proxy / captcha / JS rendering needed, no auth, no API key. They are public research APIs maintained by each museum; we honor each museum's published ToS (request rate ~1 req/sec, descriptive `User-Agent`).

### What it does

- Two source adapters (`met`, `cleveland`), one orchestrator. The `museum="both"` mode merges results (Met IDs first, then Cleveland). Each adapter is independent — one source failing doesn't block the other (gotcha #37 — independent clients per source).
- Output is one dataset record per artwork, harmonized to a single schema across both sources:
  - `source`: `"met"` | `"cleveland"`
  - `sourceId`: museum-side identifier (`objectID` / `accession_number`)
  - artist + bio, title, date display string, `dateEarliest`/`dateLatest` (integer years, nullable)
  - medium, dimensions, department, classification, culture list
  - `isPublicDomain` Boolean (image license flag)
  - `primaryImage` / `primaryImageSmall` URLs, `additionalImages[]`
  - `creditLine`, `sourceUrl`, `tags[]`
- Outputs are pushed to the default Apify dataset via `Actor.pushData(...)` per gotcha #13.
- Smoke-tested inputs `museum=met` and `museum=cleveland` both return non-null records on a real query.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `museum` | select | `"both"` | `met` | `cleveland` | `both` |
| `query` | textfield | `"van gogh"` | Free-text search. Met: maps to `?q=`. Cleveland: maps to `?q=`. |
| `hasImages` | checkbox | `true` | Met: limit to records with `primaryImage`. Cleveland: always image-present when matched. |
| `maxResults` | number (1-500) | 50 | Hard cap on total records returned across all sources. |

### Output

One dataset record per artwork. See `.actor/dataset_schema.json` for the full shape. Sample (first record from a `museum=met, query=rembrandt, maxResults=3` smoke run, abridged):

```json
{
  "source": "met",
  "sourceId": "437394",
  "title": "Aristotle with a Bust of Homer",
  "artist": "Rembrandt (Rembrandt van Rijn)",
  "artistBio": "Dutch, Leiden 1606–1669 Amsterdam",
  "date": "1653",
  "dateEarliest": 1653,
  "dateLatest": 1653,
  "medium": "Oil on canvas",
  "dimensions": "56 1/2 x 53 3/4 in. (143.5 x 136.5 cm)",
  "department": "European Paintings",
  "classification": "Paintings",
  "culture": [],
  "isPublicDomain": true,
  "primaryImage": "https://images.metmuseum.org/CRDImages/ep/original/DP-30758-001.jpg",
  "primaryImageSmall": "https://images.metmuseum.org/CRDImages/ep/web-large/DP-30758-001.jpg",
  "additionalImages": [],
  "creditLine": "Purchase, special contributions and funds given or bequeathed by friends of the Museum, 1961",
  "sourceUrl": "https://www.metmuseum.org/art/collection/search/437394",
  "tags": ["Portraits", "Paintings"]
}
```

### Limits / gotchas

- `maxResults` defaults to 50, hard cap 500 (card spec).
- Polite-use: 1 request/second per source. Concurrency is sequential per source; sources run via `Promise.allSettled`.
- Cleveland's `share_license_status` is mapped to the unified `isPublicDomain` Boolean (true when status starts with "CC0", false otherwise). Cleveland images without a permissive license are still returned with their CC URL but downstream consumers should consult `share_license_status` if their use case requires it.

### How to run locally

```bash
cd actors/museum-artworks
make install      # npm install apify SDK
make run          # local apify run -p -i .actor/input.json (full SKIP_LOCAL_RUN=1 path is recommended for CI)
```

### Build + push to Apify

```bash
SKIP_LOCAL_RUN=1 bash /Users/hermes-agent/Documents/agent-vault/Companies/apify-actor/bin/publish.sh museum-artworks
```

`bin/publish.sh` steps: (1) pre-flight auth + syntax + docker, (2) skipped via SKIP\_LOCAL\_RUN=1, (3) `apify push`, (4) cloud smoke (`apify call`), (5) inventory update.

### License

Actor code: MIT (inherited from project template).

Data attribution:

- Metropolitan Museum of Art images: CC0 (Met Open Access).
- Cleveland Museum of Art images: per `share_license_status` per record (commonly CC0 / CC-BY-SA).

### Changelog

- 0.1.0 — initial release. Met + Cleveland.

# Actor input Schema

## `museum` (type: `string`):

Which museum's collection to query. 'both' merges Met + Cleveland records.

## `query` (type: `string`):

Free-text search. Met: maps to the Met 'q' field (matches title, artistDisplayName, tags, etc.). Cleveland: maps to the 'q' URL parameter (matches title, artist, tombstone, etc.).

## `hasImages` (type: `boolean`):

Met only: skip records with no primaryImage. Cleveland always returns images for matched records.

## `maxResults` (type: `integer`):

Hard cap on total records returned across all sources (1-500). Default 50.

## Actor input object example

```json
{
  "museum": "both",
  "query": "van gogh",
  "hasImages": true,
  "maxResults": 50
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "van gogh"
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrupulous_waterbird_m4w/museum-artworks").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "van gogh" }

# Run the Actor and wait for it to finish
run = client.actor("scrupulous_waterbird_m4w/museum-artworks").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "van gogh"
}' |
apify call scrupulous_waterbird_m4w/museum-artworks --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrupulous_waterbird_m4w/museum-artworks",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oNhU7JNrJzboaAFUv/builds/xOkXAH7hMC9uQCBba/openapi.json
