# CourtListener RAG Extractor (`getascraper/courtlistener-rag-extractor`) Actor

Extract SCOTUS and U.S. federal appeals opinions from CourtListener into normalized RAG-ready JSON with fixed-token chunks, metadata, citations, and summary fallback. Built for legal AI and litigation research pipelines. $0.03 per opinion.

- **URL**: https://apify.com/getascraper/courtlistener-rag-extractor.md
- **Developed by:** [GetAScraper](https://apify.com/getascraper) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $0.67 / 1,000 dossier records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## ⚖️ CourtListener RAG extractor

### 🔍 What does CourtListener RAG Extractor do?

CourtListener RAG Extractor pulls U.S. federal court opinions from CourtListener and normalizes them into RAG-ready JSON records with fixed-token chunks, citations, and metadata. It focuses on SCOTUS and all 13 federal Courts of Appeals in v1.

Use it when you need legal AI corpora that are directly usable in LangChain, LlamaIndex, OpenAI vector stores, Qdrant, Pinecone, Weaviate, or pgvector pipelines.

### 🎯 Why use it?

- Build legal-AI retrieval corpora without writing custom ETL around CourtListener REST APIs.
- Get standardized record shape across courts: opinion ID, cluster ID, case naming, docket, filing date, citations, and source URL.
- Keep chunk size consistent for embedding and reranking workflows (512 tokens with 50-token overlap).
- Support litigation analytics and case-law search features with structured citation metadata.
- Run as an Apify Actor with scheduling, API access, and integration-ready dataset outputs.

### 🚀 How to use it?

1. Open the Actor in Apify Console.
2. Set your date range (`dateFrom`, `dateTo`) and optional court filters.
3. Optionally add a CourtListener API key from https://www.courtlistener.com/api/ for faster and richer detail retrieval.
4. Set `maxOpinions` to cap cost and runtime.
5. Run the Actor and consume the dataset output via API or download.

### 📋 Input fields

- `courtIds` (array): Court slugs. Empty means all 14 supported federal courts.
- `dateFrom` (string, required): Inclusive lower bound in `YYYY-MM-DD` format.
- `dateTo` (string, required): Inclusive upper bound in `YYYY-MM-DD` format.
- `searchQuery` (string): Optional CourtListener query syntax.
- `maxOpinions` (integer): Hard cap on returned opinions.
- `courtListenerApiKey` (secret string): Optional token for higher throughput and detail endpoint access.

### 📤 Output schema

Each dataset item is one opinion record:

```json
{
    "opinion_id": "11314034",
    "cluster_id": "10846667",
    "court": "scotus",
    "court_full": "Supreme Court of the United States",
    "case_name": "Enbridge Energy, LP v. Nessel",
    "case_name_short": null,
    "docket_number": "24-783",
    "date_filed": "2026-04-22",
    "citation_count": 21,
    "citations": ["608 U.S. ___"],
    "absolute_url": "https://www.courtlistener.com/opinion/11314034/enbridge-energy-lp-v-nessel/",
    "source": "summary",
    "chunks": [
        { "idx": 0, "text": "...", "tokens": 512 },
        { "idx": 1, "text": "...", "tokens": 213 }
    ]
}
```

### 📊 Data table

| Field            | Type   | Description                                         |
| ---------------- | ------ | --------------------------------------------------- |
| `opinion_id`     | string | CourtListener opinion ID                            |
| `court`          | string | Court slug (`scotus`, `ca1`...`cafc`)               |
| `case_name`      | string | Canonical case name                                 |
| `docket_number`  | string | Docket number                                       |
| `date_filed`     | string | Filing date                                         |
| `citation_count` | number | Citation count from cluster metadata when available |
| `source`         | string | `full_text` or `summary`                            |
| `absolute_url`   | string | Absolute CourtListener opinion URL                  |

### 💰 Pricing / cost estimation

Price target is **$0.015 per opinion**, with a **free trial of 10 results**.

| Opinions | Estimated cost |
| -------- | -------------- |
| 100      | $1.50          |
| 1,000    | $15            |
| 10,000   | $150           |

### ⭐ Enjoying CourtListener RAG Extractor?

<table width="100%">
<tr>
<td style="padding:20px 24px 14px;background:#EDEEF7;border:1px solid #EDEEF7;border-left:5px solid #2A3499;border-radius:10px 10px 0 0">
<span style="font-size:20px;letter-spacing:4px">⭐ ⭐ ⭐ ⭐ ⭐</span><br>
<span style="font-size:17px;font-weight:800;color:#1C1917">One run replaces days of manually copying opinions and building embedding-ready chunks by hand.</span><br>
<span style="font-size:14px;color:#57534E">A 5-star rating takes 10 seconds and helps other legal AI and RAG engineers find this actor. Your feedback also tells us what to build next.</span>
</td>
</tr>
<tr>
<td style="padding:0;background:#2A3499;border:1px solid #EDEEF7;border-top:none;border-radius:0 0 10px 10px;text-align:center">
<a href="https://apify.com/getascraper/courtlistener-rag-extractor/reviews" style="display:block;padding:13px 16px;color:#FFFFFF;text-decoration:none;font-weight:800;font-size:15px;letter-spacing:0.3px">★&nbsp;&nbsp;Rate this Actor on Apify</a>
</td>
</tr>
</table>

### 💡 Tips / Advanced

- Add `courtListenerApiKey` for stable throughput and richer detail endpoint access.
- Narrow with `searchQuery` for topic-specific corpora.
- Keep date windows smaller for incremental backfills.
- Start with low `maxOpinions` for schema checks, then scale.

### 🚧 Limits

- v1 supports only SCOTUS + 13 federal Courts of Appeals.
- No district or state courts in v1.
- No majority/concurrence/dissent section separation in v1.
- No citation graph extraction in v1.

### ⚠️ Legal disclaimer

This Actor extracts publicly available U.S. federal court opinions from CourtListener (operated by the Free Law Project). Output is not legal advice. Users are responsible for compliance with local professional-responsibility rules when using this data.

### ❓ FAQ

**Do I need a CourtListener API key?**
No, but it is strongly recommended. Without a key, the Actor uses conservative rate limits and may rely more heavily on summary-level fields.

**What happens when opinion detail endpoints are unavailable?**
The run continues with available search metadata and summary fallback so records still remain schema-consistent.

**Does this include citation graph relationships?**
No. v1 includes citation strings and counts, not graph topology.

### 🛟 Support

If you need feature requests or issue triage, open a ticket in this repo's Issues tab.

### 🔗 Other actors

- [arXiv scraper for RAG: papers as chunked JSON](https://apify.com/getascraper/arxiv-rag-extractor) ↗ - pulls arXiv papers into fixed-token chunks for embedding pipelines.
- [bioRxiv and medRxiv scraper for RAG: chunked JSON](https://apify.com/getascraper/biorxiv-medrxiv-rag-extractor) ↗ - extracts preprint biology and medicine papers as RAG-ready JSON.
- [PubMed Scraper for RAG: Papers as Chunked JSON](https://apify.com/getascraper/pubmed-rag-extractor) ↗ - normalizes PubMed biomedical literature into chunked records for retrieval.
- [SEC EDGAR Scraper for RAG: 10-K/10-Q/8-K as JSON](https://apify.com/getascraper/sec-edgar-rag-extractor) ↗ - converts SEC filings into structured, chunked JSON for financial AI corpora.

# Actor input Schema

## `courtIds` (type: `array`):

Select federal courts (such as 'scotus' for Supreme Court, or 'ca1' for Court of Appeals). Leave empty for all.

## `searchQuery` (type: `string`):

Optional search term to filter opinions (such as 'copyright' or 'trademark'). Leave blank for all.

## `dateFrom` (type: `string`):

Collect opinions filed on or after this date.

## `dateTo` (type: `string`):

Collect opinions filed on or before this date.

## `maxOpinions` (type: `integer`):

Set the maximum number of opinions you want to save for this run.

## `courtListenerApiKey` (type: `string`):

Paste your free CourtListener API key to speed up collection. Get one on your CourtListener account page.

## Actor input object example

```json
{
  "courtIds": [
    "scotus"
  ],
  "searchQuery": "",
  "dateFrom": "2024-01-01",
  "dateTo": "2024-06-30",
  "maxOpinions": 50
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "courtIds": [
        "scotus"
    ],
    "searchQuery": "",
    "dateFrom": "2024-01-01",
    "dateTo": "2024-06-30",
    "maxOpinions": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("getascraper/courtlistener-rag-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "courtIds": ["scotus"],
    "searchQuery": "",
    "dateFrom": "2024-01-01",
    "dateTo": "2024-06-30",
    "maxOpinions": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("getascraper/courtlistener-rag-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "courtIds": [
    "scotus"
  ],
  "searchQuery": "",
  "dateFrom": "2024-01-01",
  "dateTo": "2024-06-30",
  "maxOpinions": 50
}' |
apify call getascraper/courtlistener-rag-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=getascraper/courtlistener-rag-extractor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/JaSWZRPzubu3qQoBl/builds/3cowwrpjdfcbkZZx6/openapi.json
