Reddit Comments Scraper avatar

Reddit Comments Scraper

Pricing

from $0.35 / 1,000 comment or post rows

Go to Apify Store
Reddit Comments Scraper

Reddit Comments Scraper

Export every comment from any Reddit thread as analysis-ready rows: text, author, score, reply depth and parent post on each one — including replies hidden behind "load more comments". Point it at post links or whole subreddits. No Reddit account or login needed. Pay only for the rows you keep.

Pricing

from $0.35 / 1,000 comment or post rows

Rating

0.0

(0)

Developer

Hamza

Hamza

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Share

Turn any Reddit discussion into a clean, analysable dataset. Give it post links or whole subreddits and it returns one row per comment — the text, the author, the score, how deep in the reply chain it sits, and the parent post it belongs to. It also opens the branches Reddit hides behind "load more comments", so you get the parts of the conversation a normal copy-paste never reaches. No Reddit account, no login, nothing to install.

What you can do with it

  • Voice-of-customer research. Pull every comment on threads about your product, a competitor, or a category, and read what people say unprompted.
  • Sentiment and topic analysis at scale. Feed thousands of clean comment rows into your own model or spreadsheet instead of scraping screenshots.
  • Community and moderation reporting. Track discussion volume, reply depth and the most-upvoted answers across a set of subreddits, on a schedule.
  • Q&A and support mining. Harvest the highest-scoring answers to recurring questions ("best X for Y") to build FAQs, docs or training data.
  • Influencer and advocate discovery. See who consistently writes the top-scoring comments in the communities you care about.
  • Content research. Take the top 10 hot posts in a subreddit and collect every reply to see which angles resonate before you write.

What you get

One record per comment. Abridged real output:

{
"id": "oyzxp83",
"fullId": "t1_oyzxp83",
"type": "comment",
"body": "Ignoring dental care.\n\nPeople in their 20s can skip cleanings and think, \"My teeth are fine.\"…",
"url": "https://www.reddit.com/r/AskReddit/comments/1v32t70/people_who_are_40_what_is_a_silent_killer_habit/oyzxp83/",
"author": "lonelygayPhD",
"authorFlair": null,
"score": 25296,
"scoreHidden": false,
"controversiality": 0,
"totalAwards": 0,
"createdAt": "2026-07-22T02:20:36.000Z",
"editedAt": null,
"subreddit": "AskReddit",
"parentId": "t3_1v32t70",
"isTopLevel": true,
"isSubmitter": false,
"depth": 0,
"replyCount": 19,
"stickied": false,
"collapsed": false,
"postId": "1v32t70",
"postTitle": "People who are 40+, what is a \"silent killer\" habit that people in their 20s think is completely harmless…",
"postUrl": "https://www.reddit.com/r/AskReddit/comments/1v32t70/people_who_are_40_what_is_a_silent_killer_habit/",
"postAuthor": "The-Irumporai",
"postScore": 15757,
"postNumComments": 6251,
"postCreatedAt": "2026-07-22T02:17:17.000Z",
"threadSort": "top",
"threadPosition": 1,
"fromExpansion": false,
"scrapedAt": "2026-07-28T20:35:43.112Z"
}

Every row stands on its own — the parent post's title, link, author and score are repeated on each comment, so you can drop the dataset straight into a spreadsheet, a database or a model without joining anything back together.

Input reference

FieldTypeDefaultWhat it does
postUrlslist of stringsThreads to scrape. Full post links, share links or bare post ids (1v32t70).
subredditslist of stringsemptyOptional. Take the latest posts from these communities and scrape their comments. Accepts python, r/python or a community link.
postsPerSubredditinteger10How many posts to take from each community (max 200). Pinned announcements are skipped.
subredditSortselecthotWhich posts to take: hot, new, top, rising or controversial.
sortselecttopOrder the thread is read in: top, new, best, controversial, old or qa. Decides which comments you get first when you cap the number.
maxCommentsPerPostinteger200Stop after this many comments in one thread (max 5000). Your main cost control.
maxDepthinteger10How deep to follow reply chains. 0 keeps only top-level comments, 1 adds their direct replies (max 20).
expandAllCommentsbooleantrueOpen the collapsed "load more comments" branches. Turn off for a faster, cheaper skim of the visible thread.
maxExpansionRequestsPerPostinteger20How hard to dig into one thread. Each extra round uncovers more of the hidden replies (max 200).
minScoreintegeremptyKeep only comments with at least this many upvotes.
includeOpCommentsbooleantrueKeep comments written by the person who submitted the post.
skipDeletedbooleantrueDrop the empty [deleted] and [removed] placeholders.
includePostRecordbooleanfalseAdd one extra row per thread holding the post itself (title, body, score, media, thread statistics).
maxConcurrencyinteger3How many posts to work on at the same time (max 10).
countryselectusWhich country's view of Reddit the run should use.

You can use postUrls, subreddits, or both in the same run.

Output fields

Comment rows

FieldTypeDescription
id / fullIdstringReddit's comment id (oyzxp83) and full name (t1_oyzxp83).
typestringAlways comment for comment rows.
body / bodyHtmlstringThe comment text as plain text, and the same text with Reddit's formatting and links preserved.
url / permalinkstringDirect link to the comment.
author / authorId / authorFlairstringWho wrote it, their account id and their flair in that community.
score / scoreHiddennumber / booleanNet upvotes, and whether the community hides scores on new comments.
controversiality / totalAwardsnumberReddit's controversy flag and the number of awards received.
createdAt / editedAtdateWhen it was posted and, if applicable, edited.
subreddit / subredditIdstringCommunity it was posted in.
parentId / isTopLevelstring / booleanWhat it replies to, and whether that is the post itself.
depth / replyCountnumberHow deep in the reply chain it sits, and how many direct replies were captured. replyCount is empty on rows recovered from a hidden branch, because Reddit does not report a reply count for those.
isSubmitterbooleanTrue when the comment's author also submitted the post.
distinguished / stickied / locked / collapsedstring / booleanModerator and Reddit display flags.
postId / postTitle / postUrlstringThe parent post, repeated on every row so each one stands alone.
postAuthor / postScore / postNumComments / postCreatedAt / postFlairmixedParent-post context for filtering and grouping.
threadSort / threadPositionstring / numberThe order the thread was read in, and this comment's position within it.
fromExpansionbooleanTrue when the comment was recovered from a hidden "load more" branch.
scrapedAtdateWhen this run read the thread.

Post rows (only when includePostRecord is on) carry the full post — title, body, linkUrl, domain, score, upvoteRatio, numComments, flair, over18, spoiler, media (images, galleries, video with its streaming variants), plus a summary of how the thread went: commentsScraped, commentsReportedByReddit, expansionRequests (how many extra digging rounds were spent on it), expandableStubs, continuationStubs and expansionTruncated.

Pricing

Pay per event — you pay for results, not for run time.

  • Comment or post row (apify-default-dataset-item) — $0.0003 each ($0.30 per 1,000). Charged once for every row saved: each comment, and each post row if you turned that on.
  • Full thread expansion (full-thread-expansion) — $0.004. Charged once per post whose hidden "load more" branches were actually opened and produced extra comments. Threads with nothing hidden, and runs with expandAllComments switched off, never trigger it.

So a 500-comment thread is about $0.15 plus $0.004 for the expansion, and a 10,000-comment pull is about $3.00. If you set a budget limit, the run stops cleanly at the limit instead of overshooting it.

Limits and what this actor cannot do

  • Reddit itself caps a community feed at roughly 1,000 posts. When you seed threads from a subreddit, that is the ceiling on how many posts one community can hand over in a run, no matter how high you set postsPerSubreddit. Get more by splitting across sort orders or by feeding post links directly.
  • Very large threads are sampled, not exhausted. A thread with thousands of comments is read as far as your caps allow; maxCommentsPerPost and maxExpansionRequestsPerPost decide how far it digs. When branches are still hidden at the end, the run says so and the post row's expansionTruncated flag is set.
  • "Continue this thread" branches are reported, not followed. Once a reply chain gets very deep, Reddit stops showing it inline and offers a "continue this thread" link instead; that stretch of conversation is only reachable by scraping the sub-thread as a target of its own. The actor counts how many such branches it saw (continuationStubs) instead of pretending they were read.
  • Deleted and removed comments are gone at Reddit's end. Only the [deleted] / [removed] placeholder survives; the original text is not recoverable by anyone. Keep them with skipDeleted: false if you want the gaps in your data.
  • Private, banned and unavailable communities return nothing. They are reported and skipped; one bad target never fails the whole run. Members-only content is out of reach. Quarantined communities have not been tested.
  • No moderator-only data. Mod logs, removal reasons, reports and moderator lists are not part of what Reddit shows the public, so they are not in the output.
  • No live viewer counts. "Users here now" is not available, and post view counts always come back empty. Judge activity from comment volume and timestamps instead.
  • Scores can be hidden. Communities may hide comment scores for the first hours of a thread; those rows carry scoreHidden: true and a placeholder score.
  • Very heavy runs are paced by Reddit itself. Large jobs slow down rather than fail; keep maxConcurrency moderate if you are scraping hundreds of threads in one go.

FAQ

Do I need a Reddit account? No. You do not need a Reddit account, a login, or any credentials — just enter what you want and run it.

How fast is it? Fast. Comments arrive in large batches rather than one at a time, and several posts are worked on in parallel. In testing, the first 300 comments of a 6,000-comment thread took about five seconds, and small threads finish in well under a second each.

What is the difference between maxCommentsPerPost and maxDepth? maxCommentsPerPost limits how many comments you collect per thread; maxDepth limits how far down reply chains you follow. Set maxDepth: 0 for a clean list of top-level answers only — ideal for "ask" style threads.

Can I run it on a schedule? Yes. A common pattern is a daily run over a handful of subreddits with subredditSort: new, which keeps a rolling archive of fresh discussion. Each run is independent and always returns the current state of the threads it reads.

Why did I get fewer comments than the post's comment count? Reddit's headline count includes deleted placeholders and branches it will not hand over all at once. Raise maxCommentsPerPost and maxExpansionRequestsPerPost, set maxDepth higher, and turn skipDeleted off if you want the placeholders too. The post row's commentsScraped versus commentsReportedByReddit shows exactly how much of a thread you captured.

Can I scrape one branch of a discussion rather than a whole post? Paste the post link and use minScore or maxDepth to narrow it down; every row carries parentId, so you can rebuild any branch of the conversation yourself.

What can I export it as? Anything the platform offers — JSON, CSV, Excel or XML — or read the dataset straight into your own tooling.