← Catalog · All dump and feed prices · Apify Actor
Past the official ~1,000-post cap. I run a sample on your subreddit and window, then deliver JSON, CSV, or JSONL. You do not download a multi-terabyte torrent or keep an archive API alive.
Inbound only. I do not cold-email Store users. New buyers start here or on the Actor page.
People building a retrieval corpus who need what a subreddit said over
months, not what is on the front page this morning. Typical job: last 90
days of r/MachineLearning plus comments, chunked for a vector store.
Live scrapers stop at Reddit’s listing cap. Academic Torrents still
leaves you with a warehouse of .zst files.
| Package | Scope | Price |
|---|---|---|
| Sample | 1 subreddit, 90 days, posts + comments, up to 5,000 rows | $250 (credited if you take a dump) |
| Dump | 1–3 subreddits, named window, up to 250k rows, one delivery | $2,000–$4,000 |
| Fat dump | Large window or multi-sub pack, up to 1M rows | $5,000–$8,000 |
| Refresh | Same filters, weekly or monthly, same sink | $400–$900 / mo |
Row caps are delivery caps, not a promise the archive holds that many items. Thin windows still bill the booked package. You pay Apify usage on your account unless the quote bundles a cap. Store PPE stays $2.40/1k Gold.
Fields that actually land in a RAG pack
title, selftext,
permalink, created_iso,
score, subreddit, author
body, post_id,
parent_id, same timestamps
type = post or
comment
Not in this SKU: user-history scrape, comment-body full-text search, deleted-content recovery, or legal cover for how you use the text.
| Path | Use it when |
|---|---|
| Store Actor | You want a small or weekly pull and will start the run yourself |
| This dump | You want 50k–1M rows delivered once, with a sample first |
| Arctic Shift / Academic Torrents | You already have disk, DuckDB, and time to parse monthly dumps |
gJAFTYmczF0JICN6B: 119 items (25 posts + 94
comments) in ~38s
A paid dump still starts with a new sample on your subreddit. That proof run is not your corpus.
Longer write-up: how to get a historical Reddit RAG pack after Pushshift. Jobs, rentals, and Maps feeds live on the quote hub if Reddit is the wrong source.
Those archives are terabytes of compressed monthly files. You still pick subreddits, parse zst, join comments to posts, and cut a date window. The dump is that work already done for named communities.
No. Public Pushshift is gone. The Actor reads PullPush and Arctic Shift and fails over when one is down. Coverage follows those archives. I do not sell a complete firehose or a deleted-content vault.
Yes. Store PPE is $2.40 per 1,000 results on Gold. Use that for small or recurring pulls. The dump is for a 50k–1M row backfill you do not want to babysit.
No. RAG-only: named subreddits, a date window, posts and comments. No user-history scrape and no comment-body full-text search.
benthepythondev · benthepythondev0@gmail.com · Actor bugs stay on the Issues tab.