← Catalog · Reddit RAG dump · Actor
How to get a historical Reddit RAG pack after Pushshift (without a 4TB torrent)
18 Aug 2026 · Updated 5 Sep 2026 · benthepythondev
Public Pushshift is gone. The official Reddit listing APIs still stop around a thousand posts, which on a busy subreddit is a few weeks. The leftover options in 2026 are a volunteer API that blinks, a multi-terabyte torrent, or a small structured pull you can actually embed.
I run the third one.
The job
A retrieval system needs what r/X said, not what is hot this morning. For most packs that means:
- one or two subreddits
- a date window (90 days is the default that fits a weekly refresh)
- posts and comments
title/selftext/body/permalink/created_iso- a
typefield so you can split posts from comments
You do not need every comment Reddit ever stored. You need a window you can re-run.
What I do not recommend first
Academic Torrents and Arctic Shift monthly files are the right tool if you already have disk and a DuckDB habit. Watchful1’s dump is measured in terabytes. That is a warehouse project, not a RAG pack.
PullPush still speaks a Pushshift-shaped API. It also rate-limits and goes down. Fine behind a failover. Bad as the only pipe you sell to yourself.
Live Store scrapers are for monitoring. They see what a visitor sees today. They do not walk years of archive.
The pack I actually ship
Actor: Reddit Archive Scraper — Historical Posts for RAG
It pages PullPush and Arctic Shift, fails over when one archive is sick, and writes one row per post or comment. Public example: 90-day r/Python RAG pack.
On 5 September, build 1.0.18 returned 250 unique posts in each pagination direction in separate January 2024 r/Python checks. A comment run returned three posts and six comments, with nine result charges and three thread charges. Inspect those nine sample rows and the exact input; text excerpts are shortened to 500 characters. Subsequent documentation builds retain that runtime.
Free-tier Actor fees are $3 per 1,000 saved posts or comments, plus $0.005 for each post whose comment fetch returns eligible comments. Gold prices are $2.40 per 1,000 rows and $0.004 per non-empty thread. The start event costs $0.00005 Free or $0.00004 Gold per GB of memory, minimum one. At 512 MB, the three-post/six-comment check is $0.04205 before discounts. Set a maximum run charge in Apify.
If you want a scoped backfill delivered for you, the packaged dump starts with a $250 sample, credited toward a larger order. Service prices cover sampling, export checks and delivery; Apify usage is separate unless the quote includes a cap.
Dump prices Run the r/Python sample
A tight first input
{
"subreddits": ["Python"],
"afterDate": "2024-01-01",
"beforeDate": "2024-01-02",
"sortOrder": "oldest",
"maxPosts": 3,
"includeComments": true,
"maxCommentsPerPost": 2,
"minScore": 0
}
This is the tested sample input. The post window starts at 00:00 UTC on 1 January and ends before 00:00 UTC on 2 January; comments may have later dates. maxPosts limits posts, while comments add rows and thread fees. Check fields before increasing either cap.
Import the n8n workflow
Download reddit-archive-rag-sheets-slack.json and import it into n8n. It reads the last 14 complete UTC days, caps the run at 25 posts with three comments per post, and sets a $0.50 maximum Actor charge. The workflow starts inactive.
- Configure your Apify credential and your own Google Sheet. Use the column names in the workflow note, including
record_key. Set your subreddit in Run Reddit Archive Actor. - Run manually and inspect the output. Posts keep
selftext; comments keepbody. Sheets uses the stabletype:idkey to update existing rows during overlapping refreshes. - Configure the Slack webhook or remove that optional node, then enable the schedule after checking a small run. Credentials and destination writes must be tested in your account.
The workflow prepares Markdown for retrieval; connect your own embedding or vector-store step. The Actor deduplicates within a run, while Sheets handles repeats between runs. Archive gaps and ingestion delays can still leave missing content.
What this is not
It is not a Pushshift clone. It is not user-history search. It is not comment-body full-text search across all of Reddit. Coverage follows the public archives. If a month is missing upstream, it is missing in the pack.
Use it for research and retrieval corpora. Do not use it to harass people.
Try it
- Run the r/Python example Task.
- If the fields fit, either schedule the Actor or send the subreddit + window from the dump page.
- I will not email you first.
Affiliate if you still need an Apify account: apify.com/?fpr=qolupv. I may earn a commission; the price you pay does not change.