← Catalog · Reddit RAG dump · Actor

How to get a historical Reddit RAG pack after Pushshift (without a 4TB torrent)

18 Aug 2026 · Updated 5 Sep 2026 · benthepythondev

Public Pushshift is gone. The official Reddit listing APIs still stop around a thousand posts, which on a busy subreddit is a few weeks. The leftover options in 2026 are a volunteer API that blinks, a multi-terabyte torrent, or a small structured pull you can actually embed.

I run the third one.

The job

A retrieval system needs what r/X said, not what is hot this morning. For most packs that means:

You do not need every comment Reddit ever stored. You need a window you can re-run.

What I do not recommend first

Academic Torrents and Arctic Shift monthly files are the right tool if you already have disk and a DuckDB habit. Watchful1’s dump is measured in terabytes. That is a warehouse project, not a RAG pack.

PullPush still speaks a Pushshift-shaped API. It also rate-limits and goes down. Fine behind a failover. Bad as the only pipe you sell to yourself.

Live Store scrapers are for monitoring. They see what a visitor sees today. They do not walk years of archive.

The pack I actually ship

Actor: Reddit Archive Scraper — Historical Posts for RAG

It pages PullPush and Arctic Shift, fails over when one archive is sick, and writes one row per post or comment. Public example: 90-day r/Python RAG pack.

On 5 September, build 1.0.18 returned 250 unique posts in each pagination direction in separate January 2024 r/Python checks. A comment run returned three posts and six comments, with nine result charges and three thread charges. Inspect those nine sample rows and the exact input; text excerpts are shortened to 500 characters. Subsequent documentation builds retain that runtime.

Free-tier Actor fees are $3 per 1,000 saved posts or comments, plus $0.005 for each post whose comment fetch returns eligible comments. Gold prices are $2.40 per 1,000 rows and $0.004 per non-empty thread. The start event costs $0.00005 Free or $0.00004 Gold per GB of memory, minimum one. At 512 MB, the three-post/six-comment check is $0.04205 before discounts. Set a maximum run charge in Apify.

If you want a scoped backfill delivered for you, the packaged dump starts with a $250 sample, credited toward a larger order. Service prices cover sampling, export checks and delivery; Apify usage is separate unless the quote includes a cap.

Dump prices Run the r/Python sample

A tight first input

{
  "subreddits": ["Python"],
  "afterDate": "2024-01-01",
  "beforeDate": "2024-01-02",
  "sortOrder": "oldest",
  "maxPosts": 3,
  "includeComments": true,
  "maxCommentsPerPost": 2,
  "minScore": 0
}

This is the tested sample input. The post window starts at 00:00 UTC on 1 January and ends before 00:00 UTC on 2 January; comments may have later dates. maxPosts limits posts, while comments add rows and thread fees. Check fields before increasing either cap.

Import the n8n workflow

Download reddit-archive-rag-sheets-slack.json and import it into n8n. It reads the last 14 complete UTC days, caps the run at 25 posts with three comments per post, and sets a $0.50 maximum Actor charge. The workflow starts inactive.

  1. Configure your Apify credential and your own Google Sheet. Use the column names in the workflow note, including record_key. Set your subreddit in Run Reddit Archive Actor.
  2. Run manually and inspect the output. Posts keep selftext; comments keep body. Sheets uses the stable type:id key to update existing rows during overlapping refreshes.
  3. Configure the Slack webhook or remove that optional node, then enable the schedule after checking a small run. Credentials and destination writes must be tested in your account.

The workflow prepares Markdown for retrieval; connect your own embedding or vector-store step. The Actor deduplicates within a run, while Sheets handles repeats between runs. Archive gaps and ingestion delays can still leave missing content.

What this is not

It is not a Pushshift clone. It is not user-history search. It is not comment-body full-text search across all of Reddit. Coverage follows the public archives. If a month is missing upstream, it is missing in the pack.

Use it for research and retrieval corpora. Do not use it to harass people.

Try it

  1. Run the r/Python example Task.
  2. If the fields fit, either schedule the Actor or send the subreddit + window from the dump page.
  3. I will not email you first.

Affiliate if you still need an Apify account: apify.com/?fpr=qolupv. I may earn a commission; the price you pay does not change.