Skip to main content
The episode archive is the complete history of transcribed podcast episodes, delivered as gzip-compressed JSON Lines files. Load it once to backfill your system, then follow the episode stream from where the archive leaves off. Each line is an episode exactly as the stream delivers it with include=all, so one piece of code can handle both.
The episode archive is available by arrangement: we enable it for your organization. Authenticate with your API key, as for the stream. A key from an organization without access is rejected with 403 episode_archive_not_enabled.

What’s in it

  • Every transcribed episode. An episode is in the archive once it has a transcript, with its segments and clips if it has them.
  • Filed by publication date. The archive is divided into days and months of UTC publication date. An episode published in March 2024 is in 2024-03, whenever we ingested it.
  • One episode per line, in publication order, as the stream’s episode object with include=all. Fields that the stream adds over time appear in the archive within its refresh cycle.
  • Every file is a complete snapshot of its period. When you download a day or month again, replace your copy of it. An episode missing from the new copy was removed or moved to another day.

Load the archive

  1. List what there is. The listing gives every month with its days, their episode counts and sizes, and a version for each.
  2. Download the months or days you want. Requests are independent, so download several in parallel.
  3. Follow the stream from stream_since. The archive holds every change recorded up to as_of. Open the stream from stream_since, just before it, so nothing falls between the two. Some episodes will arrive twice; deduplicate on the episode id, as the stream already requires.
as_of and stream_since are null until the archive’s first full build has completed.

Keep your copy current

The stream carries new episodes but not history: a back catalogue we import, a correction to an old transcript, or an episode we remove. Those reach you through the archive. Once a day, list it with updated_since set to the as_of of your previous sync, then download and replace each day it returns:
With updated_since, each month lists only the days whose content changed, while its totals still describe the whole month. A day listed with "episodes": 0 no longer has any episodes; delete your copy. A day’s version changes only when its content does.

Downloads

GET /v1/podcasts/episodes/archive/download?period=… takes a month (2026-09) or a day (2026-09-29) and returns application/gzip. A month is its days concatenated, which every gzip reader (zcat, Python’s gzip, Go’s compress/gzip) reads as one stream.
  • Resuming. A day supports Range requests, so curl -C - resumes it. A month does not. On an unreliable connection, download days.
  • Unchanged files. A full download’s ETag is the period’s version; a download narrowed with include has an ETag of its own. Send the ETag you received back in If-None-Match to get 304 Not Modified when nothing changed.
  • Less data. include keeps only the relations you name, from transcript, segments and clips; the default is all of them, unlike the stream, where it is none. Every other field is unchanged. A narrowed download is assembled as it is sent, so it is slower and has no Content-Length or Range.

Replay over Server-Sent Events

GET /v1/podcasts/episodes/archive/stream?period=… sends the same episodes as Server-Sent Events, for a consumer built around the stream that wants to load history through the same code. Each episode event carries a cursor and the episode. A complete event ends a full replay, so you can tell it from a dropped connection. To resume, reconnect with cursor set to the last one you processed. If the period was rebuilt in the meantime (the daily update, shortly after midnight UTC), the replay answers with an error event instead, because the rebuild may have added episodes before your cursor; start that period again without a cursor and deduplicate on id. include works as for downloads.
For loading the whole archive, downloads are several times faster: they are compressed and can run in parallel.