> ## Documentation Index
> Fetch the complete documentation index at: https://docs.particle.pro/llms.txt
> Use this file to discover all available pages before exploring further.

# Episode archive

> Bulk download of every transcribed podcast episode, kept current daily, for loading our full history before following the stream.

The episode archive is the complete history of transcribed podcast episodes, delivered as gzip-compressed JSON Lines files. Load it once to backfill your system, then follow the [episode stream](/podcasts/stream) from where the archive leaves off. Each line is an episode exactly as the stream delivers it with `include=all`, so one piece of code can handle both.

<Note>The episode archive is available **by arrangement**: we enable it for your organization. Authenticate with your API key, as for the stream. A key from an organization without access is rejected with [`403 episode_archive_not_enabled`](/errors/episode_archive_not_enabled).</Note>

```
GET https://api.particle.pro/v1/podcasts/episodes/archive
GET https://api.particle.pro/v1/podcasts/episodes/archive/download?period=2026-09
GET https://api.particle.pro/v1/podcasts/episodes/archive/stream?period=2026-09
```

| Scope | History | Updated |
| - | - | - |
| Every transcribed episode of every podcast we cover | Everything we hold, including back catalogues imported in bulk, which the stream never carries | Daily, shortly after midnight UTC. Late episodes, imported history and corrections are folded into the days they belong to |

## What's in it

* **Every transcribed episode.** An episode is in the archive once it has a transcript, with its segments and clips if it has them.
* **Filed by publication date.** The archive is divided into days and months of UTC publication date. An episode published in March 2024 is in `2024-03`, whenever we ingested it.
* **One episode per line**, in publication order, as the stream's `episode` object with `include=all`. Fields that the stream adds over time appear in the archive within its refresh cycle.
* **Every file is a complete snapshot of its period.** When you download a day or month again, replace your copy of it. An episode missing from the new copy was removed or moved to another day.

## Load the archive

1. **List what there is.** The listing gives every month with its days, their episode counts and sizes, and a `version` for each.

   ```bash theme={"dark"}
   curl -s -H "X-API-Key: $PARTICLE_API_KEY" \
     https://api.particle.pro/v1/podcasts/episodes/archive
   ```

   ```json theme={"dark"}
   {
     "as_of": "2026-10-01T01:30:05Z",
     "stream_since": "2026-10-01T01:20:05Z",
     "months": [
       {
         "month": "2026-09",
         "episodes": 1234,
         "bytes": 56789012,
         "version": "5f0c…",
         "updated_at": "2026-10-01T01:41:12Z",
         "days": [
           { "date": "2026-09-01", "episodes": 41, "bytes": 1890123, "version": "a91e…", "updated_at": "2026-10-01T01:33:40Z" }
         ]
       }
     ]
   }
   ```

2. **Download the months or days you want.** Requests are independent, so download several in parallel.

   ```bash theme={"dark"}
   curl -s -H "X-API-Key: $PARTICLE_API_KEY" -o 2026-09.jsonl.gz \
     "https://api.particle.pro/v1/podcasts/episodes/archive/download?period=2026-09"
   zcat 2026-09.jsonl.gz | head -1 | jq '{id, title, published_at}'
   ```

3. **Follow the stream from `stream_since`.** The archive holds every change recorded up to `as_of`. Open the stream from `stream_since`, just before it, so nothing falls between the two. Some episodes will arrive twice; deduplicate on the episode `id`, as the stream already requires.

   ```
   GET /v1/podcasts/episodes/stream?milestone=fully_ingested&include=all&since=2026-10-01T01:20:05Z
   ```

`as_of` and `stream_since` are `null` until the archive's first full build has completed.

## Keep your copy current

The stream carries new episodes but not history: a back catalogue we import, a correction to an old transcript, or an episode we remove. Those reach you through the archive. Once a day, list it with `updated_since` set to the `as_of` of your previous sync, then download and replace each day it returns:

```bash theme={"dark"}
curl -s -H "X-API-Key: $PARTICLE_API_KEY" \
  "https://api.particle.pro/v1/podcasts/episodes/archive?updated_since=2026-10-01T01:30:05Z"
```

With `updated_since`, each month lists only the days whose content changed, while its totals still describe the whole month. A day listed with `"episodes": 0` no longer has any episodes; delete your copy. A day's `version` changes only when its content does.

## Downloads

`GET /v1/podcasts/episodes/archive/download?period=…` takes a month (`2026-09`) or a day (`2026-09-29`) and returns `application/gzip`. A month is its days concatenated, which every gzip reader (`zcat`, Python's `gzip`, Go's `compress/gzip`) reads as one stream.

* **Resuming.** A day supports `Range` requests, so `curl -C -` resumes it. A month does not. On an unreliable connection, download days.
* **Unchanged files.** A full download's `ETag` is the period's `version`; a download narrowed with `include` has an `ETag` of its own. Send the `ETag` you received back in `If-None-Match` to get `304 Not Modified` when nothing changed.
* **Less data.** `include` keeps only the relations you name, from `transcript`, `segments` and `clips`; the default is all of them, unlike the stream, where it is none. Every other field is unchanged. A narrowed download is assembled as it is sent, so it is slower and has no `Content-Length` or `Range`.

```python theme={"dark"}
import gzip, json

with gzip.open("2026-09.jsonl.gz", "rt") as f:
    for line in f:  # lines can be several megabytes
        episode = json.loads(line)
        upsert(episode["id"], episode)
```

## Replay over Server-Sent Events

`GET /v1/podcasts/episodes/archive/stream?period=…` sends the same episodes as Server-Sent Events, for a consumer built around the stream that wants to load history through the same code. Each `episode` event carries a `cursor` and the `episode`. A `complete` event ends a full replay, so you can tell it from a dropped connection. To resume, reconnect with `cursor` set to the last one you processed. If the period was rebuilt in the meantime (the daily update, shortly after midnight UTC), the replay answers with an `error` event instead, because the rebuild may have added episodes before your cursor; start that period again without a cursor and deduplicate on `id`. `include` works as for downloads.

```
event: episode
data: {"cursor":"MTcyNjE…","episode":{"id":"…","title":"…","transcript":{…}}}

event: complete
data: {"period":"2026-09","episodes":1234}
```

For loading the whole archive, downloads are several times faster: they are compressed and can run in parallel.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.