cursor; pass it back to get the next page, and stop when it is gone.
The envelope
Request the two most recent episodes of The Daily:Response (truncated)
data holds the page, has_more says whether another page exists, and cursor is the token for it (on the episode feed it is also the position to resume from later, so it is present even when has_more is false). Send the cursor back with the same parameters:
Response (truncated)
has_more: false and no cursor.
The loop
data, has_more, and cursor; only the parameters change. Three endpoints paginate a nested collection under their own key and are listed at the end of this page. Stop on has_more, not on the presence of cursor: the episode feed keeps returning a cursor after has_more turns false, because that cursor is the position to resume from on your next poll, and a loop that stops only when the cursor disappears would fetch empty pages forever. Once the feed says has_more: false, store the cursor and poll again on your own interval.
The cursor
- Treat it as opaque. Do not build, edit, or store cursors long term; take them from the previous response.
- Stop on
has_more. On ordinary listscursordisappears with the last page; on the episode feed it persists as your resume point, sohas_moreis the signal on both. - Keep the parameters the same. A cursor continues the request that produced it. To change a filter, start over without a cursor.
- Cursors belong to their endpoint. A cursor from
/v1/podcasts/episodesdoes not continue/v1/podcasts/mentions. - Persist before you advance. In an import, write a page’s records before saving its cursor as the checkpoint, and key writes by canonical ids so a replayed page is harmless. A walk over a changing catalog is not a snapshot.
- Budget the walk. Each page is one request. Cap the pages a job may fetch, and treat a hit cap as a partial result to report rather than a silent stop.
- A cursor the server cannot read is rejected with
422(validation_error) on the episode feed and the advertising placements list, and treated as the first page on the other lists. Either way, take cursors only from responses; if a walk fails or restarts unexpectedly, check that the cursor was passed through unchanged.
What limit means
limit is the page size, from 1 to 100 (default 25). Ask for 100 when you intend to read everything, so the walk takes as few requests as possible. Each page is one request for rate limiting and metering.
A few endpoints scoped to one episode return the whole set when limit is omitted, because the set is small: segments, speakers, entities, and topics of an episode. GET /v1/podcasts/episodes/{id}/transcript/words accepts a limit up to 5,000.
Which endpoints paginate
Every endpoint that returns a list: podcasts, episodes, the episode feed, search results, mentions, clips, segments, guests, rankings, ratings, sponsors and ad placements, publishers, companies and their people, entities, topics, and alerts and their matches. Their pages share the envelope above. Three endpoints paginate a nested collection under their own key rather thandata, with the same has_more and cursor semantics: the entity mentions of one episode (GET /v1/podcasts/episodes/{id}/transcript/mentions, collection entities), the word-level transcript (GET /v1/podcasts/episodes/{id}/transcript/words, collection words), and the matches embedded in one alert delivery (GET /v1/alerts/deliveries/{id}, collection matches with matches_has_more and matches_cursor). Read the collection key from the response and the loop above applies unchanged.
Endpoints that return one object (a podcast, an episode, a company, a summary, a timeseries) return it whole. Timeseries endpoints such as GET /v1/podcasts/mentions/timeseries return every bucket in the window in one response (up to 1,000 buckets, so pick the interval to fit the window), which is why they are the right tool for “how often over time” questions instead of paging through mentions.
Related
- Concepts for the identifier and error conventions the loop relies on
- Track a company across podcasts for a walk that ends in an alert instead of a loop