Pagination

Parallel and serial pagination

Paged operations

Some GitHub API operations return their results one page at a time. For instance, there are many thousands of gists, but if we call list_public we only see the first 30:

api = GhApi()
gists = await api.gists.list_public()
len(gists)
30

That’s because this operation takes two optional parameters, per_page, and page:

api.gists.list_public

List public gists

Docs: https://docs.github.com/rest/gists/gists#list-public-gists

Parameters: - since (str, optional): Only show results that were last updated after the given time. This is a timestamp in ISO 8601 format: YYYY-MM-DDTHH:MM:SSZ. - per_page (int, default: 30): The number of results per page (max 100). For more information, see “Using pagination in the REST API.” - page (int, default: 1): The page number of the results to fetch. For more information, see “Using pagination in the REST API.”

This is a common pattern for list_* operations in the GitHub API. One way to get more results is to increase per_page:

len(await api.gists.list_public(per_page=100))
100

However, per_page has a maximum of 100, so if you want more, you’ll have to pass page= to get pages beyond the first. An easy way to iterate through all pages is to use paged, which returns an async generator:


source

paged

def paged(
    oper, *args, per_page:int=30, max_pages:int=9999, **kwargs
):

Convert operation oper(*args,**kwargs) into an async iterator, requesting pages serially until one comes back empty

We’ll demonstrate this using the repos.list_for_org method:

api.repos.list_for_org

List organization repositories

Docs: https://docs.github.com/rest/repos/repos#list-organization-repositories

Parameters: - org (str, required): The organization name. The name is not case sensitive. - direction (str, optional): The order to sort by. Default: asc when using full_name, otherwise desc. - type (str, default: ‘all’): Specifies the types of repositories you want returned. - sort (str, default: ‘created’): The property to sort the results by. - per_page (int, default: 30): The number of results per page (max 100). For more information, see “Using pagination in the REST API.” - page (int, default: 1): The page number of the results to fetch. For more information, see “Using pagination in the REST API.”

repos = await api.repos.list_for_org(org='fastai')
len(repos),repos[0].name
(30, 'docs')

To convert this operation into a Python iterator, pass the operation itself, along with any arguments (either keyword or positional) to paged. Note how the function and arguments are passed separately:

repos = paged(api.repos.list_for_org, org='fastai')

The object returned from paged is an async generator, so iterate through it with async for:

async for page in repos: print(len(page), page[0].name)
30 docs
30 nbdev_template
30 hugo
30 course22
4 lm-hackers

Getting pages in parallel

Rather than requesting each page one at a time, we can save some time by getting all the pages we need in parallel.


source

GhApi.last_page

def last_page():

Parse RFC 5988 link header from most recent operation, and extract the last page

To help us know the number of pages needed, we can use last_page, which uses the link header we just looked at to grab the last page from GitHub.

We will need multiple pages to get all the repos in the github organization, even if we get 100 at a time:

await api.repos.list_for_org('github', per_page=100)
api.last_page()
6

source

pages

async def pages(
    oper, n_pages, *args, per_page:int=100, **kwargs
):

Get n_pages pages from oper(*args,**kwargs), in parallel

pages by default passes per_page=100 to the operation.

Let’s look at some examples. To get all the pages for the repos in the github organization in parallel, we can use this:

gh_repos = (await pages(api.repos.list_for_org, api.last_page(), 'github')).concat()
len(gh_repos)
554

If you already know ahead of time the number of pages required, there’s no need to call last_page. For instance, the GitHub docs specify that we can get at most 3000 gists:

gists = (await pages(api.gists.list_public, 30)).concat()
len(gists)
3000

GitHub ignores the per_page parameter for some API calls, such as listing public events, which it limits to 8 pages of 30 items per page. To retrieve all pages in these cases, you need to explicitly pass the lower per page limit:

await api.activity.list_public_events()
api.last_page()
10
evts = (await pages(api.activity.list_public_events, api.last_page(), per_page=30)).concat()
len(evts)
292

Sync clients

On a GhApi(sync=True) client (see core’s “Sync usage” section) endpoint calls return pages directly, so pagination gets sync twins. sync_paged is paged as a plain generator:


source

sync_paged

def sync_paged(
    oper, *args, per_page:int=30, max_pages:int=9999, **kwargs
):

paged for a sync=True client: request pages serially until one comes back empty

sapi = GhApi(sync=True)
pgs = L(sync_paged(sapi.repos.list_for_org, org='fastai'))
test_eq(len(pgs[0]), 30)
assert len(pgs) > 1

sync_pages mirrors pages, fetching a known number of pages in parallel; with no event loop to gather on, it uses a thread pool instead, which is safe because every request creates its own HTTP client. last_page works unchanged on a sync client, since it only parses the stored link header.


source

sync_pages

def sync_pages(
    oper, n_pages, *args, per_page:int=100, n_workers:int=16, **kwargs
):

Get n_pages pages from oper(*args,**kwargs) on a sync=True client, in parallel via threads

gists_s = sync_pages(sapi.gists.list_public, 3).concat()
test_eq(len(gists_s), 300)

GH Notifications


source

gh_notifs

async def gh_notifs(
    days:int=5, reasons:tuple=('mention', 'review_requested', 'author', 'assign'), include_read:bool=False,
    per_page:int=100
):

Get GitHub notifications from the past days, formatted as a summary string

await gh_notifs(2, include_read=True)
'**2 notifications** (past 2 days; mention, review_requested, author, assign)\n\n- [PullRequest #5](https://github.com/AnswerDotAI/fastcflare/pull/5) AnswerDotAI/fastcflare — Use new fastspec `AttrDict` response (assign)\n- [PullRequest #14](https://github.com/AnswerDotAI/fastspec/pull/14) AnswerDotAI/fastspec — fix: walk anyOf/oneOf in _schema_props_required (review_requested)'