api = GhApi()Pagination
Paged operations
Some GitHub API operations return their results one page at a time. For instance, there are many thousands of gists, but if we call list_public we only see the first 30:
gists = await api.gists.list_public()
len(gists)30
That’s because this operation takes two optional parameters, per_page, and page:
api.gists.list_publicList public gists
Docs: https://docs.github.com/rest/gists/gists#list-public-gists
Parameters: - since (str, optional): Only show results that were last updated after the given time. This is a timestamp in ISO 8601 format: YYYY-MM-DDTHH:MM:SSZ. - per_page (int, default: 30): The number of results per page (max 100). For more information, see “Using pagination in the REST API.” - page (int, default: 1): The page number of the results to fetch. For more information, see “Using pagination in the REST API.”
This is a common pattern for list_* operations in the GitHub API. One way to get more results is to increase per_page:
len(await api.gists.list_public(per_page=100))100
However, per_page has a maximum of 100, so if you want more, you’ll have to pass page= to get pages beyond the first. An easy way to iterate through all pages is to use paged, which returns an async generator:
paged
def paged(
oper, *args, per_page:int=30, max_pages:int=9999, **kwargs
):Convert operation oper(*args,**kwargs) into an async iterator, requesting pages serially until one comes back empty
We’ll demonstrate this using the repos.list_for_org method:
api.repos.list_for_orgList organization repositories
Docs: https://docs.github.com/rest/repos/repos#list-organization-repositories
Parameters: - org (str, required): The organization name. The name is not case sensitive. - direction (str, optional): The order to sort by. Default: asc when using full_name, otherwise desc. - type (str, default: ‘all’): Specifies the types of repositories you want returned. - sort (str, default: ‘created’): The property to sort the results by. - per_page (int, default: 30): The number of results per page (max 100). For more information, see “Using pagination in the REST API.” - page (int, default: 1): The page number of the results to fetch. For more information, see “Using pagination in the REST API.”
repos = await api.repos.list_for_org(org='fastai')
len(repos),repos[0].name(30, 'docs')
To convert this operation into a Python iterator, pass the operation itself, along with any arguments (either keyword or positional) to paged. Note how the function and arguments are passed separately:
repos = paged(api.repos.list_for_org, org='fastai')The object returned from paged is an async generator, so iterate through it with async for:
async for page in repos: print(len(page), page[0].name)30 docs
30 nbdev_template
30 hugo
30 course22
4 lm-hackers
Link header (RFC 5988)
GitHub tells us how many pages are available using the link header. Unfortunately the pypi LinkHeader library appears to no longer be maintained, so we’ve put a refactored version of it here.
parse_link_hdr
def parse_link_hdr(
header
):Parse an RFC 5988 link header, returning a dict from rels to a tuple of URL and attrs dict
Here’s an example of a link header with just one link:
parse_link_hdr('<http://example.com>; rel="foo bar"; type=text/html'){'foo bar': ('http://example.com', {'type': 'text/html'})}
links = parse_link_hdr('<http://example.com>; rel="foo bar"; type=text/html')
link = links['foo bar']
test_eq(link[0], 'http://example.com')
test_eq(link[1]['type'], 'text/html')Let’s test it on the headers we received on our last call to GitHub. You can access the last call’s headers in `recv_hdrs’:
api.recv_hdrs['Link']'<https://api.github.com/organizations/20547620/repos?per_page=30&page=5>; rel="prev", <https://api.github.com/organizations/20547620/repos?per_page=30&page=5>; rel="last", <https://api.github.com/organizations/20547620/repos?per_page=30&page=1>; rel="first"'
Here’s what happens when we parse that:
parse_link_hdr(api.recv_hdrs['Link']){'prev': ('https://api.github.com/organizations/20547620/repos?per_page=30&page=5',
{}),
'last': ('https://api.github.com/organizations/20547620/repos?per_page=30&page=5',
{}),
'first': ('https://api.github.com/organizations/20547620/repos?per_page=30&page=1',
{})}
Getting pages in parallel
Rather than requesting each page one at a time, we can save some time by getting all the pages we need in parallel.
GhApi.last_page
def last_page():Parse RFC 5988 link header from most recent operation, and extract the last page
To help us know the number of pages needed, we can use last_page, which uses the link header we just looked at to grab the last page from GitHub.
We will need multiple pages to get all the repos in the github organization, even if we get 100 at a time:
await api.repos.list_for_org('github', per_page=100)
api.last_page()6
pages
async def pages(
oper, n_pages, *args, per_page:int=100, **kwargs
):Get n_pages pages from oper(*args,**kwargs), in parallel
pages by default passes per_page=100 to the operation.
Let’s look at some examples. To get all the pages for the repos in the github organization in parallel, we can use this:
gh_repos = (await pages(api.repos.list_for_org, api.last_page(), 'github')).concat()
len(gh_repos)554
If you already know ahead of time the number of pages required, there’s no need to call last_page. For instance, the GitHub docs specify that we can get at most 3000 gists:
gists = (await pages(api.gists.list_public, 30)).concat()
len(gists)3000
GitHub ignores the per_page parameter for some API calls, such as listing public events, which it limits to 8 pages of 30 items per page. To retrieve all pages in these cases, you need to explicitly pass the lower per page limit:
await api.activity.list_public_events()
api.last_page()10
evts = (await pages(api.activity.list_public_events, api.last_page(), per_page=30)).concat()
len(evts)292
Sync clients
On a GhApi(sync=True) client (see core’s “Sync usage” section) endpoint calls return pages directly, so pagination gets sync twins. sync_paged is paged as a plain generator:
sync_paged
def sync_paged(
oper, *args, per_page:int=30, max_pages:int=9999, **kwargs
):paged for a sync=True client: request pages serially until one comes back empty
sapi = GhApi(sync=True)
pgs = L(sync_paged(sapi.repos.list_for_org, org='fastai'))
test_eq(len(pgs[0]), 30)
assert len(pgs) > 1sync_pages mirrors pages, fetching a known number of pages in parallel; with no event loop to gather on, it uses a thread pool instead, which is safe because every request creates its own HTTP client. last_page works unchanged on a sync client, since it only parses the stored link header.
sync_pages
def sync_pages(
oper, n_pages, *args, per_page:int=100, n_workers:int=16, **kwargs
):Get n_pages pages from oper(*args,**kwargs) on a sync=True client, in parallel via threads
gists_s = sync_pages(sapi.gists.list_public, 3).concat()
test_eq(len(gists_s), 300)GH Notifications
gh_notifs
async def gh_notifs(
days:int=5, reasons:tuple=('mention', 'review_requested', 'author', 'assign'), include_read:bool=False,
per_page:int=100
):Get GitHub notifications from the past days, formatted as a summary string
await gh_notifs(2, include_read=True)'**2 notifications** (past 2 days; mention, review_requested, author, assign)\n\n- [PullRequest #5](https://github.com/AnswerDotAI/fastcflare/pull/5) AnswerDotAI/fastcflare — Use new fastspec `AttrDict` response (assign)\n- [PullRequest #14](https://github.com/AnswerDotAI/fastspec/pull/14) AnswerDotAI/fastspec — fix: walk anyOf/oneOf in _schema_props_required (review_requested)'