Skip to content

API documentation

Bulk PDF downloads

Load every report in your licence once, then keep the collection current by moving only what changed. This guide covers the first load, the change feed and a complete Python client.

API version 2.4.0 · Updated 2 October 2026

How a sync works

  1. First load. Walk GET /reports page by page. Store each report's metadata and download its PDF.
  2. Incremental sync. On a schedule, read GET /reports/changes from where you stopped. Apply each change and download only PDFs whose hash changed.
  3. Reconciliation. Once a month, walk GET /reports again and flag any report you hold that is no longer listed.

Use report_id as your primary key. Use content_sha256 to decide whether a PDF needs downloading.

The first load

GET /reports lists every report in your licence, ordered by report_id, up to 200 per page. Follow page_info.next_cursor until has_more is false. You can narrow the load with filters such as type, fcyear or company_key.

Shell
curl -s "https://api.sustainabilityreports.com/api/v2/reports?limit=200" -H "Authorization: Bearer $SRC_API_KEY"
curl -s "https://api.sustainabilityreports.com/api/v2/reports?limit=200&cursor=<next_cursor>" -H "Authorization: Bearer $SRC_API_KEY"

Note the time just before you start. Your first incremental sync starts from that moment, so nothing that changes during the load is missed. Reports with has_file: false have metadata but no PDF yet.

GET /reports/{report_id}/download returns a signed HTTPS link to the PDF. Fetch the PDF from that link with a plain GET, without your API key.

  • A link is valid until expires_at: up to 15 minutes, and never beyond your key's own expiry. Request a fresh link for each attempt rather than storing links.
  • The link is a credential. Do not log it, share it or embed it in an end-user product.
  • Verify every file against content_sha256 before you keep it. A metadata update does not change the PDF, so an unchanged hash means no download.
  • Each link request is recorded against your key and counts as a download in GET /meta/usage. If that record cannot be written, the API returns 503 with Retry-After and no link.

Staying current with the change feed

GET /reports/changes lists every report in your licence that was added, updated or withdrawn after a point in time, ordered by change time.

  1. First run: GET /reports/changes?since=2026-10-01T00:00:00Z (the time your first load started), then follow next_cursor until has_more is false.
  2. Store the last next_cursor. It is present on every response, including an empty page.
  3. Next run: GET /reports/changes?cursor=<stored cursor>. A cursor takes precedence over since.
GET /reports/changes
{
  "data": [
    {"change": "upsert", "report_id": 529769, "report_key": "XMIL_G_2025_AR2", "company_key": "XMIL_G",
     "changed_at": "2026-10-08T03:12:44Z", "content_sha256": "a1d0c6e8...", "report": { "...": "full Report object" }},
    {"change": "withdrawn", "report_id": 98765, "report_key": "XMIL_G_2025_AR1", "company_key": "XMIL_G",
     "changed_at": "2026-10-08T04:01:02Z", "content_sha256": null, "report": null}
  ],
  "page_info": {"count": 2, "has_more": false, "next_cursor": "eyJ1IjoiMjAyNi0xMC0wOFQwNDowMTowMloiLCJpIjo5ODc2NX0"}
}
  • upsert: a report was added or changed. The full report is attached. Download the PDF again only if content_sha256 differs from the one you hold.
  • withdrawn: we no longer serve the report, for example a duplicate or a misattributed document. Remove or flag it. Ignore withdrawals of reports you never held.
  • Delivery is at least once: the same change can appear twice. Apply changes idempotently by report_id.
  • Changes appear after a settling delay of at least 15 minutes, longer while a large catalogue update is still being written. The feed lags real time slightly but never skips a change that was in progress.
  • When companies are added to your licence, their existing reports come through the feed at that moment. You do not need a separate pull.
  • Now and then a catalogue-wide maintenance update re-stamps many records at once. These appear as upserts with an unchanged content_sha256, so they cost no downloads.

Monthly reconciliation

Two kinds of change are not in the feed:

  • a report removed outright from our catalogue, which is rare, for example a duplicate merged into another record;
  • a report that leaves your scope because it was reclassified to another fiscal year or report type.

Once a month, walk GET /reports and flag any report_id you hold that no longer appears. If you prefer to check a known list, POST /reports/lookup resolves up to 1,000 report_key values per call. A key that was renamed resolves to the current report.

Volume and fair use

  • Report PDFs average a few megabytes, so a licence covering thousands of companies can amount to hundreds of gigabytes. Run large first loads between 00:00 and 06:00 UTC where you can.
  • Use up to 8 parallel connections and stay below 10 requests per second from one IP address. See rate limits.
  • There are no per-call or per-download fees. Fair use means syncing and downloading what you need for your licensed purposes. Downloading unchanged PDFs again and again is not.
  • Set a User-Agent that names your integration, for example example-sync/1.0.
  • Your licence sets out how you may use and share the reports. Signed links and the dataset itself are not for redistribution.

A complete Python client

A minimal sync client using requests. It handles pagination, rate limits, transient errors and network failures, and verifies every PDF against content_sha256 before it keeps it. Partial files never replace good ones. Run full_load() once, then weekly_sync() on a schedule.

sync.py
import hashlib
import os
import pathlib
import tempfile
import time
from datetime import datetime, timezone

import requests

BASE = "https://api.sustainabilityreports.com/api/v2"
OUT = pathlib.Path("pdfs")
S = requests.Session()
S.headers.update({
    "Authorization": f"Bearer {os.environ['SRC_API_KEY']}",
    "User-Agent": "example-sync/1.0",
})
TRANSIENT = {429, 500, 502, 503, 504}


class SyncError(Exception):
    pass


def call(method, path, **kw):
    """One API call, retried on 429, transient 5xx and network errors."""
    for attempt in range(6):
        try:
            r = S.request(method, BASE + path, timeout=60, **kw)
        except requests.RequestException:
            time.sleep(2 ** attempt)
            continue
        if r.status_code in TRANSIENT:
            time.sleep(int(r.headers.get("Retry-After", 2 ** attempt)))
            continue
        if r.status_code >= 400:
            raise SyncError(f"{method} {path}: HTTP {r.status_code} {r.text[:300]}")
        return r.json()
    raise SyncError(f"{method} {path}: gave up after retries")


def pages(path, params):
    """Yield items; the final next_cursor is returned as StopIteration.value."""
    params = dict(params)
    while True:
        page = call("GET", path, params=params)
        yield from page["data"]
        cursor = page["page_info"].get("next_cursor")
        if not page["page_info"]["has_more"]:
            return cursor
        params = {k: v for k, v in params.items() if k != "since"}
        params["cursor"] = cursor


def sha256_of(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(1 << 20), b""):
            h.update(chunk)
    return h.hexdigest()


def download_pdf(report):
    """Fetch, verify and atomically store one PDF. Never logs the signed URL."""
    OUT.mkdir(exist_ok=True)
    target = OUT / report["filename"]
    expected = report.get("content_sha256")
    if target.exists() and expected and sha256_of(target) == expected:
        return target                                       # unchanged, skip
    for attempt in range(4):
        link = call("GET", f"/reports/{report['report_id']}/download")  # fresh link each try
        tmp = None
        try:
            with requests.get(link["url"], stream=True, timeout=300) as r:  # no API key here
                if r.status_code != 200:
                    raise SyncError(f"PDF fetch for report {report['report_id']}: HTTP {r.status_code}")
                fd, tmp = tempfile.mkstemp(dir=OUT, suffix=".part")
                with os.fdopen(fd, "wb") as f:
                    for chunk in r.iter_content(1 << 20):
                        f.write(chunk)
        except (requests.RequestException, SyncError):
            if tmp and os.path.exists(tmp):
                os.remove(tmp)                              # no partial files left behind
            time.sleep(2 ** attempt)
            continue
        if expected and sha256_of(tmp) != expected:
            os.remove(tmp)
            time.sleep(2 ** attempt)
            continue
        os.replace(tmp, target)                             # atomic
        return target
    raise SyncError(f"PDF for report {report['report_id']} failed verification or download")


def full_load():
    started = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
    for rep in pages("/reports", {"limit": 200}):
        store_metadata(rep)                                 # your code: upsert by report_id
        if rep["has_file"]:
            download_pdf(rep)
    save_state(since=started, cursor=None)                  # your code: persist both


def weekly_sync():
    state = load_state()                                    # your code
    params = {"cursor": state["cursor"]} if state["cursor"] else {"since": state["since"]}
    gen = pages("/reports/changes", {**params, "limit": 100})
    while True:
        try:
            item = next(gen)
        except StopIteration as stop:
            save_state(since=state["since"], cursor=stop.value)  # resume point for next time
            return
        if item["change"] == "withdrawn":
            mark_withdrawn(item["report_id"])               # your code; ignore unknown ids
        else:
            store_metadata(item["report"])
            if item["report"]["has_file"]:
                download_pdf(item["report"])

store_metadata, mark_withdrawn, save_state and load_state are yours to write, against whatever store you use. The first weekly_sync() starts from the moment full_load() began; the feed is idempotent, so the overlap is harmless.

Need help with a large first load?

Email [email protected] before you start a load of more than a few hundred gigabytes, and tell us the egress IP addresses you will use.