How a sync works
- First load. Walk
GET /reportspage by page. Store each report's metadata and download its PDF. - Incremental sync. On a schedule, read
GET /reports/changesfrom where you stopped. Apply each change and download only PDFs whose hash changed. - Reconciliation. Once a month, walk
GET /reportsagain and flag any report you hold that is no longer listed.
Use report_id as your primary key. Use content_sha256 to decide whether a PDF needs downloading.
The first load
GET /reports lists every report in your licence, ordered by report_id, up to 200 per page. Follow page_info.next_cursor until has_more is false. You can narrow the load with filters such as type, fcyear or company_key.
curl -s "https://api.sustainabilityreports.com/api/v2/reports?limit=200" -H "Authorization: Bearer $SRC_API_KEY"
curl -s "https://api.sustainabilityreports.com/api/v2/reports?limit=200&cursor=<next_cursor>" -H "Authorization: Bearer $SRC_API_KEY"Note the time just before you start. Your first incremental sync starts from that moment, so nothing that changes during the load is missed. Reports with has_file: false have metadata but no PDF yet.
Signed download links
GET /reports/{report_id}/download returns a signed HTTPS link to the PDF. Fetch the PDF from that link with a plain GET, without your API key.
- A link is valid until
expires_at: up to 15 minutes, and never beyond your key's own expiry. Request a fresh link for each attempt rather than storing links. - The link is a credential. Do not log it, share it or embed it in an end-user product.
- Verify every file against
content_sha256before you keep it. A metadata update does not change the PDF, so an unchanged hash means no download. - Each link request is recorded against your key and counts as a download in
GET /meta/usage. If that record cannot be written, the API returns503withRetry-Afterand no link.
Staying current with the change feed
GET /reports/changes lists every report in your licence that was added, updated or withdrawn after a point in time, ordered by change time.
- First run:
GET /reports/changes?since=2026-10-01T00:00:00Z(the time your first load started), then follownext_cursoruntilhas_moreis false. - Store the last
next_cursor. It is present on every response, including an empty page. - Next run:
GET /reports/changes?cursor=<stored cursor>. A cursor takes precedence oversince.
{
"data": [
{"change": "upsert", "report_id": 529769, "report_key": "XMIL_G_2025_AR2", "company_key": "XMIL_G",
"changed_at": "2026-10-08T03:12:44Z", "content_sha256": "a1d0c6e8...", "report": { "...": "full Report object" }},
{"change": "withdrawn", "report_id": 98765, "report_key": "XMIL_G_2025_AR1", "company_key": "XMIL_G",
"changed_at": "2026-10-08T04:01:02Z", "content_sha256": null, "report": null}
],
"page_info": {"count": 2, "has_more": false, "next_cursor": "eyJ1IjoiMjAyNi0xMC0wOFQwNDowMTowMloiLCJpIjo5ODc2NX0"}
}upsert: a report was added or changed. The full report is attached. Download the PDF again only ifcontent_sha256differs from the one you hold.withdrawn: we no longer serve the report, for example a duplicate or a misattributed document. Remove or flag it. Ignore withdrawals of reports you never held.- Delivery is at least once: the same change can appear twice. Apply changes idempotently by
report_id. - Changes appear after a settling delay of at least 15 minutes, longer while a large catalogue update is still being written. The feed lags real time slightly but never skips a change that was in progress.
- When companies are added to your licence, their existing reports come through the feed at that moment. You do not need a separate pull.
- Now and then a catalogue-wide maintenance update re-stamps many records at once. These appear as upserts with an unchanged
content_sha256, so they cost no downloads.
Monthly reconciliation
Two kinds of change are not in the feed:
- a report removed outright from our catalogue, which is rare, for example a duplicate merged into another record;
- a report that leaves your scope because it was reclassified to another fiscal year or report type.
Once a month, walk GET /reports and flag any report_id you hold that no longer appears. If you prefer to check a known list, POST /reports/lookup resolves up to 1,000 report_key values per call. A key that was renamed resolves to the current report.
Volume and fair use
- Report PDFs average a few megabytes, so a licence covering thousands of companies can amount to hundreds of gigabytes. Run large first loads between 00:00 and 06:00 UTC where you can.
- Use up to 8 parallel connections and stay below 10 requests per second from one IP address. See rate limits.
- There are no per-call or per-download fees. Fair use means syncing and downloading what you need for your licensed purposes. Downloading unchanged PDFs again and again is not.
- Set a
User-Agentthat names your integration, for exampleexample-sync/1.0. - Your licence sets out how you may use and share the reports. Signed links and the dataset itself are not for redistribution.
A complete Python client
A minimal sync client using requests. It handles pagination, rate limits, transient errors and network failures, and verifies every PDF against content_sha256 before it keeps it. Partial files never replace good ones. Run full_load() once, then weekly_sync() on a schedule.
import hashlib
import os
import pathlib
import tempfile
import time
from datetime import datetime, timezone
import requests
BASE = "https://api.sustainabilityreports.com/api/v2"
OUT = pathlib.Path("pdfs")
S = requests.Session()
S.headers.update({
"Authorization": f"Bearer {os.environ['SRC_API_KEY']}",
"User-Agent": "example-sync/1.0",
})
TRANSIENT = {429, 500, 502, 503, 504}
class SyncError(Exception):
pass
def call(method, path, **kw):
"""One API call, retried on 429, transient 5xx and network errors."""
for attempt in range(6):
try:
r = S.request(method, BASE + path, timeout=60, **kw)
except requests.RequestException:
time.sleep(2 ** attempt)
continue
if r.status_code in TRANSIENT:
time.sleep(int(r.headers.get("Retry-After", 2 ** attempt)))
continue
if r.status_code >= 400:
raise SyncError(f"{method} {path}: HTTP {r.status_code} {r.text[:300]}")
return r.json()
raise SyncError(f"{method} {path}: gave up after retries")
def pages(path, params):
"""Yield items; the final next_cursor is returned as StopIteration.value."""
params = dict(params)
while True:
page = call("GET", path, params=params)
yield from page["data"]
cursor = page["page_info"].get("next_cursor")
if not page["page_info"]["has_more"]:
return cursor
params = {k: v for k, v in params.items() if k != "since"}
params["cursor"] = cursor
def sha256_of(path):
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
return h.hexdigest()
def download_pdf(report):
"""Fetch, verify and atomically store one PDF. Never logs the signed URL."""
OUT.mkdir(exist_ok=True)
target = OUT / report["filename"]
expected = report.get("content_sha256")
if target.exists() and expected and sha256_of(target) == expected:
return target # unchanged, skip
for attempt in range(4):
link = call("GET", f"/reports/{report['report_id']}/download") # fresh link each try
tmp = None
try:
with requests.get(link["url"], stream=True, timeout=300) as r: # no API key here
if r.status_code != 200:
raise SyncError(f"PDF fetch for report {report['report_id']}: HTTP {r.status_code}")
fd, tmp = tempfile.mkstemp(dir=OUT, suffix=".part")
with os.fdopen(fd, "wb") as f:
for chunk in r.iter_content(1 << 20):
f.write(chunk)
except (requests.RequestException, SyncError):
if tmp and os.path.exists(tmp):
os.remove(tmp) # no partial files left behind
time.sleep(2 ** attempt)
continue
if expected and sha256_of(tmp) != expected:
os.remove(tmp)
time.sleep(2 ** attempt)
continue
os.replace(tmp, target) # atomic
return target
raise SyncError(f"PDF for report {report['report_id']} failed verification or download")
def full_load():
started = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
for rep in pages("/reports", {"limit": 200}):
store_metadata(rep) # your code: upsert by report_id
if rep["has_file"]:
download_pdf(rep)
save_state(since=started, cursor=None) # your code: persist both
def weekly_sync():
state = load_state() # your code
params = {"cursor": state["cursor"]} if state["cursor"] else {"since": state["since"]}
gen = pages("/reports/changes", {**params, "limit": 100})
while True:
try:
item = next(gen)
except StopIteration as stop:
save_state(since=state["since"], cursor=stop.value) # resume point for next time
return
if item["change"] == "withdrawn":
mark_withdrawn(item["report_id"]) # your code; ignore unknown ids
else:
store_metadata(item["report"])
if item["report"]["has_file"]:
download_pdf(item["report"])store_metadata, mark_withdrawn, save_state and load_state are yours to write, against whatever store you use. The first weekly_sync() starts from the moment full_load() began; the feed is idempotent, so the overlap is harmless.
Need help with a large first load?
Email [email protected] before you start a load of more than a few hundred gigabytes, and tell us the egress IP addresses you will use.
