Page by page
Every page has a clear boundary and a stable PDF page number, even when the printed page label differs.
Sustainability reports are hard to work with: dense text, complex tables, charts and images, all locked inside PDFs. We convert them into text you can search, parse and feed to a model, and we keep track of which page everything came from. The original PDF comes along, so you can always check.
Every page has a clear boundary and a stable PDF page number, even when the printed page label differs.
Headings, tables, figures, captions and footnotes are labelled, so your code can treat each one differently.
Every order pairs the Markdown with the original PDF and a 50-field metadata record.
Choose a page to compare the source PDF with the Markdown produced by our pipeline.
HMC Capital · 2022 Sustainability Report
Previewing PDF pages 1–5 of 52

:::: {.page n=1}
# 2022 Sustainability Report
Creating Healthy Communities
::: {#1-i-1 .image}
A group of diverse people standing in a circle outdoors, holding each other's shoulders, looking towards the horizon. The image is bathed in warm, golden sunlight, suggesting a community or team effort.
:::
::::Our conversion process is built specifically for sustainability reports. Every page goes through vision OCR. Every figure then gets a second look from a separate vision model, and every table gets its own correction pass.
Our vision OCR process reads each page and identifies text, reading order, tables and visual regions across complex report layouts.
Each detected figure receives a separate vision-model pass for a useful description and type. Tables receive targeted transcription and correction, preserving their evidence and units where available.
We turn the enriched result into page-addressable Markdown, then check page sequence, typed elements, table retention and source binding. Files that don’t pass are held back for review instead of being delivered.
Good to know: we also use these Markdown files in our own data extraction jobs. Our extraction pipeline achieved 98.7% value accuracy in SRC-Bench v1: 226 of 229 reported values matched expert annotations across seven indicators in 50 reports.
Every element is labelled so your code can select the content it needs: charts, tables, images or document structure. These classes are part of the file, rather than inferred again each time you use it.
We use Pandoc fenced divs such as ::: {#7-bc-1 .bar-chart}. The class identifies the element type; the ID identifies the first bar chart on PDF page 7. Page containers use :::: {.page n=7 label=5}, where n is the PDF page position and label, when available, is the printed page number.
.bar-chart.pie-chart.trend-chart.scatter-plot.materiality-matrix.map.infographic.diagram.chemical.other.table.image.page.toc.caption.footnote{page}-{type}-{counter}; the counter resets for each type on each page. Captions, footnotes and the table of contents use their class without a visual ID. Inline images use spans such as [company logo]{#7-i-2 .image}.$…$ or $$…$$. Repeated page headers and footers are removed by default, and columns are linearised in reading order.On request, we can retain headers and footers, provide tables as embedded HTML, omit visual tags, or add element geometry and enhancement metadata. Custom transformations are quoted separately.
Certified files pass checks for provenance, page sequence, typed elements, table evidence, duplication, character safety and Pandoc structure. OCR body text is not editorially rewritten; model-generated figure descriptions and difficult table transcriptions remain explicitly typed.
Complex layouts can still introduce errors, particularly in rotated tables, stylised charts and ambiguous reading order. If you find systematic issues, share examples so we can investigate and agree a correction.
One UTF-8 file per report, with document provenance, page containers and typed content. Standard Markdown is enriched with Pandoc attributes for precise structure.
The unmodified source reports accompany the conversion so your team can inspect layouts, figures and quoted passages.
Report and company details, identifiers and filenames connect the Markdown to the right source. See the metadata field guide.
Prices per report, in USD. Markdown is an add-on to Bulk Reports.
Bulk Reports price per report, in USD
| Tier | Order size | Markdown (PDF included)* | How to order | |
|---|---|---|---|---|
| 1 | Up to 1,000 | $0.60 | $1.20 | Download Portal |
| 2 | 1,001–10,000 | $0.45 | $0.90 | Request quote |
| 3 | More than 10,000 | $0.30 | $0.60 | Request quote |
* Markdown can’t be ordered in the Download Portal yet. Contact us and we’ll set it up.
Academic and non-profit: 30% off for eligible academic and non-profit organisations.
Build your Bulk Reports enquiry and select Markdown files in the Add-ons step.
Request quote