Skip to content

Turn reports into
structured text.

Ready for LLM pipelinesPage-level traceability

Overview

Sustainability reports are hard to work with: dense text, complex tables, charts and images, all locked inside PDFs. We convert them into text you can search, parse and feed to a model, and we keep track of which page everything came from. The original PDF comes along, so you can always check.

Page by page

Every page has a clear boundary and a stable PDF page number, even when the printed page label differs.

Typed content

Headings, tables, figures, captions and footnotes are labelled, so your code can treat each one differently.

Source included

Every order pairs the Markdown with the original PDF and a 50-field metadata record.

Explore a report

Choose a page to compare the source PDF with the Markdown produced by our pipeline.

HMC Capital · 2022 Sustainability Report
Previewing PDF pages 1–5 of 52

Original PDF
1 / 5
Original page 1 of HMC Capital 2022 Sustainability Report
Processed Markdownpage 1.md
:::: {.page n=1}

# 2022 Sustainability Report

Creating Healthy Communities

::: {#1-i-1 .image}
A group of diverse people standing in a circle outdoors, holding each other's shoulders, looking towards the horizon. The image is bathed in warm, golden sunlight, suggesting a community or team effort.
:::

::::

Process

Our conversion process is built specifically for sustainability reports. Every page goes through vision OCR. Every figure then gets a second look from a separate vision model, and every table gets its own correction pass.

1

Read every page

Our vision OCR process reads each page and identifies text, reading order, tables and visual regions across complex report layouts.

2

Reprocess the visuals

Each detected figure receives a separate vision-model pass for a useful description and type. Tables receive targeted transcription and correction, preserving their evidence and units where available.

3

Render and validate

We turn the enriched result into page-addressable Markdown, then check page sequence, typed elements, table retention and source binding. Files that don’t pass are held back for review instead of being delivered.

Good to know: we also use these Markdown files in our own data extraction jobs. Our extraction pipeline achieved 98.7% value accuracy in SRC-Bench v1: 226 of 229 reported values matched expert annotations across seven indicators in 50 reports.

Taxonomy

Every element is labelled so your code can select the content it needs: charts, tables, images or document structure. These classes are part of the file, rather than inferred again each time you use it.

We use Pandoc fenced divs such as ::: {#7-bc-1 .bar-chart}. The class identifies the element type; the ID identifies the first bar chart on PDF page 7. Page containers use :::: {.page n=7 label=5}, where n is the PDF page position and label, when available, is the printed page number.

Charts

.bar-chart
Bar, column and stacked-bar charts.
.pie-chart
Pie, donut and ring charts.
.trend-chart
Line charts and time series.
.scatter-plot
Scatter and bubble plots.

Other visualisations

.materiality-matrix
Materiality grids and stakeholder/impact plots.
.map
Geographic maps and regional breakdowns.
.infographic
Mixed visual and text compositions.
.diagram
Process flows, organisation charts and conceptual diagrams.
.chemical
Chemical structures and molecular diagrams.
.other
Visualisations outside the other categories.

Tables and images

.table
Data tables, with units and year ranges where available.
.image
Photographs, logos, icons and decorative graphics.

Document structure

.page
PDF page containers with a sequential page number.
.toc
The document’s table of contents.
.caption
Figure or image captions.
.footnote
Footnote text, including lists and multiple paragraphs.
File structure and formatting
  • Document provenance. YAML frontmatter identifies the report and conversion. Certified V3 files also carry the source OCR, converter and validator versions and hashes. The companion metadata includes the original PDF’s SHA-256.
  • Stable references. Visual IDs follow {page}-{type}-{counter}; the counter resets for each type on each page. Captions, footnotes and the table of contents use their class without a visual ID. Inline images use spans such as [company logo]{#7-i-2 .image}.
  • Tables. Standard Markdown pipe tables retain units and year ranges as attributes where available. Merged cells carry explicit row/column-span annotations, with blank placeholders for covered cells.
  • Text and equations. UTF-8 preserves original-language characters. Headings, lists and links retain their structure; equations use $…$ or $$…$$. Repeated page headers and footers are removed by default, and columns are linearised in reading order.
  • Tool compatibility. Pandoc can parse the typed structure into an abstract syntax tree (AST). Other Markdown tools may require an extension or conversion to preserve attributes.
Customisation and quality checks

On request, we can retain headers and footers, provide tables as embedded HTML, omit visual tags, or add element geometry and enhancement metadata. Custom transformations are quoted separately.

Certified files pass checks for provenance, page sequence, typed elements, table evidence, duplication, character safety and Pandoc structure. OCR body text is not editorially rewritten; model-generated figure descriptions and difficult table transcriptions remain explicitly typed.

Complex layouts can still introduce errors, particularly in rotated tables, stylised charts and ambiguous reading order. If you find systematic issues, share examples so we can investigate and agree a correction.

Delivery

Structured Markdown

One UTF-8 file per report, with document provenance, page containers and typed content. Standard Markdown is enriched with Pandoc attributes for precise structure.

Original PDFs

The unmodified source reports accompany the conversion so your team can inspect layouts, figures and quoted passages.

50-field metadata

Report and company details, identifiers and filenames connect the Markdown to the right source. See the metadata field guide.

Pricing

Prices per report, in USD. Markdown is an add-on to Bulk Reports.

Bulk Reports price per report, in USD

Tier
1
Order size
Up to 1,000
PDF
$0.60
Markdown (PDF included)*
$1.20
How to order
Download Portal
Tier
2
Order size
1,001–10,000
PDF
$0.45
Markdown (PDF included)*
$0.90
How to order
Request quote
Tier
3
Order size
More than 10,000
PDF
$0.30
Markdown (PDF included)*
$0.60
How to order
Request quote

* Markdown can’t be ordered in the Download Portal yet. Contact us and we’ll set it up.

Academic and non-profit: 30% off for eligible academic and non-profit organisations.

Getting started

Request a quote

Build your Bulk Reports enquiry and select Markdown files in the Add-ons step.

Request quote

FAQ

What are Markdown files?

Each report converted into plain text with its structure intact: headings, tables and labelled figures. Page markers show exactly which PDF page each part came from. They’re made for text analysis and AI workflows.

Why use Markdown instead of feeding PDFs directly to a model?

Depending on the tool, a PDF upload may use only extracted text and miss chart content, image-based text, table structure or footnotes. Our Markdown makes these elements explicit through vision OCR, typed figure descriptions, corrected tables and page references. It gives extraction jobs consistent, reusable input without parsing the PDF again each time. The original PDF stays available to check the source.

How are charts and tables processed?

Every page goes through vision OCR. Each figure then gets a second pass from a separate vision model that classifies and describes it, and tables get their own correction pass. Before delivery, every file is checked for structure and against its source PDF.

What is included with Markdown?

Always the original PDF and the 50-field metadata.

How are Markdown files priced?

Double the PDF price: $1.20 per report for up to 1,000 reports, $0.90 from 1,001 to 10,000, and $0.60 above 10,000, PDFs included. For now, you order Markdown through us rather than the Download Portal. See the full pricing table.