This is a benchmark dataset of government PDFs and corresponding OCR.
142 billion tokens of text extracted from 9 million United States government PDFs.
GovText is a public-domain dataset of PDFs harvested from the End of Term Web Archive (EOT), released together with two independent text renderings per document, the embedded text and OCR text produced by olmOCR, plus per-page OCR metadata and full web provenance.
Token counts use the o200k tokenizer and are computed over the pages on which each extractor recovered text.
Everything lives under a single public S3 prefix and can be read with the AWS CLI or any
S3-compatible client using --no-sign-request.
An equivalent HTTPS endpoint is available for tools that prefer plain URLs — useful for
DuckDB, curl, and range reads:
https://s3.us-west-2.amazonaws.com/us-west-2.opendata.source.coop/govscape/eota-ocr/...
The first four are per document, addressed by pdf_id. The parquet file is a single
dataset-level table covering all documents.
pdf_id is the 32-character base32-encoded SHA-1 payload digest recorded in the EOT CDX index, for example:
2222QNYHGDQLSRDX3UKLN6TOTX5AQU6E
This is the join key across every artifact, and it matches the digest column of
complete_cdx.parquet. Because it is a content digest, identical PDFs served from different URLs collapse to one pdf_id. This is how the corpus is deduplicated.
extracted_text/, ocr_text/)Both are gzipped tarballs holding one UTF-8 .txt file per page, named
<pdf_id>_<page-number>.txt with page numbers starting at 1.
A page absent from an archive means that extractor recovered no text for it.
ocr_metadata/)One JSON object per document, in Dolma record format:
metadata:
attributes — each is an array with one entry per page, in page order:
pdf_page_numbers lets you slice text back into pages without touching the tarballs — the character ranges are half-open and index directly into text.
cdx/complete_cdx.parquet)One row per capture: 34,257,302 rows covering every time the crawler saw one of these PDFs, across every EOT crawl.
filename, offset, and length together are enough to range-read the original record straight out of the EOT WARCs.
Captures by crawl:
The parquet covers all five crawls, but GovText itself contains only documents that appear
in the 2020 crawl. Rows from other crawls are included so you can trace a document's
history backward and forward in time. Filter on filename LIKE '%EOT-2020%' to restrict to
the crawl the corpus was built from.
The End of Term Web Archive is a multi-institutional effort that has captured the web presence of the U.S. federal government at the end of every presidential administration since 2008. GovText is derived from its 2020 crawl.
PDFs were identified and extracted as follows:
application/pdf or whose
URL carries a .pdf extension.Each surviving PDF was then rendered twice:
pypdfium2 (v4.30) reads the text objects already encoded in a page's content stream,
and returns nothing where no text layer exists.olmOCR rasterizes each page and transcribes it with a fine-tuned vision-language
model, producing linearized text.attributes.is_table and attributes.is_diagram are useful filters for identifying pages where this is likeliest.Released under CC0 1.0 Universal. The underlying End of Term Web Archive materials are U.S. government works in the public domain.
Please reach out with questions, and please let us know if you make use of this data set! Akshay Mehta — amehta48@bu.edu