The FTW benchmark data includes all the chips and associated imagery to train field boundary models. This repo is starting as a demonstration of how to organize it, and will evolve to be alpha and beta releases of the 2.0 dataset. This catalog is a 'git-backed' Portolan catalog, backed by https://github.com/fieldsoftheworld/benchmark-data-catalog
Every claim in this file is quoted from a source or measured from the data.
Training and evaluation data for field boundary delineation models, one
collection per country or region. Each collection holds chip items on the FTW
grid (a fixed, MGRS-derived grid, so a cell id means the same square of ground
in every collection). Each chip item carries label masks — instance, semantic
2-class, semantic 3-class, and the DECODE boundary and distance layers — and,
where the collection has imagery, two Sentinel-2 scenes (planting and harvest)
as child items alongside it. Those scenes are either stored (a four-band
GeoTIFF clipped to the chip, published with it) or linked (the whole scene
COG on the source STAC API, for the reader to window); the coverage table below
says which, per collection, and some collections carry no imagery at all. Every
chip carries a ftw:split assignment of train, val or test, made in
spatial blocks.
Collection-level assets carry the field polygons the masks were cut from, their
boundary lines, the chips table with split assignments and field coverage, and
items.parquet, a stac-geoparquet mirror of every item.
The field boundaries come from the harmonized field boundary catalog; each collection's Source link below names the exact source collection.
Read items.parquet first. It lists every item — chips and season scenes —
with its bbox, split and asset hrefs, so a loader can select chips without
walking item JSON. It is queryable in place over HTTPS:
Each collection's own agent guide
(one per collection, at https://source.coop/ftw/benchmark-data/<id>/AGENTS.md)
documents its columns, its data quality notes and worked queries with their
results. The STAC entry point is
catalog.json; the
same tree is browsable in the
Portolan data browser.
Links in these documents are absolute: source.coop URLs are pages for people,
data.source.coop URLs are the bytes.
3 collections · 17,120 chips · 1,553,660 field polygons
stored means a four-band GeoTIFF clipped to the chip and published with it; linked means the chip's season item points at the whole Sentinel-2 scene on the source STAC API, for a reader to window. A chip with no scene still carries its masks and its split — pair it with imagery of your own, on the footprint in items.parquet.
Every number above is measured, not declared: chips and splits from each
collection's <id>_chips.parquet, field polygons from the <id>_fields
table the masks were cut from, imagery from the season child items on disk,
and crop labels from the ftw:hcat_dominant_code property in items.parquet.
Do not assume a collection carries imagery or crop labels because its
neighbours do — check the coverage table, or the collection's own assets.
This edition is 2.0.0-alpha.1, and while it is in alpha collections are
updated in place — the updated stamp in a collection's collection.json is
when it was last rebuilt. Every collection is rebuildable from its recipe in
the source repository
with ftw-dataset-tools.
Version 1.0 of the benchmark, superseded by this one, remains at
kerner-lab/fields-of-the-world;
the paper is Fields of The World (2024).
Each collection carries the license of the field boundary data it was cut from (the License column above). Derived masks and imagery chips are published under the same terms as their source. Sentinel-2 imagery is Copernicus data, free and open.
| 5,103 |
| 4,078/512/513 |
| 2,857 stored |
| License |
| harmonized/si |
| browse |