The entire Sentinel 2 archive translated into stac-geoparquet enabling scalable bulk access using any tool that understands Parquet. Syncs daily with Earth Search API and uses its COG's as the assets.
Tile-month aggregates over the sentinel-2-c1-l2a item index, which is
Sentinel-2 Collection 1, generated by
tools/s2_stats.py --collection sentinel-2-c1-l2a and joined by mgrs_tile.
The collection publishes three parquet products and one tileset, with the same
columns the stats collection builds from
sentinel-2-l2a. Only the tile column grouped on differs, because Collection 1
supplies it as the index's _tile.
The products:
mgrs-monthly.parquet -- one row per MGRS tile per month: scene
count, minimum and median eo:cloud_cover, the id and UTC date of that
tile-month's least-cloudy scene, and the mean and maximum percent of the
tile its scenes fill (mean_cover/max_cover, from
100 - s2:nodata_pixel_percentage; NULL when no scene carried it). Every
percent is an integer 0-100 (round()). Sorted by
(mgrs_tile, year, month) in 50k-row groups, so a filter on one tile is
a range read of one row group plus a small footer, not the whole file.
This is the cheap first stop for "which month has a cloud-free scene here".months/YYYY-MM.parquet -- one file per month present in the table
(not listed as assets; the pattern is the contract), holding that
month's rows with the paint columns only: mgrs_tile, scene_count, min_cloud_cover, median_cloud_cover, mean_cover, max_cover. Sorted by
mgrs_tile, one row group. The explorer app fetches one of these per
month it paints. The newest month that has a slice is the
max(year, month) row of timeline.parquet; a month between the oldest
and newest with no rows has no file (a 404 means "no tile-months").timeline.parquet -- one row per (year, month) over all tiles:
tile_count, summed scene_count, and the minimum min_cloud_cover. A
few KB; read it whole for the global timeline and the month span.mgrs.pmtiles -- one polygon per MGRS tile, on the vector layer
mgrs with an mgrs_tile attribute, generated with gpio pmtiles create
(tippecanoe underneath). The polygon is the envelope of that tile's scene
footprints, not the true MGRS grid cell, and it never changes on a daily
stat refresh: only the parquet products are recomputed, so the app's map
layer and its statistics update independently.A Sentinel-2 scene whose bbox spans more than 20 degrees of longitude has
wrapped the antimeridian and is reported as [-180, ..., 180, ...] --
averaging or enveloping that bbox draws a false polygon across the whole
globe. Those scenes are dropped from the footprint aggregation that builds
mgrs.pmtiles, but they are still counted normally in
mgrs-monthly.parquet: the exclusion is a geometry-rendering fix, not a
data-quality filter.
The speed is layout, not infrastructure. Each product is shaped for exactly one access pattern:
mgrs-monthly.parquet
is sorted by mgrs_tile, so one tile's rows sit together in one or two
row groups. A reader checks the footer's per-group tile ranges, fetches
only the matching groups with HTTP range requests, and moves tens of
kilobytes instead of the ~21 MB file.months/YYYY-MM.parquet
slice contains one month's rows with only the paint columns, ~100-150 KB.
A client fetches it whole in one request; there is nothing to prune.The scene explorer reads all three this way with hyparquet, a small pure-JS parquet reader, in the page: whole-file fetches for the small products, footer-pruned range reads for the big one. No server or query engine sits between the browser and the bucket. The item index uses the same idea at larger scale: sort by the query key, and the footer statistics route each read to a few small row groups.
One tile's history (a range read of one or two row groups):
Every tile for one month (a ~120 KB file):
Which months exist, and the newest:
The table covers every
month the sentinel-2-c1-l2a index holds, from 2015-10 on; the newest
month is the max(year, month) row of timeline.parquet. The daily
refresh recomputes every year its created lookback touches and rewrites
those months' slices and the timeline; tools/make_stats_collection.py --collection sentinel-2-c1-l2a
measures this collection's temporal extent, row count and updated from
the timeline at each publish, so the collection never has to be edited by
hand to follow the table. Because ESA's Collection 1 reprocessing is still
filling old years in (with recent created timestamps, which the index's
refresh looks back on), a month's scene_count can rise long after the
month ended; expect old months to change, not only the current year.
Aggregated from the sentinel-2-c1-l2a
item index, itself mirrored from the
Earth Search STAC API. Contains
modified Copernicus Sentinel data, under the
Copernicus Sentinel Data Terms and Conditions
(SPDX CC-BY-SA-3.0-IGO).