The entire Sentinel 2 archive translated into stac-geoparquet enabling scalable bulk access using any tool that understands Parquet. Syncs daily with Earth Search API and uses its COG's as the assets.
Guidance for AI agents and automated clients querying this collection.
One rule governs each edit to this file. A claim here is either quoted from a source or measured from the data. If you cannot point at where a fact came from, it does not belong in this file. An agent acting on an invented column name or an invented join key produces a confident wrong answer, and nothing downstream catches it.
One row per Sentinel-2 L2A scene in the AWS Earth Search item index,
republished as partitioned GeoParquet. The imagery stays in the public
sentinel-cogs bucket on AWS, and the assets column of each row already
includes its COG URLs.
A year's archive is one of three shapes, and the year tells you which.
2015-2018 are one items.parquet each. From 2019 the archive is split
by the UTM zone of s2:mgrs_tile (the leading one or two digits of the tile
id): 2019 and 2020 into four files, and from 2021 into eight, whose
boundaries sit inside the four (every quartile edge is also an octant edge):
The boundaries are fixed (both tiers balance the 2018 row distribution: the
four at 27/26/22/25%, the eight at 10–15% each; the eight exist because a
quarter of a 2021-sized year no longer built inside one CI job), so the part
that holds a tile is known from the tile id and the year alone: 31UFU is
zone 31, in z21-35.parquet for 2019–2020 and z21-31.parquet from 2021.
The current year also has live.parquet, the tail fetched daily since the
last consolidation. Its daily rebuild drops every id the year's archive parts
already hold, so no two parts of a year overlap. A glob over the year reads
each scene once, whichever shape the year has, and a client can always
dedupe on id anyway; it is safe and removes nothing:
hive_partitioning = true exposes year as an INTEGER column that is not
stored in the files. Filter on it first; it is the only filter that skips
whole files. The glob goes through the anonymous s3:// door because DuckDB
expands * only where it can list the store: over plain https:// it stops
with "Globs (*) for generic HTTP file are not supported", so an https://
read names one part (below). There is no zone= directory and no zone
column: the zone split is a file name inside the year, and you use it by
choosing the file.
Filter in this order. Each step removes more data than the next one can.
year IN (…) — partition pruning, skips whole files."s2:mgrs_tile" = '31UFU' — the spatial join key, and the cheapest spatial
filter there is. A tile id is stable for the life of the mission._month = 8 or a datetime range — rows are sorted by month first, so a
month filter prunes row groups inside the file."eo:cloud_cover" < 10 — the usual last cut.Column names with a colon are not identifiers. Quote them: "eo:cloud_cover",
not eo:cloud_cover.
Read only the part containing your tile's zone. The glob above opens every
part of the year and lets the s2:mgrs_tile filter discard the rest of it by
row-group statistics: seven footers read for nothing (measured 2026-09-21,
11.9 s against 11.5 s for the single-part form below, same eight rows).
Row-group size differs by vintage: parts published through 2023 carry
~100k-row groups (a tile lookup reads ~14 MB per group it touches); parts
from 2024 on carry ~6k-row groups (~0.85 MB per hit). Sort order also
differs by vintage: through 2025 rows are ordered (_month, _hilbert); from
2026 they are ordered (_month, s2:mgrs_tile, _hilbert), so one tile's scenes
for a month occupy one small row group and a tile lookup is one range
request. All vintages read identically, differing only in bytes and requests
per hit, and the older years are rebuilt when a larger machine allows. When
you know the tile, skip them entirely: pick the file by zone and year. Zone
31 in 2021 is in z21-31.parquet, so the same query touches one file:
To choose the file in code, parse the zone with ^\d{1,2} and take the range
that contains it from the year's tier: the four ranges for 2019–2020, the
eight from 2021; for a year before 2019 the file is items.parquet regardless
of zone. Add live.parquet for the current year. A bounding box spans the
zone ranges its longitudes fall in (each UTM zone is six degrees wide), so a
regional bbox query names one or two parts; a global query globs.
The per-year item (year=YYYY/YYYY.json) lists every part of that year as its
own asset, with its own row count and time range, if you would rather discover
than assume.
Use a ST_Intersects filter on geometry when you have a real polygon and no
tile id. It works, and it is slower than the tile filter, because it has to
decode geometries the tile filter never reads.
datetime is TIMESTAMP WITH TIME ZONE and every value is UTC. DuckDB renders
it in the session time zone, so run SET TimeZone='UTC' before you compare a
rendered string to a date, or a scene acquired at 23:50 UTC reports the next
day west of Greenwich and the previous day east of it.
The published column list is tools/s2_schema.py in the source repository, and
the collection's table:columns is generated from it. Both carry a description
per column; that is the authority, and this section covers what a description
cannot say on its own.
Spatial. geometry is the scene footprint in CRS84. bbox is the same
footprint as [w, s, e, n]. s2:mgrs_tile is the MGRS tile id, such as
31UFU, and it is the join key. Scenes over one place share it across all
years.
Time. datetime is the acquisition instant, UTC. s2:generation_time is
when ESA processed the product, not when the satellite looked.
The added columns are query helpers rather than STAC. They exist so that readers can prune:
_month — month(datetime), 1 to 12. The first sort key._hilbert — ST_Hilbert(geometry, world bounds). The last sort key.Each part file is written sorted by month first, so a month filter prunes row
groups inside a year. Hilbert position sorts last, which narrows each row
group's bounding box so a spatial filter prunes as well. Parts published
through 2025 are (_month, _hilbert), and parts from 2026 are
(_month, s2:mgrs_tile, _hilbert), which puts one tile's month in one row
group (see Row-group size above).
Earth Search publishes neither of these two columns. Do not pass them on as
STAC properties, and read _hilbert only as a position on a space-filling
curve, which orders rows and measures nothing.
Everything else is the upstream STAC property under its upstream name,
unchanged. Two of them are reconstructed when Earth Search omits them, and
never invented: sat:relative_orbit is parsed from the _R(\d{3})_ group of
s2:product_uri, and s2:mean_solar_zenith is computed as
90 - view:sun_elevation.
s2:granule_id and sat:orbit_state are NULL on newer items. Earth Search
stopped publishing those two properties, and this mirror leaves an absent value
absent rather than inventing one. A query that requires either silently drops
every recent scene. Use s2:product_uri in place of s2:granule_id, and
derive the orbit direction from sat:relative_orbit if you truly need it.
Several s2:*_percentage columns are NULL on some items for the same reason:
the upstream item did not carry them.
The record starts in November 2016, when Earth Search produced its first L2A Cloud-Optimized GeoTIFFs. No items exist for 2015 and most of 2016. The years 2017 and 2018 are partial, and the record is complete from about December 2018.
That is what Earth Search and the sentinel-cogs bucket serve, not a gap
introduced here. Do not report "no scenes in early 2016" as an observation
about Sentinel-2 — the mission was acquiring; this index does not carry those
items.
Check the per-year items (year=YYYY/YYYY.json) for each year's measured row
count and time range before you conclude anything about a period.
assets columnassets is a VARCHAR holding the upstream STAC assets object verbatim, as a
compact JSON string. It is a string on purpose: deeply nested structs make a
Parquet file hard for some readers to open, and a string keeps every reader's
schema flat. You parse it in exchange.
Keys present on a current L2A scene:
Each asset object carries href, type, title, roles, and for the raster
assets gsd, eo:bands, raster:bands, proj:shape and proj:transform.
The proj:* values are per scene and correct only for that scene.
The collection's item_assets block mirrors the stable part of this — the band
metadata per key — for STAC clients that expect it at collection level. It is
documentation. The per-item assets value is the authority, it is the only
place the hrefs exist, and it carries keys (cloud, snow,
product_metadata) that the upstream item_assets template omits.
Do not build an asset URL from a template. Read the href.
A scene can be fetched more than once: the daily refresh re-reads a five-day
window, and a reprocessed product keeps its id. Rows are deduped by id,
keeping the highest s2:generation_time (NULLS LAST), when each year is
built. So id is unique within a part. The parts of a year do not overlap:
a scene has one tile, and a tile one zone, and the daily live.parquet
rebuild drops every id the year's archive parts hold. The one exception is
a reprocessed product, which keeps its id with a newer s2:generation_time:
live keeps it beside the archive's older copy until the next consolidation
replaces that copy, so dedupe on id keeping the highest
s2:generation_time when you glob the current year. Otherwise id is unique
within a year. Across the whole table, treat id as unique and report it if
you ever find otherwise.
This collection does not filter, reclassify or interpolate anything Earth
Search publishes. It adds _month, _hilbert and the flat thumbnail_url,
and no other column. No STAC API stands in front of it, so query the Parquet
directly over HTTP range requests, without a key.
Structural links resolve relative to the object that carries them. Objects here
carry no self link, so a client tracks its own location.
z53-60.parquet| 53–60 |