The entire Sentinel 2 archive translated into stac-geoparquet enabling scalable bulk access using any tool that understands Parquet. Syncs daily with Earth Search API and uses its COG's as the assets.
Guidance for AI agents and automated clients querying this collection.
One rule governs each edit to this file. A claim here is either
quoted from a source or measured from the data. If you cannot point at where a
fact came from, it does not belong in this file. An agent acting on an invented
column name or an invented join key produces a confident wrong answer, and
nothing downstream catches it.
Every year from 2015 on is published (the collection.json counted
30,402,025 rows, 2015-10 to 2026-09, when read on 2026-09-22). This is the
uniform record: query it first, and use
sentinel-2-l2a when you need Earth
Search's original index instead.
One row per Sentinel-2 Collection 1 L2A scene in the AWS Earth Search item
index (collection sentinel-2-c1-l2a: ESA's reprocessing of the archive to
one processing baseline), republished as partitioned GeoParquet. The imagery
is not here: it is in the public e84-earth-search-sentinel-data bucket on
AWS, and every COG URL is already in the assets column of the row that
describes it.
A year is one items.parquet. There is no zone split and no zone column:
the archive of a year, whatever its size, is that one file. Any year may
also carry live files, one per month of it (live-01.parquet to
live-12.parquet): the tail the daily refresh has appended since the last
fold, each scene in the month it was acquired in, so a day of refresh
rewrites only the months it fetched. A fold every month or two merges them
into items.parquet and empties each one. A year refreshed before those
files existed may also carry live.parquet with zero rows: it held the
whole tail as one file, and this catalog never deletes a published file.
Glob the year (year=YYYY/*.parquet) rather than naming the parts, and you
do not have to know which months are there.
The files of a year do not overlap, with one exception: a scene ESA
reprocessed keeps its id with a newer s2:generation_time, and the live
part holds the newer copy beside the archive's older one until the next
fold. Dedupe on id keeping the highest s2:generation_time when you glob
a year; it is safe and removes nothing else. The QUALIFY below is that
dedupe; without it the count is high by the number of scenes reprocessed
since the last fold, which is a fine approximation for a coverage question
and the wrong answer for an inventory one:
hive_partitioning = true exposes year as an INTEGER column that is not
stored in the files. Filter on it first; it is the only filter that skips
whole files. The glob goes through the anonymous s3:// door because DuckDB
expands * only where it can list the store: over plain https:// it stops
with "Globs (*) for generic HTTP file are not supported", so an https://
read names the file (below).
Query pattern
Filter in this order. Each step removes more data than the next one can.
year IN (…) — partition pruning, skips whole files.
_tile = '31UFU' — the spatial join key, the first sort key, and the
cheapest spatial filter there is: a tile's year is one contiguous run,
so this filter prunes to the one or two row groups that hold it. A tile
id is stable for the life of the mission.
A datetime range — the second sort key, so it trims the tile's run;
on its own (no tile) it prunes nothing, because every month is in
nearly every row group. _month still works as a filter, with the same
caveat.
"eo:cloud_cover" < 10 — the usual last cut.
Column names with a colon are not identifiers. Quote them: "eo:cloud_cover",
not eo:cloud_cover. _tile, _month and _hilbert need no quotes.
Row groups. Each file is sorted (_tile, datetime), which is tile-major,
in uniform row groups at a target of 6,000 rows (6,144 as DuckDB writes them).
One tile's year is one contiguous run in acquisition order, so a
tile-and-window query admits the one or two groups covering that run, at one
range request each and a few hundred KB, whatever the window. This is the
layout the search-latency experiments behind issue #9 measured as fastest
for tile-window searches; the month-major order of the older
sentinel-2-l2a parts scatters a tile's year across its twelve month
sections and admits a group per month instead. A month filter alone does
not prune here (every month is in nearly every group), and a bbox filter
inside a tile has no Hilbert locality to lean on; both are correct, just
not cheap. Lead with the tile.
Use an ST_Intersects filter on geometry when you have a real polygon and no
tile id. It works, and it is slower than the tile filter, because it has to
decode geometries the tile filter never reads. proj:centroid (a struct of
lat, lon) is a cheap point stand-in when the footprint is more than you
need.
datetime is TIMESTAMP WITH TIME ZONE and every value is UTC. DuckDB renders
it in the session time zone, so run SET TimeZone='UTC' before you compare a
rendered string to a date, or a scene acquired at 23:50 UTC reports the next
day west of Greenwich and the previous day east of it.
Schema
The published column list is tools/s2c1_schema.py in the source repository,
and the collection's table:columns is generated from it, as is this table.
That is the authority; the notes after it cover what a description cannot
say on its own.
Column
Type
Description
thumbnail_url
string
Preview JPEG on the e84-earth-search-sentinel-data bucket.
type
string
Always 'Feature'.
stac_version
string
STAC version of the source item.
stac_extensions
list
Extension schema URIs of the source item.
id
string
Earth Search item id, such as S2B_T31UET_20260921T105030_L2A.
bbox
list
Item bounding box [w, s, e, n], CRS84.
links
STRUCT(href VARCHAR, rel VARCHAR, title VARCHAR, "type" VARCHAR)[]
Source item links (canonical et al.); paging links are stripped.
collection
string
Always 'sentinel-2-c1-l2a'.
datetime
timestamp[us, tz=UTC]
Acquisition datetime, UTC.
Spatial.geometry is the scene footprint in CRS84. bbox is the same
footprint as [w, s, e, n]. There is no s2:mgrs_tile in Collection 1
items: the tile arrives as grid:code (MGRS-31UET) and as the three
mgrs:* parts, and this mirror adds _tile, the bare id (31UET,
grid:code without its MGRS- prefix). _tile is THE join key: scenes
over one place share it across all years, and it is the same id the
sentinel-2-l2a collection calls s2:mgrs_tile and the stats-c1
collection calls mgrs_tile.
Time.datetime is the acquisition instant, UTC. s2:generation_time is
when ESA processed the product, not when the satellite looked. created
and updated are when Earth Search ingested and last touched the item;
created is what the daily refresh looks back on, because reprocessed old
scenes arrive with an old datetime and a new created.
The added columns are query helpers rather than STAC. They exist so that
readers can prune and join:
_tile — the MGRS tile id, above. The first sort key; datetime (an
upstream column) is the second, and together they are the whole order.
_month — month(datetime), 1 to 12. A filter convenience only; not a
sort key in this collection (it is the first sort key in
sentinel-2-l2a, so the column is kept here with the same meaning).
_hilbert — ST_Hilbert(geometry, world bounds). Kept for parity with
sentinel-2-l2a; not a sort key here.
Earth Search publishes none of these three columns. Do not pass them on as
STAC properties, and read _hilbert only as a position on a space-filling
curve, which orders rows and measures nothing.
Both nested upstream objects arrive as strings.processing:software is
the upstream name-to-version map as a compact JSON string, because its keys
change between baselines and a struct would need one schema per baseline.
assets is the JSON string described below. proj:centroid and links are the only nested columns.
Each remaining column is the upstream STAC property under its upstream
name, unchanged. This mirror reconstructs and derives no column from another
field.
Null columns
s2:dark_features_percentage is NULL on every scene processed at baseline
05.11 or later — ESA dropped the class from the scene classification, and
the column is present on baselines 05.00–05.10 only (measured on one item
per year across 2017–2026 when the schema was frozen). A query that requires
it silently drops every recently processed scene. s2:processing_baseline
says which case a row is.
A property the schema does not know (something Earth Search adds after the
freeze) is dropped from the row and counted by the fetch, never invented.
Coverage
Earth Search's sentinel-2-c1-l2a held 30,391,138 items when it was counted
on 2026-09-21, by year of acquisition: 2015 200 · 2016 222 · 2017 24,664 ·
2018 1,329,973 · 2019 3,166,919 · 2020 4,013,725 · 2021 4,055,717 · 2022
283,705 · 2023 4,249,728 · 2024 4,369,942 · 2025 5,065,361 · 2026 3,830,982
(to date). The earliest item is acquired 2015-10-22.
Those counts move. ESA's reprocessing is ongoing, so the thin years (2015–
2017, and 2022 at a fraction of its neighbours) are years it has not
reached or finished, and they grow with recent created timestamps. That
is the source's state, not a gap introduced here. Do not report "few scenes
in 2022" as an observation about Sentinel-2 — the mission was acquiring at
full rate; the index has not caught up. Check the per-year items
(year=YYYY/YYYY.json) for each year's measured row count and time range
before you conclude anything about a period, and expect a year's count to
be higher next month.
The assets column
assets is a VARCHAR holding the upstream STAC assets object verbatim, as a
compact JSON string. It is a string on purpose: deeply nested structs make a
Parquet file hard for some readers to open, and a string keeps every reader's
schema flat. You parse it in exchange.
Keys present on a Collection 1 scene (the 23 keys of the upstream
item_assets template, all present on the frozen sample item):
There are no JPEG 2000 assets in Collection 1, unlike sentinel-2-l2a.
Every href sits under one directory per scene:
https://e84-earth-search-sentinel-data.s3.us-west-2.amazonaws.com/sentinel-2-c1-l2a/<zone>/<band>/<sq>/<year>/<month>/<id>/
(zone, latitude band and grid square being the three parts of _tile, the
month unpadded), with the band files named B04.tif, TCI.tif, SCL.tif,
CLD_20m.tif, SNW_20m.tif and L2A_PVI.jpg. That directory also
holds <id>.json, the canonical STAC item (the row's links carry it as
rel:canonical).
Each asset object carries href, type, title, roles, and for the raster
assets gsd, eo:bands, raster:bands, proj:shape and proj:transform.
The proj:* values are per scene and correct only for that scene.
The collection's item_assets block mirrors the stable part of this — the band
metadata per key — for STAC clients that expect it at collection level. It is
documentation. The per-item assets value is the authority, and it is the only
place the hrefs exist.
Do not build an asset URL from a template. Read the href.
Dedupe
A scene can be fetched more than once: the daily refresh re-reads a lookback
window on created, and a reprocessed product keeps its id. Rows are
deduped by id when each year file is built and again at each fold. That
dedupe retains the highest s2:generation_time, with NULLS LAST. So id is
unique within a file. Between folds, a reprocessed scene can sit in a live part
with a newer s2:generation_time than the copy in items.parquet, so
dedupe on id keeping the highest s2:generation_time when you glob a
year (the QUALIFY in the glob snippet above). Across the whole table,
treat id as unique and report it if you ever
find otherwise.
Cadence
Three clocks, from the workflows in the repository's .github/workflows/
and the cluster scripts in tools/rails/:
Daily, 03:42 UTC (refresh-daily, its Collection 1 job): the last
five days of Earth Search by created, appended to the live part of
every (year, month) the slice touches -- in steady state one file --
with the ids the year's items.parquet already holds dropped. The year
items and the collection are restamped from the published files on the
same run, so table:row_count and the extents describe the bucket, not
the commit.
Daily, on the same run: the stats collection
(stats-c1) has every touched year recomputed
and spliced into mgrs-monthly.parquet.
Every month or two, and at year end, by a person on the RAILS
cluster (tools/rails/fold_live.sbatch): every live part of a year is
merged into its items.parquet and emptied. There is no monthly
consolidation on GitHub for this collection. Between folds a year's files
together are the year; the dedupe above makes the count exact.
The daily job runs only while the repository variable C1_LIVE_ENABLED is
set. With it unset this collection changes only when a person commits and
publishes it.
Scope of the mirror
This collection does not filter, reclassify or interpolate anything Earth
Search publishes. It adds _month, _hilbert, _tile and the flat
thumbnail_url, and no other column. No STAC API stands in front of it, so
query the Parquet directly over HTTP range requests, without a key.
Structural links resolve relative to the object that carries them. Objects here
carry no self link, so a client tracks its own location.