Icechunk for GOBAI-O2 https://www.ncei.noaa.gov/access/metadata/landing-page/bin/iso?id=gov.noaa.nodc:0259304
🌐 View data in browser · 💻 Data access (code) · 📦 Data access (NCEI)
An Icechunk store of the complete GOBAI-O2 v2.3 monthly dataset: global gridded ocean-interior dissolved oxygen, its uncertainty, temperature and salinity, 2004-01 through 2024-12, on a 1° × 1° × 58-level pressure grid.
https://data.source.coop/fish-pace/gobai-o2/monthly
Unlike the other stores in this repository, this one is materialized: it holds real Zarr v3 chunks, not byte-range references. Nothing is fetched from NCEI at read time and the store is self-contained. The source is a single 12 GB contiguous, uncompressed NetCDF file, which offers no chunk boundaries worth referencing; rechunked and Zstd-compressed it is 5.59 GB in 12,858 objects.
The dataset is complete and static — GOBAI-O2 v2.3 is a finished archive version, not a growing feed — so this store needs no update pipeline. A later GOBAI-O2 version would be a new store.
No install, no account, no download — the viewer streams chunks straight from the store:
The viewer is gridlook, a WebGL globe for
cloud-hosted Zarr and Icechunk stores, published alongside the data at
gobai-o2/viewer/. The store to open is in the URL fragment after #, so any
store can be swapped into the same link — nothing about the viewer is specific to this
dataset. Give it a moment on first load: it fetches the store's metadata before drawing.
It is a browser rendering a multi-gigabyte store over the network, so treat it as a look, not an analysis. For anything quantitative use the code path below.
Needs icechunk >= 2.1 and xarray. No credentials, no account, and — because the chunks
are materialized — no virtual-chunk authorization step.
icechunk 1.x will not work.
icechunk.http_storagedoes not exist in icechunk 1.x (it arrived in 2.0). Every icechunk 2.x release needs Python 3.12 or newer, so on an older Pythonpip install icechunkquietly installs 1.1.x, and the code below fails withAttributeError: module 'icechunk' has no attribute 'http_storage'. Check what you have:python -c "import icechunk; print(icechunk.__version__)".
The snapshot that built the store is also tagged, so readonly_session(tag="v2.3") pins
you to it regardless of later commits.
Longitudes run 20.5 → 379.5, not 0 → 360 or −180 → 180: the grid starts at 20.5°E and
continues past the dateline into a second lap. This comes from the source file and is kept
as is. Selecting the Pacific near 200 is therefore sel(lon=200), and a plot of a full
map is centred on the Atlantic. To get a conventional axis:
Land and below-bottom cells are NaN (about 39 % of the surface layer). There is no
_FillValue or missing_value to unmask — the source declares none, and float NaN
survives the round trip.
GOBAI-O2 (Gridded Ocean Biogeochemistry from Artificial Intelligence — Oxygen) reconstructs ocean interior dissolved oxygen by training machine-learning models on shipboard bottle and profiling-float oxygen observations, then applying them to temperature and salinity fields from the Roemmich and Gilson (2009) Argo climatology. The method is described in Sharp et al. (2023).
All four are float32 on (time, pres, lat, lon). temp and sal are the Roemmich and
Gilson (2009) fields the oxygen reconstruction was driven with, carried along for
convenience — they are not an independent product.
Chunks are (time: 14, pres: 2, lat: 73, lon: 120) — 0.98 MB of float32 each, Zstd
compressed. That is a deliberate compromise between the two access patterns, neither of
which it favours: a global map at one time and pressure touches 6 chunks, a 252-month time
series at one point and depth touches 18, and a full-depth profile at one time touches 29.
Measured anonymously over the public endpoint, those three reads take about 0.4 s, 0.8 s
and 1.3 s.
The viewer is published by publish_viewer.py in the
GitHub repository
(python publish_viewer.py --product gobai-o2 --build ~/gridlook); it is a plain static
build of gridlook, and holds no data of its own.
The notebook that built it sits next to this README:
gobai-o2-monthly-icechunk-sc.ipynb (also in the GitHub repository
https://github.com/ocean-icechunks/icechunks, directory gobai-o2-monthly/). It downloads or
streams the NCEI file, adds CF and ACDD metadata, rechunks, writes with
Dataset.to_zarr, and validates the published store against the source.
It runs end to end without credentials and without writing anything: RUN_WRITE
defaults to False, so it opens the source, builds the same metadata and encoding, skips
the write, and validates. Set RUN_WRITE = True (with Source Cooperative write
credentials) to rebuild. requirements.txt beside it lists the package floors; Python 3.12
or newer is required, because icechunk 2.x is published requires_python = ">=3.12".
The source file is one 12 GB NetCDF at
https://www.ncei.noaa.gov/data/oceans/archive/arc0207/0259304/5.5/data/0-data/GOBAI-O2-v2.3.nc
(accession 0259304, archive version 5.5, 12,207,313,813 bytes). NCEI's landing-page
download button currently redirects in a loop, but that archive path serves the file
directly and honours HTTP range requests, so the notebook can read it without downloading
it — engine="netcdf4" plus a #mode=bytes URL suffix. Downloading it first is still much
faster for a full rebuild, because the source is contiguous and uncompressed and the write
reads it in a strided pattern.
Metadata only. No data value was altered. The source file carries per-variable
descriptions but no global attributes at all, and free-text units ("micromoles per kilogram", "degrees Celcius", "N/A" for salinity). The build added Conventions,
title, summary, source, history, creator and citation fields, the DOI, the license
and product_version; CF standard_name, units and axis attributes on the coordinates;
CF standard names and udunits-style units on the variables; and ancillary_variables
linking uncer to oxy.
Both are in the repository's ancestry(), and every code block on this page was executed
against the live store before it was published.
Code. The notebook and helpers are released under Apache-2.0 and are free to use, copy, adapt and redistribute, commercially or not — no attribution required.
Data. The data is not ours. GOBAI-O2 is released by its authors under CC0 1.0, which waives all copyright and imposes no legal obligation to cite. Scholarly practice still asks that you do:
Sharp, J. D., Fassbender, A. J., Carter, B. R., Johnson, G. C., Schultz, C., & Dunne, J. P. (2022). GOBAI-O2: A Global Gridded Monthly Dataset of Ocean Interior Dissolved Oxygen Concentrations Based on Shipboard and Autonomous Observations (NCEI Accession 0259304). NOAA National Centers for Environmental Information. https://doi.org/10.25921/z72m-yz67
and cite the method paper:
Sharp, J. D., Fassbender, A. J., Carter, B. R., Johnson, G. C., Schultz, C., & Dunne, J. P. (2023). GOBAI-O2: temporally and spatially resolved fields of ocean interior dissolved oxygen over nearly 2 decades. Earth System Science Data, 15, 4481–4518. https://doi.org/10.5194/essd-15-4481-2023