This dataset contains GPS-located land-cover samples that can be used to train and validate AI models that generate detailed, accurate maps, with a focus on coffee and cocoa production systems. The data were collected across Colombia and Ghana through expert interpretation of high-resolution satellite imagery. Each sample is labelled and quality-controlled to represent the land-cover types. The classification scheme includes four main classes: coffee, cocoa, tree crops, and seasonal agriculture.
This dataset is part of the Sample Earth initiative, a global effort to build open, high-quality reference data for improving the accuracy and inclusiveness of land-cover maps.
This dataset contains GPS-located land-cover samples that can be used to train and validate AI models that generate detailed, accurate maps, with a focus on coffee and cocoa production systems. The data were collected across Colombia and Ghana through expert interpretation of high-resolution satellite imagery. Each sample is labelled and quality-controlled to represent the land-cover types. The classification scheme includes four main classes: coffee, cocoa, tree crops, and seasonal agriculture.
While the primary goal is to distinguish coffee and cocoa systems from other land uses, the dataset also supports broader applications such as agricultural monitoring, deforestation analysis, ecosystem-service mapping, land-use planning, and suitability modelling.
By providing transparent, well-validated training data, this dataset contributes to Sample Earth's broader objective: strengthening AI-based land monitoring tools and supporting global efforts, including the EU Deforestation Regulation (EUDR), to ensure sustainable, deforestation-free agricultural supply chains.
The dataset is designed to grow continuously, incorporating new commodities, timeframes, and countries over time.
This dataset was produced with the support of Google.

Both GeoPackages use the WGS 84 (EPSG:4326) coordinate reference system and are licensed under CC BY 4.0.
Each GeoPackage contains four tables:
land-cover-legendAttribute table containing the land cover classification legend.
sampling-areaPoint geometry layer representing the centre of each sampling area cell. Each cell defines a 600 m × 600 m area within which imagery interpretation was conducted.
polygon-collectionMultiPolygon geometry layer containing exhaustive land cover delineations within each sampling area. Around each selected sampling area centre, within the 600 m extent, all cocoa, coffee, tree crops (orchards/bamboo), and seasonal agriculture pixels were mapped as polygons. Therefore, any non-annotated pixel within a sampling area can be assigned to the remaining (other) land cover categories. In other words, the polygon collection represents a complete annotation of the target classes within each cell.
point-collectionPoint geometry layer containing GPS reference points interpreted from high-resolution satellite imagery, collected within a 600–1200 m square buffer around the sampling area centre for each land cover category. These points complement the polygon collection by providing point-based reference samples across the classified categories.
A QGIS project is provided to visualise the dataset with pre-configured styling and layer ordering. Open the .qgz project file in QGIS (version 3.x or later) to explore the data with the appropriate symbology and basemap.
If you'd like to collect additional points and have them validate cleanly against the legend, configure the attribute form of point-collection (or polygon-collection) so that land_cover_id is picked from a drop-down sourced from land-cover-legend.
point-collection → Properties → Attributes Form.land_cover_id.land-cover-legendland_cover_id (the integer stored in the point)name (the label shown to the surveyor)sample_uuid, cell_uuid: Text Edit, with a default expression of uuid() for sample_uuid and a Value Relation to sampling-area.cell_uuid for cell_uuid.shade_level: Range (Integer, min 0 max 3) or a Value Map if you have fixed categories.recently_planted: Checkbox.created, updated: Date/Time, with default expression now() and, for updated, Apply default value on update ticked.comment: Text Edit, multi-line.With this setup, opening the identify → form view on a new point gives surveyors a drop-down of legend names rather than raw integers, while the stored value remains the integer land_cover_id. This keeps the file consistent with the existing data and the schema documented above.
imagery_date) was used as the basis for visual interpretation.You can reuse the GeoPackages directly in your own QGIS project. Simply drag each .gpkg into the Browser or use Layer → Add Layer → Add Vector Layer…. Each file exposes its four tables as separate layers: land-cover-legend, sampling-area, polygon-collection, and point-collection.
QGIS join: this makes the legend's name and description appear as virtual fields on every polygon/point feature in the attribute table and identify tool.
land-cover-legend, sampling-area, polygon-collection, point-collection.polygon-collection → Properties → Joins.land-cover-legendland_cover_idland_cover_idlegend_) so columns appear as name/description instead of land-cover-legend_name.name and description.point-collection..qgz project file.The joins are saved in the QGIS project, not in the GeoPackage, so the .gpkg files remain unchanged and portable.
We are actively looking for collaborators to expand this dataset to new geographies and crop types. If you are a researcher, practitioner, or organisation working on land cover mapping, remote sensing, or agricultural monitoring and would like to contribute, whether by collecting new reference data, validating existing annotations, or integrating this dataset into your workflows, we would love to hear from you. Please reach out to us at T.Vantalon@cgiar.org.
A companion dataset covering Vietnam and Ghana with an extended 10-class / 68-sub-class legend is available on the Harvard Dataverse:
Vantalon, T. et al. (2025). Sample Earth: Land Cover Reference Dataset. Harvard Dataverse. doi:10.7910/DVN/U7HWY1
If you use this dataset, please cite it as:
Vantalon, T. et al. (2025). Sample Earth: Tree Crop Reference Dataset. Alliance of Bioversity International and CIAT. doi:10.60480/mjsk-yg88
All data is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
This work was conducted by the Alliance of Bioversity International and CIAT and supported by Google.
| Datetime |
| Timestamp of record creation |
updated | Datetime | Timestamp of last update |
| Datetime |
| Timestamp of record creation |
updated | Datetime | Timestamp of last update |