Agent guide: harvesting datasets from catalogs
List datasets inside registered catalogs via their public APIs. Human narrative: harvest.md. Per type: scientific IRs, domain scientific, opendata, geoportals, indicators, metadata, other. Shared protocols: harvest-protocols.md. Incremental: harvest-incremental.md. Identifiers: harvest-identifiers.md. EO: harvest-earthdata.md. Biodiversity: harvest-biodiversity.md. Viewers: harvest-viewers.md. Output: harvest-output.md.
This is not catalog discovery (discover.md) and not registry query (query.md).
Goal
For catalogs the user named (or a scoped DuckDB selection), produce dataset identifiers with:
- catalog
uid/id/link software.id- native dataset id and/or DOI/handle (harvest-identifiers.md)
- the type filter you applied
- skip counts (publications, files, showcases)
Do not write dataset YAML into this repository. Do not invent uid for datasets.
Before calling APIs
- Read llms.txt if you have not already.
- Resolve catalogs from exports (
datasets.duckdb/full.parquet). Useendpoints[]when present. - Confirm
software.id. Ifcustom, do not guess a CKAN/DSpace filter — inspect the live API once or stop. Do not filter exports on asoftware.idthat is missing fromdata/software/.
SELECT id, uid, name, link,
software.id AS software_id,
endpoints
FROM catalogs
WHERE id = 'examplegov'
OR lower(link) LIKE '%example.gov%';
Order of work
- Identify (OAI Identify, CKAN
status_show, Dataverseinfo/version, DSpace/server/api). - List type vocabularies (OAI
ListSets, search facets, CKANfqtypes). - Apply the platform dataset filter from the harvest guides. Do not page an unfiltered IR search.
- Paginate with the documented cursor (
start,page,resumptionToken). Small page size. Incremental later: harvest-incremental.md. - Drop publications, theses, files-under-datasets, showcases, harvest sources (keep vs drop).
- Emit one JSON record per kept dataset plus skip counts (harvest-output.md).
Platform shortcuts
Open the harvest heading from software-index.md. Do not invent filters for custom.
If software.id is | Dataset filter (then follow the harvest page) |
|---|---|
dataverse | /api/search?q=*&type=dataset |
dspace / dspacecris | f.entityType=Dataset or OAI ListSets |
invenio / inveniordm | /api/records?q=metadata.resource_type.type:dataset |
ckan / dkan | package_search (packages, not resources) |
ekan | /jsonapi/dataset/dataset or /data.json |
opendatasoft | /api/explore/v2.1/catalog/datasets |
socrata | /api/catalog/v1?only=datasets |
geonetwork | CSW GetRecords; hierarchyLevel dataset/series |
stacserver | /collections (items only if that is the grain) |
arcgisserver | /arcgis/rest/services?f=pjson — not GPServer |
pxweb | /api/v1/ tables (type: t), not folders |
pxstat | PxStat.Data.Cube_API.ReadCollection matrices, not widgets |
fenix | FAOSTAT groupsanddomains — not observation cubes |
tabnet | .def tables — not CGI sessions |
sparkmap | Public hub layers — not CARES HQ as a second copy |
g3wsuite | Published project WMS — not /admin |
sentinelhub | STAC collections — not Process jobs or EO Browser tiles |
geusmap | One harvest per mapname WMS/WFS layers |
resourcecontracts | /contract/resources — not per-clause chrome |
gxopendata | Tenant dataset list — not apply-gateway flows |
converis | Datasets only — not publications or persons |
labkey | Studies / published folders — not assay runs |
synapse | Projects and dataset entities — not every file |
xnat | Projects — not imaging sessions |
omero | Projects/screens — not images |
kadi4mat | Records and collections — not file blobs |
edal | DOI datasets — not a single landing as a seed |
nomad | Published entries/uploads — not calculation files |
intermine | Experiments/datasets — not gene pages |
gringlobal | Accession catalog exports — not each accession HTML page |
plutof | Datasets / DOI records — not occurrences or UNITE |
jgi | Genome/transcriptome projects — not gene pages or IMG/GOLD |
cbioportal | Studies — not mutation/CNA rows |
esasciencearchive | TAP tables — not FITS files |
odweb | /odweb/ datasets — not the parent CMS |
opengov | Named reports on {org}.opengov.com, not Highcharts series |
imfnsdp | SDMX categories/series on the country page — not DSBB |
archipelago | Solr /search digital objects — not Drupal nodes |
redivis | /api/v1/organizations/{org}/datasets — not tables or workflows |
custom | harvest-other.md decision tree |
If the filter returns zero hits, inspect one unfiltered sample and ListSets / facets before concluding the catalog has no data. Local labels include Forschungsdaten, Research Data, and numeric WEKO3 item types.
Accept / reject (dataset records)
Accept when the API object is a dataset (or data collection) with a stable id.
Reject: article, thesis, poster, presentation, person, project, org unit, CKAN resource row, Dataverse type=file, showcase, harvest source, login-only metadata, WMS tiles, ArcGIS GPServer, STAC items when collections are the grain, PxWeb folders, SDMX codelists-as-datasets, aggregator duplicates of source portals, IPUMS extracts, DHIS2 analytics cells, OpenAIRE publications.
Do not
- Walk
data/entities/**/*.yamlto find catalogs - Add harvested datasets as registry YAML
- Bypass
401/403or guess API keys (Pure/ws/api/often needs a key — use public OAI/sitemap instead) - Unscoped crawls or HTML scrapers for the whole web
- Treat Idra/federations as a substitute for harvesting member catalogs unless the user asked for the federation view
Related
- harvest.md
- harvest-scientific.md
- harvest-scientific-domain.md
- harvest-opendata.md
- harvest-geoportals.md
- harvest-indicators.md
- harvest-metadata.md
- harvest-other.md
- harvest-protocols.md
- harvest-incremental.md
- harvest-earthdata.md
- harvest-biodiversity.md
- harvest-viewers.md
- harvest-identifiers.md
- harvest-output.md
- software-index.md
- apidetect.md
- query.md
- discover.md