Skip to main content

Agent guide: harvesting datasets from catalogs

List datasets inside registered catalogs via their public APIs. Human narrative: harvest.md. Per type: scientific IRs, domain scientific, opendata, geoportals, indicators, metadata, other. Shared protocols: harvest-protocols.md. Incremental: harvest-incremental.md. Identifiers: harvest-identifiers.md. EO: harvest-earthdata.md. Biodiversity: harvest-biodiversity.md. Viewers: harvest-viewers.md. Output: harvest-output.md.

This is not catalog discovery (discover.md) and not registry query (query.md).

Goal​

For catalogs the user named (or a scoped DuckDB selection), produce dataset identifiers with:

  • catalog uid / id / link
  • software.id
  • native dataset id and/or DOI/handle (harvest-identifiers.md)
  • the type filter you applied
  • skip counts (publications, files, showcases)

Do not write dataset YAML into this repository. Do not invent uid for datasets.

Before calling APIs​

  1. Read llms.txt if you have not already.
  2. Resolve catalogs from exports (datasets.duckdb / full.parquet). Use endpoints[] when present.
  3. Confirm software.id. If custom, do not guess a CKAN/DSpace filter — inspect the live API once or stop. Do not filter exports on a software.id that is missing from data/software/.
SELECT id, uid, name, link,
software.id AS software_id,
endpoints
FROM catalogs
WHERE id = 'examplegov'
OR lower(link) LIKE '%example.gov%';

Order of work​

  1. Identify (OAI Identify, CKAN status_show, Dataverse info/version, DSpace /server/api).
  2. List type vocabularies (OAI ListSets, search facets, CKAN fq types).
  3. Apply the platform dataset filter from the harvest guides. Do not page an unfiltered IR search.
  4. Paginate with the documented cursor (start, page, resumptionToken). Small page size. Incremental later: harvest-incremental.md.
  5. Drop publications, theses, files-under-datasets, showcases, harvest sources (keep vs drop).
  6. Emit one JSON record per kept dataset plus skip counts (harvest-output.md).

Platform shortcuts​

Open the harvest heading from software-index.md. Do not invent filters for custom.

If software.id isDataset filter (then follow the harvest page)
dataverse/api/search?q=*&type=dataset
dspace / dspacecrisf.entityType=Dataset or OAI ListSets
invenio / inveniordm/api/records?q=metadata.resource_type.type:dataset
ckan / dkanpackage_search (packages, not resources)
ekan/jsonapi/dataset/dataset or /data.json
opendatasoft/api/explore/v2.1/catalog/datasets
socrata/api/catalog/v1?only=datasets
geonetworkCSW GetRecords; hierarchyLevel dataset/series
stacserver/collections (items only if that is the grain)
arcgisserver/arcgis/rest/services?f=pjson — not GPServer
pxweb/api/v1/ tables (type: t), not folders
pxstatPxStat.Data.Cube_API.ReadCollection matrices, not widgets
fenixFAOSTAT groupsanddomains — not observation cubes
tabnet.def tables — not CGI sessions
sparkmapPublic hub layers — not CARES HQ as a second copy
g3wsuitePublished project WMS — not /admin
sentinelhubSTAC collections — not Process jobs or EO Browser tiles
geusmapOne harvest per mapname WMS/WFS layers
resourcecontracts/contract/resources — not per-clause chrome
gxopendataTenant dataset list — not apply-gateway flows
converisDatasets only — not publications or persons
labkeyStudies / published folders — not assay runs
synapseProjects and dataset entities — not every file
xnatProjects — not imaging sessions
omeroProjects/screens — not images
kadi4matRecords and collections — not file blobs
edalDOI datasets — not a single landing as a seed
nomadPublished entries/uploads — not calculation files
intermineExperiments/datasets — not gene pages
gringlobalAccession catalog exports — not each accession HTML page
plutofDatasets / DOI records — not occurrences or UNITE
jgiGenome/transcriptome projects — not gene pages or IMG/GOLD
cbioportalStudies — not mutation/CNA rows
esasciencearchiveTAP tables — not FITS files
odweb/odweb/ datasets — not the parent CMS
opengovNamed reports on {org}.opengov.com, not Highcharts series
imfnsdpSDMX categories/series on the country page — not DSBB
archipelagoSolr /search digital objects — not Drupal nodes
redivis/api/v1/organizations/{org}/datasets — not tables or workflows
customharvest-other.md decision tree

If the filter returns zero hits, inspect one unfiltered sample and ListSets / facets before concluding the catalog has no data. Local labels include Forschungsdaten, Research Data, and numeric WEKO3 item types.

Accept / reject (dataset records)​

Accept when the API object is a dataset (or data collection) with a stable id.

Reject: article, thesis, poster, presentation, person, project, org unit, CKAN resource row, Dataverse type=file, showcase, harvest source, login-only metadata, WMS tiles, ArcGIS GPServer, STAC items when collections are the grain, PxWeb folders, SDMX codelists-as-datasets, aggregator duplicates of source portals, IPUMS extracts, DHIS2 analytics cells, OpenAIRE publications.

Do not​

  • Walk data/entities/**/*.yaml to find catalogs
  • Add harvested datasets as registry YAML
  • Bypass 401/403 or guess API keys (Pure /ws/api/ often needs a key — use public OAI/sitemap instead)
  • Unscoped crawls or HTML scrapers for the whole web
  • Treat Idra/federations as a substitute for harvesting member catalogs unless the user asked for the federation view