Skip to main content

Harvesting datasets from catalog APIs

This registry stores catalogs (portals, geoportals, repositories). It does not store the datasets inside those catalogs. To list or index datasets, harvest the catalog’s public API.

Two different jobs share the word harvest:

JobWhat you wantWhere to go
List catalogs in this registryCountry, type, software.id, endpoints[]query-examples.md, agents/query.md
List datasets inside a catalogRecords from the remote API, filtered to datasetsThis page, then the platform guides
Find catalogs not yet registeredNew portal URLsdiscovery.md

Coding agents: agents/harvest.md. Production harvesting for the Dateno stack lives in reaper — these pages are the human/API recipes, not a crawler in this repository.

Guides​

GuideUse when
Scientific repositoriesDSpace, Invenio, EPrints, Pure, Esploro, DLCM, easydb, Haplo, Omega-PSIR, Hyrax, Archipelago, LabKey, Synapse, XNAT, OMERO, Kadi4Mat, e!DAL, NOMAD, and other IRs that mix publications, theses, software, and datasets
Domain scientific repositoriesIPT, Biodiv, THREDDS, ERDDAP, FROST-Server, GeoNature, Specify, DaCHS, Breedbase, Tripal, VEuPathDB, MassBank, ioChem-BD, ESGF, InterMine, GRIN-Global, PlutoF, JGI, cBioPortal
Open data portalsCKAN, OpenDataSoft, Socrata, RUDI, GIS Open Data Portal, ResourceContracts, RDF Online Repository, Guangxi, ODWeb, OpenGov — packages vs resources vs contracts vs reports
GeoportalsGeoNetwork CSW, GeoNode, ArcGIS, STAC, OGC API, OneGeo Suite, PRODIGE, G3W-SUITE, CubeWerx, M.App Enterprise, mviewer, Isogeo, Geocortex, QGIS Server — layers vs services vs tiles
Indicators and microdataPxWeb / PxStat tables, DGBAS Web tables, SDMX dataflows, OpenSDG / Goal Tracker, IMF NSDP, NADA studies, DHIS2, TabNet, SIDRA, FENIX, SparkMap, eDatos, Cancer-Rates.info, HCI, Virtual LMI, Fingertips, KOSIS, e-Stat, UNdata, Comtrade Plus, Our World in Data, IPUMS
Metadata catalogsFAIR Data Point DCAT, Aristotle MDR, Fusion Registry, Metadata Browser, CBD Clearing-House, WMO OSCAR
Search, ML, API, marketplacesAggregators, OpenAIRE, OpenML, Hugging Face, CodaLab, API directories, marketplaces, custom
ProtocolsOAI-PMH, CSW, DCAT, STAC, SDMX, OGC, ArcGIS REST — grain that is shared across products
Incremental harvestsfrom=, metadata_modified, STAC datetime, checkpoints
Earth observationTHREDDS, ERDDAP, STAC collections, Open Data Cube, Copernicus, openEO, Sentinel Hub, ESGF, ESA Science Archive
Biodiversity and genomicsIPT, Symbiota, ALA, GBIF datasets, Ensembl species, Breedbase, Tripal, VEuPathDB, PlutoF, InterMine, JGI, cBioPortal
Map viewersQWC2, Masterportal, Lizmap, mviewer, WebEWID, GISPLAN, Wagmap, Trimble Locus / Louhi / Landfolio, Spatial Suite, Spectrum Spatial Analyst, Hajk, myCarta, KortInfo, IntraMaps, LocalMaps, GEUSMAP, GISApp, GeneGIS PAGIS, GisMaster, GeoPortale.cloud, LDP SIT, GFMaplet, HyG Mapgis, iObčina, iShare, Cadcorp, StatMap Earthlight — layers not tiles
Dataset identifiersNative id + catalog uid; DOI/handle; do not mint cdi######## for datasets
Harvest outputJSON record shape, skip counts, empty-harvest checklist

Pick a guide: mixed IR → scientific. CKAN/Socrata-like → opendata. CSW/STAC/ArcGIS catalog → geoportals. Map UI only → viewers. Gridded EO → earthdata. IPT/Symbiota/GBIF/Breedbase → biodiversity. Tables/dataflows → indicators. FDP/MDR → metadata. from= / checkpoints → incremental. Shared OAI/CSW/DCAT grain → protocols. Aggregators/ML/custom → other. JSON records / empty results → output.

Do not apply IR publication filters to WMS or PxWeb. Do not treat CSW service records or STAC items as datasets unless that is the catalog grain.

Workflow​

  1. Pick catalogs from exports, not by walking YAML. Prefer status = active, api = true, and a known software.id.
  2. Read endpoints[] on the record. Those URLs were probed for this registry. If empty, use the platform default paths in the guides — still GET only, on that host.
  3. Identify, then filter to datasets (server-side query or OAI set), then paginate.
  4. Store dataset identifiers (native id, DOI/handle when present) plus the catalog uid (harvest-identifiers.md). Emit the output record. Do not write dataset YAML into this repository.
  5. Stop on 401 / 403. Do not guess API keys or follow login forms.
  6. Later runs: reuse the same filter with a date/token checkpoint (harvest-incremental.md).
SELECT id, uid, name, link,
software.id AS software_id,
endpoints
FROM catalogs
WHERE catalog_type = 'Scientific data repository'
AND status = 'active'
AND software.id IN (
'dspace', 'dspacecris', 'invenio', 'inveniordm', 'eprints',
'hyrax', 'pure', 'esploro', 'opus', 'elsevierdigitalcommons',
'weko3', 'dabar', 'opensciencesi', 'phaidra', 'figshare', 'haplo', 'worktribe', 'mycore',
'omegapsir', 'converis', 'dlcm', 'easydb', 'djehuty', 'islandora',
'ipt', 'thredds', 'erddap', 'radar', 'frostserver', 'geonature',
'yoda', 'redivis', 'symbiota', 'breedbase', 'tripal', 'veupathdb',
'massbank', 'iochembd', 'esgf', 'labkey', 'synapse', 'xnat',
'omero', 'kadi4mat', 'edal', 'nomad', 'intermine', 'gringlobal',
'plutof', 'jgi', 'cbioportal', 'esasciencearchive', 'archipelago',
'talkbank', 'clld', 'dachs', 'specify', 'divaportal', 'huggingface',
'databus', 'dialnetcris'
)
LIMIT 50;

Filter software.id only with values that exist in data/software/ (currently 467 definitions). Do not invent an id that is missing from software_ids.yaml.

For open data or geo, change catalog_type and the software.id list (ckan, geonetwork, stacserver, …). Nested software / endpoints are STRUCT / STRUCT[] in DuckDB (ai-consumers.md).

Why scientific repositories need extra filters​

Open-data portals (CKAN, Socrata) list datasets as the primary object. Institutional repositories and CRIS portals list research outputs: journal articles, theses, presentations, code, and — sometimes — datasets.

If you page /api/records or OAI ListRecords with no type filter, most hits are publications. The scientific guide is the filter cookbook.

Prefer server-side filters (search q=, facet, OAI setSpec). Client-side dc:type matching is the fallback when the API has no type parameter. Vocabularies differ per campus — always ListSets / inspect one sample record before a full crawl.

Keep vs drop (shared vocabulary)​

Keep a record when its type is clearly research data (including data papers only if they deposit data files — prefer the dataset record).

Keep (examples)Drop (examples)
Dataset, DataSet, Research Data, Forschungsdaten, ResearchDataArticle, Journal Article, Review
Data collection, DataCollection, DatabaseThesis, Dissertation, Doctoral thesis, Master thesis
Tabular data, Geospatial data, Census dataConference paper, Presentation, Poster, Lecture
COAR c_ddb1 (dataset)COAR c_6501 (journal article), c_46ec (thesis)
DataCite resourceTypeGeneral=DatasetDataCite Text, Image, Audiovisual, Other
Figshare item_type=3 (dataset), 4 (fileset)Figshare paper, poster, presentation, thesis

Also drop: user accounts, projects, org units, researcher profiles, harvest-source records, individual files when a parent dataset record exists (Dataverse type=file vs type=dataset).

Other catalog types use a different grain: CSW dataset/series not services; STAC collections not items; CKAN packages not resources; PxWeb / PxStat tables not folders; TabNet .def tables not CGI sessions; FENIX domains not observation cubes; IMF NSDP SDMX categories/series not the DSBB directory; ResourceContracts contracts not clause chrome; ODWeb /odweb/ datasets not the parent CMS; IPT archives not occurrences; PlutoF datasets not occurrences; InterMine experiments/datasets not gene pages; cBioPortal studies not mutation rows; XNAT projects not sessions; OMERO projects/screens not images; ESA TAP tables not FITS files; map viewers named layers not tiles (harvest-protocols.md).

Software/code and models are not datasets unless the catalog types them as data. Index them separately if you need them.

Politeness​

  • GET public URLs only. Short timeout. One or two probes to learn paging, then polite page size (10–100).
  • Honor Retry-After and back off on 429.
  • Do not write internet-wide scanners in this repository. Harvest named catalogs from the registry.
  • Prefer OAI-PMH ListRecords with a set and resumption tokens over scraping HTML.
  • After a catalog is registered here, python scripts/apidetect.py detect-single CATALOG_ID --dryrun can fill endpoints[] — that is not a dataset crawl.

What not to store here​

  • Dataset-level YAML, CKAN packages, Dataverse studies, STAC items
  • Copies of harvested JSON in data/entities/
  • API keys, cookies, or session tokens