Harvesting datasets from catalog APIs
This registry stores catalogs (portals, geoportals, repositories). It does not store the datasets inside those catalogs. To list or index datasets, harvest the catalog’s public API.
Two different jobs share the word harvest:
| Job | What you want | Where to go |
|---|---|---|
| List catalogs in this registry | Country, type, software.id, endpoints[] | query-examples.md, agents/query.md |
| List datasets inside a catalog | Records from the remote API, filtered to datasets | This page, then the platform guides |
| Find catalogs not yet registered | New portal URLs | discovery.md |
Coding agents: agents/harvest.md. Production harvesting for the Dateno stack lives in reaper — these pages are the human/API recipes, not a crawler in this repository.
Guides
| Guide | Use when |
|---|---|
| Scientific repositories | DSpace, Invenio, EPrints, Pure, Esploro, DLCM, easydb, Haplo, Omega-PSIR, Hyrax, Archipelago, LabKey, Synapse, XNAT, OMERO, Kadi4Mat, e!DAL, NOMAD, and other IRs that mix publications, theses, software, and datasets |
| Domain scientific repositories | IPT, Biodiv, THREDDS, ERDDAP, FROST-Server, GeoNature, Specify, DaCHS, Breedbase, Tripal, VEuPathDB, MassBank, ioChem-BD, ESGF, InterMine, GRIN-Global, PlutoF, JGI, cBioPortal |
| Open data portals | CKAN, OpenDataSoft, Socrata, RUDI, GIS Open Data Portal, ResourceContracts, RDF Online Repository, Guangxi, ODWeb, OpenGov — packages vs resources vs contracts vs reports |
| Geoportals | GeoNetwork CSW, GeoNode, ArcGIS, STAC, OGC API, OneGeo Suite, PRODIGE, G3W-SUITE, CubeWerx, M.App Enterprise, mviewer, Isogeo, Geocortex, QGIS Server — layers vs services vs tiles |
| Indicators and microdata | PxWeb / PxStat tables, DGBAS Web tables, SDMX dataflows, OpenSDG / Goal Tracker, IMF NSDP, NADA studies, DHIS2, TabNet, SIDRA, FENIX, SparkMap, eDatos, Cancer-Rates.info, HCI, Virtual LMI, Fingertips, KOSIS, e-Stat, UNdata, Comtrade Plus, Our World in Data, IPUMS |
| Metadata catalogs | FAIR Data Point DCAT, Aristotle MDR, Fusion Registry, Metadata Browser, CBD Clearing-House, WMO OSCAR |
| Search, ML, API, marketplaces | Aggregators, OpenAIRE, OpenML, Hugging Face, CodaLab, API directories, marketplaces, custom |
| Protocols | OAI-PMH, CSW, DCAT, STAC, SDMX, OGC, ArcGIS REST — grain that is shared across products |
| Incremental harvests | from=, metadata_modified, STAC datetime, checkpoints |
| Earth observation | THREDDS, ERDDAP, STAC collections, Open Data Cube, Copernicus, openEO, Sentinel Hub, ESGF, ESA Science Archive |
| Biodiversity and genomics | IPT, Symbiota, ALA, GBIF datasets, Ensembl species, Breedbase, Tripal, VEuPathDB, PlutoF, InterMine, JGI, cBioPortal |
| Map viewers | QWC2, Masterportal, Lizmap, mviewer, WebEWID, GISPLAN, Wagmap, Trimble Locus / Louhi / Landfolio, Spatial Suite, Spectrum Spatial Analyst, Hajk, myCarta, KortInfo, IntraMaps, LocalMaps, GEUSMAP, GISApp, GeneGIS PAGIS, GisMaster, GeoPortale.cloud, LDP SIT, GFMaplet, HyG Mapgis, iObčina, iShare, Cadcorp, StatMap Earthlight — layers not tiles |
| Dataset identifiers | Native id + catalog uid; DOI/handle; do not mint cdi######## for datasets |
| Harvest output | JSON record shape, skip counts, empty-harvest checklist |
Pick a guide: mixed IR → scientific. CKAN/Socrata-like → opendata. CSW/STAC/ArcGIS catalog → geoportals. Map UI only → viewers. Gridded EO → earthdata. IPT/Symbiota/GBIF/Breedbase → biodiversity. Tables/dataflows → indicators. FDP/MDR → metadata. from= / checkpoints → incremental. Shared OAI/CSW/DCAT grain → protocols. Aggregators/ML/custom → other. JSON records / empty results → output.
Do not apply IR publication filters to WMS or PxWeb. Do not treat CSW service records or STAC items as datasets unless that is the catalog grain.
Workflow
- Pick catalogs from exports, not by walking YAML. Prefer
status = active,api = true, and a knownsoftware.id. - Read
endpoints[]on the record. Those URLs were probed for this registry. If empty, use the platform default paths in the guides — still GET only, on that host. - Identify, then filter to datasets (server-side query or OAI set), then paginate.
- Store dataset identifiers (native id, DOI/handle when present) plus the catalog
uid(harvest-identifiers.md). Emit the output record. Do not write dataset YAML into this repository. - Stop on
401/403. Do not guess API keys or follow login forms. - Later runs: reuse the same filter with a date/token checkpoint (harvest-incremental.md).
SELECT id, uid, name, link,
software.id AS software_id,
endpoints
FROM catalogs
WHERE catalog_type = 'Scientific data repository'
AND status = 'active'
AND software.id IN (
'dspace', 'dspacecris', 'invenio', 'inveniordm', 'eprints',
'hyrax', 'pure', 'esploro', 'opus', 'elsevierdigitalcommons',
'weko3', 'dabar', 'opensciencesi', 'phaidra', 'figshare', 'haplo', 'worktribe', 'mycore',
'omegapsir', 'converis', 'dlcm', 'easydb', 'djehuty', 'islandora',
'ipt', 'thredds', 'erddap', 'radar', 'frostserver', 'geonature',
'yoda', 'redivis', 'symbiota', 'breedbase', 'tripal', 'veupathdb',
'massbank', 'iochembd', 'esgf', 'labkey', 'synapse', 'xnat',
'omero', 'kadi4mat', 'edal', 'nomad', 'intermine', 'gringlobal',
'plutof', 'jgi', 'cbioportal', 'esasciencearchive', 'archipelago',
'talkbank', 'clld', 'dachs', 'specify', 'divaportal', 'huggingface',
'databus', 'dialnetcris'
)
LIMIT 50;
Filter software.id only with values that exist in data/software/ (currently 467 definitions). Do not invent an id that is missing from software_ids.yaml.
For open data or geo, change catalog_type and the software.id list (ckan, geonetwork, stacserver, …). Nested software / endpoints are STRUCT / STRUCT[] in DuckDB (ai-consumers.md).
Why scientific repositories need extra filters
Open-data portals (CKAN, Socrata) list datasets as the primary object. Institutional repositories and CRIS portals list research outputs: journal articles, theses, presentations, code, and — sometimes — datasets.
If you page /api/records or OAI ListRecords with no type filter, most hits are publications. The scientific guide is the filter cookbook.
Prefer server-side filters (search q=, facet, OAI setSpec). Client-side dc:type matching is the fallback when the API has no type parameter. Vocabularies differ per campus — always ListSets / inspect one sample record before a full crawl.
Keep vs drop (shared vocabulary)
Keep a record when its type is clearly research data (including data papers only if they deposit data files — prefer the dataset record).
| Keep (examples) | Drop (examples) |
|---|---|
| Dataset, DataSet, Research Data, Forschungsdaten, ResearchData | Article, Journal Article, Review |
| Data collection, DataCollection, Database | Thesis, Dissertation, Doctoral thesis, Master thesis |
| Tabular data, Geospatial data, Census data | Conference paper, Presentation, Poster, Lecture |
COAR c_ddb1 (dataset) | COAR c_6501 (journal article), c_46ec (thesis) |
DataCite resourceTypeGeneral=Dataset | DataCite Text, Image, Audiovisual, Other |
Figshare item_type=3 (dataset), 4 (fileset) | Figshare paper, poster, presentation, thesis |
Also drop: user accounts, projects, org units, researcher profiles, harvest-source records, individual files when a parent dataset record exists (Dataverse type=file vs type=dataset).
Other catalog types use a different grain: CSW dataset/series not services; STAC collections not items; CKAN packages not resources; PxWeb / PxStat tables not folders; TabNet .def tables not CGI sessions; FENIX domains not observation cubes; IMF NSDP SDMX categories/series not the DSBB directory; ResourceContracts contracts not clause chrome; ODWeb /odweb/ datasets not the parent CMS; IPT archives not occurrences; PlutoF datasets not occurrences; InterMine experiments/datasets not gene pages; cBioPortal studies not mutation rows; XNAT projects not sessions; OMERO projects/screens not images; ESA TAP tables not FITS files; map viewers named layers not tiles (harvest-protocols.md).
Software/code and models are not datasets unless the catalog types them as data. Index them separately if you need them.
Politeness
- GET public URLs only. Short timeout. One or two probes to learn paging, then polite page size (
10–100). - Honor
Retry-Afterand back off on429. - Do not write internet-wide scanners in this repository. Harvest named catalogs from the registry.
- Prefer OAI-PMH
ListRecordswith asetand resumption tokens over scraping HTML. - After a catalog is registered here,
python scripts/apidetect.py detect-single CATALOG_ID --dryruncan fillendpoints[]— that is not a dataset crawl.
What not to store here
- Dataset-level YAML, CKAN packages, Dataverse studies, STAC items
- Copies of harvested JSON in
data/entities/ - API keys, cookies, or session tokens
Related
- harvest-scientific.md
- harvest-scientific-domain.md
- harvest-opendata.md
- harvest-geoportals.md
- harvest-indicators.md
- harvest-metadata.md
- harvest-other.md
- harvest-protocols.md
- harvest-incremental.md
- harvest-earthdata.md
- harvest-biodiversity.md
- harvest-viewers.md
- harvest-identifiers.md
- harvest-output.md
- agents/harvest.md
- apidetect.md
- when-to-use.md