Harvesting metadata catalogs
Metadata catalogs publish DCAT/RDF catalogs, data-element registries, or SDMX structure. Harvest catalog and dataset metadata, not file bytes and not every code list as a dataset unless that is the product.
Overview: harvest.md. Finding installations: discovery-metadata.md. GET only. Stop on 401/403.
What to keep
| Keep | Drop |
|---|---|
| FAIR Data Point Catalog and child Dataset IRIs | Binary distributions as extra datasets (they are files) |
| Aristotle stewarded public objects the user asked for (datasets vs data elements) | Staff-only workflow items |
| Fusion Registry dataflows (dataset analog) | Codelists and DSDs unless harvesting structure |
| Metadata Browser public dataset records | Terminology-only hits if you want datasets |
If the live product is CKAN, GeoNetwork, or Dataverse, use those harvest guides instead of this page.
FAIR Data Point (fairdatapoint)
FDP is RDF DCAT. Start at the catalog root with RDF Accept headers.
GET https://host/
Accept: text/turtle
Also try application/ld+json. Follow dcat:dataset (and nested dcat:catalog) IRIs. Harvest each Dataset resource once. Drop dcat:Distribution as a separate dataset (keep URLs as files on the parent). Cap recursion through child FDPs so you do not crawl the whole federation.
HTML fdp-client alone is not a harvest. Swagger /swagger-ui documents the API — use it to find catalog/dataset paths. Register/harvest the FDP root, not a single dataset IRI as the catalog.
Keep: FAIR Data Point Catalog and child Dataset IRIs. Drop: dcat:Distribution as a separate dataset (keep URLs as files on the parent) and the public index as if it were every child FDP.
Docs: docs.fairdatapoint.org. Public index: home.fairdatapoint.org (do not re-harvest the index as if it were every child FDP).
Aristotle MDR (aristotlemdr)
GET https://host/api/v4/
GET https://host/api/v4/metadata/
GET https://host/api/v4/metadata
This is a metadata registry (object classes, data elements, value domains). Those are not research-data files.
- If the user wants datasets: keep Aristotle types that represent a dataset/distribution (names vary by stewardship model) and skip data-element noise.
- If the user wants metadata objects: page
/api/v4/metadata/and record native ids — still do not write them into this registry’s YAML.
Skip login-only stewardship UIs.
Keep: Aristotle types that represent a dataset/distribution when the user wants datasets; otherwise page /api/v4/metadata/ for metadata objects. Drop: staff-only workflow items and login-only stewardship UIs.
Fusion Registry (fusionregistry)
SDMX structural metadata.
GET https://host/ws/public/sdmxapi/rest
GET https://host/ws/rest
List dataflows as the harvest grain for “datasets”. Harvest DSDs/codelists only when the job is a structure crawl. Do not confuse this with PxWeb/.Stat observation APIs (harvest-indicators.md) or with Fusion Data Browser series catalogs (fusiondatabrowser).
Keep: Fusion Registry dataflows. Drop: codelists and DSDs unless harvesting structure.
Detection also probes origin /ws/rest and /ws/public/sdmxapi/rest when the catalog link is a /FusionRegistry (or /FusionMetadataRegistry) mount.
Metadata Browser (mwmb)
Public MetadataWorks UI. There may be no stable open list API. Harvest the public dataset/standard listing if a JSON/search endpoint exists in endpoints[]. Skip terminology-only pages when the user asked for datasets. One deployment = one catalog harvest scope.
Keep: public dataset/standard listing if a JSON/search endpoint exists in endpoints[].
Drop: terminology-only pages when the user asked for datasets.
GET https://host/api/
DataHub (datahubproject)
GET https://host/api/graphql
Keep: dataset / data-product metadata entities. Drop: users, glossary terms, and lineage edges as datasets. GraphQL is often authenticated (401) — stop rather than scraping the SPA. One harvest scope per public DataHub catalog. Distinct from datahub.io (CKAN).
CEDAR Workbench (cedar)
Harvest the public template / metadata-instance catalog, not the marketing site.
GET https://cedar.metadatacenter.org
Keep: published metadata templates and filled experiment metadata records. Drop: login, template-designer chrome, and the metadatacenter.org homepage.
CBD Clearing-House (cbdchm)
Public records of the CBD clearing-house realms (CHM, ABSCH, BCH).
GET https://chm.cbd.int/en/
GET https://absch.cbd.int/en/
GET https://bch.cbd.int/en/
The shared API is documented at docs.cbddev.xyz. Production calls go to api.cbd.int.
Keep: published national records, certificates, decisions, and country profiles. Drop: login, draft submission workflows, and the developer documentation portal. One harvest scope per realm. Bioland national portals are not this software.
WMO OSCAR (oscar)
Station, satellite, and instrument metadata. Read endpoints are public; writes require an account.
GET https://oscar.wmo.int/surface/rest/api/search/station
GET https://space.oscar.wmo.int/apidoc/
OSCAR/Surface search results are paged (stationSearchResults, page, items). OSCAR/Space record JSON is described in the API doc on that host.
Keep: stations, satellites, instruments, and observation variables. Drop: user accounts, gap-analysis chrome, and the WMO Weather Radar Database (wrd.mgm.gov.tr), which is not OSCAR.
DCAT without an FDP
Many open-data sites expose /catalog.xml, /data.json, or DCAT-AP. That harvest belongs with harvest-opendata.md (dcat:Dataset only). Protocol details: harvest-protocols.md. Use this page when software.id is fairdatapoint, aristotlemdr, fusionregistry, mwmb, datahubproject, cedar, cbdchm, or oscar.