Skip to main content

Harvesting metadata catalogs

Metadata catalogs publish DCAT/RDF catalogs, data-element registries, or SDMX structure. Harvest catalog and dataset metadata, not file bytes and not every code list as a dataset unless that is the product.

Overview: harvest.md. Finding installations: discovery-metadata.md. GET only. Stop on 401/403.

What to keep​

KeepDrop
FAIR Data Point Catalog and child Dataset IRIsBinary distributions as extra datasets (they are files)
Aristotle stewarded public objects the user asked for (datasets vs data elements)Staff-only workflow items
Fusion Registry dataflows (dataset analog)Codelists and DSDs unless harvesting structure
Metadata Browser public dataset recordsTerminology-only hits if you want datasets

If the live product is CKAN, GeoNetwork, or Dataverse, use those harvest guides instead of this page.

FAIR Data Point (fairdatapoint)​

FDP is RDF DCAT. Start at the catalog root with RDF Accept headers.

GET https://host/
Accept: text/turtle

Also try application/ld+json. Follow dcat:dataset (and nested dcat:catalog) IRIs. Harvest each Dataset resource once. Drop dcat:Distribution as a separate dataset (keep URLs as files on the parent). Cap recursion through child FDPs so you do not crawl the whole federation.

HTML fdp-client alone is not a harvest. Swagger /swagger-ui documents the API — use it to find catalog/dataset paths. Register/harvest the FDP root, not a single dataset IRI as the catalog.

Keep: FAIR Data Point Catalog and child Dataset IRIs. Drop: dcat:Distribution as a separate dataset (keep URLs as files on the parent) and the public index as if it were every child FDP.

Docs: docs.fairdatapoint.org. Public index: home.fairdatapoint.org (do not re-harvest the index as if it were every child FDP).

Aristotle MDR (aristotlemdr)​

GET https://host/api/v4/
GET https://host/api/v4/metadata/
GET https://host/api/v4/metadata

This is a metadata registry (object classes, data elements, value domains). Those are not research-data files.

  • If the user wants datasets: keep Aristotle types that represent a dataset/distribution (names vary by stewardship model) and skip data-element noise.
  • If the user wants metadata objects: page /api/v4/metadata/ and record native ids — still do not write them into this registry’s YAML.

Skip login-only stewardship UIs.

Keep: Aristotle types that represent a dataset/distribution when the user wants datasets; otherwise page /api/v4/metadata/ for metadata objects. Drop: staff-only workflow items and login-only stewardship UIs.

Fusion Registry (fusionregistry)​

SDMX structural metadata.

GET https://host/ws/public/sdmxapi/rest
GET https://host/ws/rest

List dataflows as the harvest grain for “datasets”. Harvest DSDs/codelists only when the job is a structure crawl. Do not confuse this with PxWeb/.Stat observation APIs (harvest-indicators.md) or with Fusion Data Browser series catalogs (fusiondatabrowser).

Keep: Fusion Registry dataflows. Drop: codelists and DSDs unless harvesting structure.

Detection also probes origin /ws/rest and /ws/public/sdmxapi/rest when the catalog link is a /FusionRegistry (or /FusionMetadataRegistry) mount.

Metadata Browser (mwmb)​

Public MetadataWorks UI. There may be no stable open list API. Harvest the public dataset/standard listing if a JSON/search endpoint exists in endpoints[]. Skip terminology-only pages when the user asked for datasets. One deployment = one catalog harvest scope.

Keep: public dataset/standard listing if a JSON/search endpoint exists in endpoints[]. Drop: terminology-only pages when the user asked for datasets.

GET https://host/api/

DataHub (datahubproject)​

GET https://host/api/graphql

Keep: dataset / data-product metadata entities. Drop: users, glossary terms, and lineage edges as datasets. GraphQL is often authenticated (401) — stop rather than scraping the SPA. One harvest scope per public DataHub catalog. Distinct from datahub.io (CKAN).

CEDAR Workbench (cedar)​

Harvest the public template / metadata-instance catalog, not the marketing site.

GET https://cedar.metadatacenter.org

Keep: published metadata templates and filled experiment metadata records. Drop: login, template-designer chrome, and the metadatacenter.org homepage.

CBD Clearing-House (cbdchm)​

Public records of the CBD clearing-house realms (CHM, ABSCH, BCH).

GET https://chm.cbd.int/en/
GET https://absch.cbd.int/en/
GET https://bch.cbd.int/en/

The shared API is documented at docs.cbddev.xyz. Production calls go to api.cbd.int.

Keep: published national records, certificates, decisions, and country profiles. Drop: login, draft submission workflows, and the developer documentation portal. One harvest scope per realm. Bioland national portals are not this software.

WMO OSCAR (oscar)​

Station, satellite, and instrument metadata. Read endpoints are public; writes require an account.

GET https://oscar.wmo.int/surface/rest/api/search/station
GET https://space.oscar.wmo.int/apidoc/

OSCAR/Surface search results are paged (stationSearchResults, page, items). OSCAR/Space record JSON is described in the API doc on that host.

Keep: stations, satellites, instruments, and observation variables. Drop: user accounts, gap-analysis chrome, and the WMO Weather Radar Database (wrd.mgm.gov.tr), which is not OSCAR.

DCAT without an FDP​

Many open-data sites expose /catalog.xml, /data.json, or DCAT-AP. That harvest belongs with harvest-opendata.md (dcat:Dataset only). Protocol details: harvest-protocols.md. Use this page when software.id is fairdatapoint, aristotlemdr, fusionregistry, mwmb, datahubproject, cedar, cbdchm, or oscar.