Skip to main content

Harvesting search engines, ML catalogs, and other types

Catalog types without a dedicated high-volume harvest page: data search engines, ML catalogs, API directories, marketplaces, dataset lists, and software.id: custom. Type rules: catalog-types.md. Finding them: discovery-other.md. Overview: harvest.md.

GET public URLs only. Stop on 401/403. Do not invent a new software.id.

Data search engines​

Aggregators search other catalogs (Idra, national dataset search, harvested unions).

Prefer harvesting the source portals listed in this registry. Harvesting the aggregator duplicates those datasets and loses the true uid of the source catalog.

If the user explicitly wants the federation view:

  • Idra (idra): /Idra/api/v1/ dataset search. Keep federated datasets. Record the source catalog URL when Idra provides it.
  • OpenAIRE — below. Prefer source IRs.
  • Aleph (aleph): see below.
  • custom search engines: use their documented search API; store source identifiers.

Do not add harvested member catalogs as new registry YAML (that is discovery.md).

OpenAIRE (openaire)​

EXPLORE / CONNECT gateways over the OpenAIRE Graph. Filter exports on software.id = 'openaire'.

GET https://api.openaire.eu/search/datasets

Keep: Graph datasets (research products typed as dataset). Drop: publications, software, and org units. For a CONNECT community portal, use that gateway’s search/API with the community filter — do not dump the whole European graph. Prefer harvesting source IRs from this registry when you need publisher-level ids. The Graph data-source harvest list is openaire-sync.md. Stop on 401.

Aleph (aleph)​

OCCRP-style investigative collections.

GET https://host/api/2/collections

Keep: collections that are dataset corpora. Drop: entity/document search hits (/api/2/entities) as datasets. Page collection results; do not crawl every PDF. Skip aleph.occrp.org if you only needed the existing registry record.

Machine learning catalogs​

OpenML, Galaxy, and similar hubs list datasets (or data libraries), not tasks, runs, or user spaces. Regional challenge platforms (Zindi, AIcrowd, SIGNATE, Grand Challenge) are usually custom — harvest the hub’s public dataset list, not every competition submission.

OpenML (openmlorg)​

GET https://www.openml.org/api/v1/json/data/list/limit/100/offset/0

Keep: datasets (data). Drop: tasks, flows, runs, and setups. Do not harvest every OpenML task as a dataset. Skip cloning openml.org if you only needed the existing registry record — harvest contents when the user asked.

Kaggle (kaggle)​

GET https://www.kaggle.com/api/v1/datasets/list

Keep: datasets. Drop: notebooks, competitions, and models. The hub is one registered catalog. The list API often needs a Kaggle token — if it 401s, harvest the public HTML datasets catalog only.

Hugging Face (huggingface)​

GET https://huggingface.co/api/datasets

Keep: datasets. Drop: models, Spaces, and per-user repos. The Hub is usually one registered catalog.

CodaLab Competitions (codalab)​

GET https://competitions.codalab.org/competitions/

Keep: public competitions that publish dataset bundles, or the hub’s public dataset list if exposed. Drop: submissions, leaderboards, and per-user workspaces. One registered hub.

Codabench (codabench)​

GET https://www.codabench.org/

Keep: public benchmarks / competitions and any public dataset catalog pages. /api/datasets/ is often 403 — stop and use the HTML catalog. Drop: submissions and login-only datasets. One registered hub.

Grand Challenge (grandchallenge)​

GET https://grand-challenge.org/challenges/
GET https://grand-challenge.org/archives/

Keep: public challenges and archives (dataset collections), including a hub /data/ page that lists datasets. Drop: algorithms as extra catalogs, reader studies, submissions, per-user workspaces, and track subdomains of an install already registered. /api/v1/ may return 403 to non-browser clients — stop and use the HTML lists. One catalog per installation.

EvalAI (evalai)​

GET https://eval.ai/

Keep: public challenges and any advertised dataset bundles. Drop: submissions, leaderboards as extra datasets, and login-only drafts. One registered hub.

Galaxy (galaxy)​

Public data libraries are the dataset catalog. Histories and workflow runs are not. Prefer harvest-scientific.md unless catalog_type is Machine learning catalog.

GET https://host/api/libraries

Keep: public data libraries. Drop: histories, workflows, and job outputs. Stop on 401 for user workspaces.

Hugging Face (huggingface) has its own harvest section above. Kaggle and Papers with Code are usually one registered hub each. Harvest their public dataset APIs if asked; do not add per-user spaces as catalogs.

API catalogs​

The product is a list of APIs, not datasets. Harvest API entries (name, docs URL, publisher) only when the user wants an API inventory. Store the API’s stable id + the catalog uid (harvest-identifiers.md). A CKAN Action API stays an open-data harvest (harvest-opendata.md) unless the catalog itself is an API directory on CKAN.

Azure API Management (azureapim)​

Harvest the developer portal's API list. One portal is one catalog.

GET https://host/apis

Keep: API products listed in the developer portal. Drop: subscription keys, user profiles, and test-console calls.

WSO2 API Manager (wso2apimanager)​

Harvest the developer portal's API list. One portal is one catalog.

GET https://host/devportal/apis

Keep: APIs listed in the developer portal (name, docs URL, publisher). Drop: subscription keys, application registrations, and user profiles. Unauthenticated /devportal REST endpoints vary by version — stop on 401 rather than probing admin APIs.

Data marketplaces​

Public catalog of datasets for sale or license. Harvest public listing APIs only (dataset id, title, landing URL). Stop on paywalls and 401. Do not scrape prices, buyer lists, or sample files into this repository. access_mode on the catalog record is often restricted — still harvest metadata if it is public.

Datasets lists​

HTML tables, spreadsheets, GitHub inventories (catalog_type: Datasets list). One row / bullet with a dataset title + URL = one dataset. Skip the wrapping README as a dataset. No CMS API — parse the published file the catalog link points at, not a site-wide scrape.

Squarespace (squarespace)​

Filter exports on software.id = 'squarespace'. One harvest scope per site.

No data API. Keep the public data / indicator pages (and the files they link) listed in the site navigation or data page. Grain is the published table, report, or file — not each marketing page. Drop blog, about, and contact chrome.

Keep: public data / indicator pages and linked files. Drop: marketing and blog chrome.

Jekyll (jekyll)​

Filter exports on software.id = 'jekyll'. One harvest scope per project site.

No data API; Jekyll sites are static HTML. Keep the public dataset listing pages and the release/download files they link. Grain is the dataset or benchmark entry, not each docs post. Drop documentation, blog, and paper pages. Many Jekyll catalogs also have a GitHub repo — release assets there count as the dataset files.

Keep: public dataset listing pages and linked release files. Drop: docs, blog, and paper pages.

Custom software (custom)​

About one catalog in eight has no shared product ID. Do not guess CKAN, DSpace, or GeoNetwork filters.

Decision tree

  1. Confirm software.id is custom in exports (or two independent fingerprints failed in discovery.md).
  2. GET endpoints[] if any. Use those URLs first.
  3. Protocol fallback (stop at the first public list):
    • DCAT / data.json / /catalog.json — keep dcat:Dataset only (harvest-protocols.md)
    • OAI-PMH Identify → ListSets → dataset-named sets (harvest-scientific.md)
    • CSW GetRecords / STAC /collections / OGC API /collections?f=json (harvest-geoportals.md, harvest-protocols.md)
    • CKAN-shaped /api/3/action/package_search only if the JSON is actually CKAN ("help" + "success") — then the record should not stay custom
  4. If the site is an HTML table, spreadsheet, or GitHub inventory, parse that file only (Datasets lists).
  5. Stop when none of those exist. Do not HTML-scrape the whole CMS from this repository’s workflows. Report login on 401/403. A correct empty harvest is valid.

Worked examples (high-count custom families)

FamilyRegistry grainHarvestStop
EMBL-EBI resources (www.ebi.ac.uk/pride, /ena, /gwas, …)One YAML per resource, not ebi.ac.ukThat resource’s public dataset/study list or documented RESTThe EBI homepage, gene pages, every file
NCBI (/sra, /geo, /datasets, PubChem, …)One YAML per databaseThat database’s public dataset/accession catalogncbi.nlm.nih.gov chrome, BLAST jobs
Hugging Face / Kaggle / Papers with CodeOne hub catalogPublic dataset API/listNotebooks, models, competitions, user spaces
Zindi / AIcrowd / SIGNATEOne challenge hubPublic dataset listSubmissions, leaderboards
Municipal Excel / HTML inventoryThe file the link points atRows with title + URLThe wrapping CMS

Do not invent a new software.id for a one-off.

Keep: the first public dataset list from endpoints[] or the protocol fallback (DCAT, OAI dataset set, CSW/STAC/OGC collections, or a published HTML/CSV inventory). Drop: CMS-wide scrapes, guessed CKAN/DSpace/GeoNetwork filters, gene pages, and 401/403 login walls.