Harvesting datasets from scientific repositories
Institutional repositories and CRIS portals mix publications, theses, software, and research data. Harvest the public API, then filter to datasets. Overview and keep/drop vocabulary: harvest.md. Finding installations: discovery-scientific.md. Domain stacks: discovery-scientific-domain.md.
Replace https://host with the catalog link origin (no trailing slash unless the path needs it). GET only. Stop on 401/403.
Use endpoints[] from the registry when present (apidetect.md). Paths below are the defaults those maps probe.
| Page | Use when |
|---|---|
| This page | Institutional repositories and CRIS (Dataverse, DSpace, Invenio, EPrints, Pure, Converis, Omega-PSIR, Archipelago, RADAR, Yoda, Redivis, LabKey, Synapse, BioUML, XNAT, OMERO, Kadi4Mat, e!DAL, NOMAD, META-SHARE, Gen3, TR32DB, Index Data Keystone, …) |
| Domain repositories | IPT, Symbiota, THREDDS, ERDDAP, Breedbase, Tripal, VEuPathDB, MassBank, ioChem-BD, ESGF, ALA, SciCat-adjacent stacks, CLLD, TalkBank, Pathway Tools, IBDC |
All software.id values: software-index.md.
Mixed vs dataset-native
| Class | software.id (typical) | Filter needed? |
|---|---|---|
| Mixed IR / CRIS | dspace, dspacecris, invenio, inveniordm, eprints, hyrax, samvera, islandora, archipelago, opus, mycore, phaidra, weko3, dabar, opensciencesi, pure, esploro, elsevierdigitalcommons, figshare, haplo, worktribe, omegapsir, converis, librecat, vufind, divaportal, keystoneils | Yes — publications dominate |
| Dataset-native | dataverse, radar, yoda, redivis, instdb, labkey, synapse, xnat, omero, kadi4mat, edal, nomad, gen3 on this page; IPT/THREDDS/Breedbase/ESGF/InterMine/cBioPortal and similar on harvest-scientific-domain.md | Little or none — still skip files, occurrences, and login-only rows |
OAI-PMH fallback (any IR)
When REST search has no type filter, use OAI-PMH (harvest-protocols.md).
GET https://host/oai/request?verb=Identify(DSpace) or the Identify URL inendpoints[].verb=ListSets— keepsetSpecvalues that mean data (ResearchData,doc-type:researchdata,datasets,Dataset,Forschungsdaten). Ignorecom_/col_sets that are the whole repository.verb=ListRecordswithmetadataPrefix=oai_dcandsetequal to thatsetSpec. FollowresumptionToken.- If no dataset set exists, harvest
oai_dcand keep records whosedc:type(or DataCiteresourceTypeGeneral, or COAR URI) matches the keep list.
Common Identify paths: /oai?verb=Identify, /oai/request?verb=Identify, /cgi/oai2?verb=Identify, /oai2d, /ws/oai?verb=Identify, /api/oai?verb=Identify.
Do not treat ListIdentifiers titles as datasets. Do not harvest metadataPrefix=marc21 as a substitute for type.
Dataverse (dataverse)
Native search already distinguishes objects. Prefer datasets, not files or sub-dataverses.
List datasets:
GET https://host/api/search?q=*&type=dataset&per_page=100&start=0
Page with start. total_count is in the JSON envelope.
Also useful: /api/info/version (dataverseapi), OAI /oai?verb=Identify.
Drop: type=file (file hits under a dataset), type=dataverse (collections), /dataset.xhtml?persistentId= as a crawl seed (that is one record). Harvest the installation root from the registry, then this search API.
Keep: type=dataset search hits.
Docs: guides.dataverse.org.
DSpace 7+ (dspace)
DSpace items are publications, theses, and datasets in one index.
Unfiltered (do not use as the crawl): /server/api/discover/search/objects
Worked example A — DSpace 7 entity type
Filter to dataset entities (DSpace-CRIS / configurable entities):
GET https://host/server/api/discover/search/objects?dsoType=ITEM&f.entityType=Dataset,equals&size=100&page=0
Some campuses name the entity ResearchData or Product. Inspect facets once:
GET https://host/server/api/discover/search/objects?dsoType=ITEM&size=0
Read _embedded.searchResult.page and facet values for entityType / dc.type. If there is no entity type, filter Solr-style:
GET https://host/server/api/discover/search/objects?dsoType=ITEM&query=dc.type:Dataset&size=100
Try Forschungsdaten, Research Data, and Dataset — values are local.
Worked example B — classic OAI ListSets (dc.type)
DSpace 6 and 7 fallback when REST has no entity type:
GET https://host/oai/request?verb=Identify
GET https://host/server/oai/request?verb=Identify
GET https://host/oai/request?verb=ListSets
GET https://host/oai/request?verb=ListRecords&metadataPrefix=oai_dc&set=col_123456789_4
Keep setSpec values whose name is dataset / research data / Forschungsdaten. Ignore com_ community sets that are the whole repository. Then ListRecords with that set. If no dataset set exists, harvest oai_dc and keep records whose dc:type matches the keep list.
DSpace 6 REST: /rest/items has no reliable type filter. Prefer OAI as above, or skip 6.x hosts without a dataset collection.
Drop: dsoType=COMMUNITY / COLLECTION, researcher Person / OrgUnit / Project (CRIS), bitstream URLs.
Keep: items filtered to Dataset / ResearchData (REST entity type, dc.type, or OAI dataset set).
DSpace-CRIS (dspacecris)
Same REST/OAI as DSpace. Prefer f.entityType=Dataset,equals (or the campus ResearchData entity). Drop CRIS Person, OrgUnit, and Project objects.
Keep: Dataset / ResearchData entities via REST or OAI.
Drop: CRIS Person, OrgUnit, and Project objects.
GET https://host/server/api/discover/search/objects?f.entityType=Dataset,equals
Invenio (invenio)
Classic Invenio (not RDM). /api/records returns all record types.
GET https://host/api/records?size=25
Filter to datasets the same way as InvenioRDM, then confirm the UI is not InvenioRDM-branded. Keep: resource_type dataset. Drop: publication, presentation, poster, image, video, lesson, other. software is not a dataset. OAI is often /oai2d?verb=Identify; some classic installs also expose /oai?verb=Identify.
InvenioRDM (inveniordm)
/api/records returns all record types (publication, dataset, software, poster, …).
GET https://host/api/records?q=metadata.resource_type.type:dataset&size=100&page=1
If that query returns zero but the UI has a Dataset facet, try:
GET https://host/api/records?q=metadata.resource_type.id:dataset&size=100
GET https://host/api/records?type=dataset&size=100
Follow links.next. Inspect hits.hits[].metadata.resource_type.
Drop: publication, presentation, poster, image, video, lesson, other unless you explicitly want those corpora. software is not a dataset.
Keep: resource_type dataset (or type=dataset).
OAI is often /oai2d?verb=Identify. Skip zenodo.org if you only need institutional instances already in the registry.
Docs: inveniordm.docs.cern.ch.
HAL (hal)
CCSD HAL open archive. Search/OAI live on api.archives-ouvertes.fr, not necessarily on the portal host. Filter to datasets; publications dominate.
GET https://api.archives-ouvertes.fr/search/?q=*:*&wt=json&rows=0
GET https://api.archives-ouvertes.fr/oai/hal/?verb=Identify
Do not copy those API hosts onto every *.hal.science tenant as if they were local. Keep: docType_s:DATA (or equivalent dataset type). Drop: publications, theses, and conference papers. One harvest scope per public HAL portal.
Keep: dataset deposits. Drop: publications, theses, and conference papers.
Mendeley Data (mendeleydata)
Elsevier hub data.mendeley.com. Harvest the public dataset catalog. Do not crawl every DOI landing page.
GET https://data.mendeley.com/research-data/
Keep: public datasets. Drop: Mendeley Desktop libraries and per-dataset files as crawl seeds. One registered hub.
EPrints (eprints)
Every eprint has a type (article, thesis, dataset, monograph, …).
Browse/export by type:
GET https://host/cgi/exportview/type/dataset/JSON/dataset.js
Search: /cgi/search/archive/advanced with type=dataset (parameter names vary; confirm on one host).
OpenSearch description is /cgi/opensearchdescription. Detection types it opensearch.
GET https://host/cgi/opensearchdescription
REST: /rest/eprint/ plus the numeric eprint id (.xml) is per-record. For a crawl, OAI is better:
GET https://host/cgi/oai2?verb=ListSets
GET https://host/cgi/oai2?verb=ListRecords&metadataPrefix=oai_dc&set=DATASET_SET
If there is no dataset set, ListRecords and keep dc:type = dataset / Dataset.
Drop: article, thesis, book, conference_item, exhibition, performance.
Keep: eprints with type=dataset (exportview, OAI set, or dc:type).
Index Data Keystone (keystoneils)
HTML catalog. Search uses CCL query parameters, not a dataset JSON API.
GET https://host/search?search_type=simple&wf_step=init&cclterm1=&cclfield1=term
Record pages reuse that query string with display_mode=detail. There is no stable list URL.
The suite can expose Z39.50, SRU/SRW, and OAI-PMH. Many installs do not. e-Locus returns 404 for /oai/request. When Identify succeeds, use the OAI fallback and keep dc:type dataset / research data.
Keep: items typed as a dataset or research data. Drop: theses, dissertations, articles, books, and digitized monograph pages. University of Crete e-Locus and Anemi are publication and digitized-document libraries; do not harvest them as dataset catalogs unless a dataset collection is present.
Samvera Hyrax (hyrax)
Blacklight JSON catalog. Work types include GenericWork, Dataset, Etd, Image, FileSet.
GET https://host/catalog.json?f[human_readable_type_sim][]=Dataset&per_page=100&page=1
If that facet is empty, try f[resource_type_sim][]=Dataset or f[has_model_ssim][]=Dataset. FileSets are files, not datasets.
Optional OAI-PMH when the Blacklight OAI plugin is enabled:
GET https://host/catalog/oai?verb=Identify
Then ListSets / ListRecords with the OAI fallback.
Keep: Blacklight works typed Dataset. Drop: FileSets, GenericWork/Etd/Image unless they are the data product.
Samvera (samvera)
Same Blacklight harvest as Hyrax when the UI is Samvera without Hyrax branding.
GET https://host/catalog.json?f[human_readable_type_sim][]=Dataset&per_page=100&page=1
Same optional OAI as Hyrax: /catalog/oai?verb=Identify.
Keep: Dataset works (same grain as Hyrax). Drop: FileSets and publication-only work types.
Islandora (islandora) is Drupal+Fedora: harvest the public JSON:API or Solr only when a dataset content model / collection exists. Prefer Islandora over raw fedora /fcrepo/rest. See Islandora.
Clowder (clowder)
REST API under /api (Swagger UI at /swagger/). Use GET only — Clowder returns 404 to HEAD.
GET https://host/api/status
GET https://host/api/datasets?limit=100
GET https://host/api/datasets/{id}
GET https://host/api/datasets/{id}/metadata.jsonld
/api/status is anonymous and returns version plus dataset/file counts. /api/datasets lists datasets visible to the caller — anonymous on open instances, 401 Not authorized where login is required (then harvest only with an API key, or skip). Collections (/api/collections) group datasets; spaces (/api/spaces) are access-control groupings, not datasets.
Keep: datasets, with their files as resources. Drop: spaces and collections as dataset rows (keep as grouping metadata), files without a parent dataset.
OPUS (opus)
German IRs. The dataset document type is usually researchdata / ResearchData.
GET https://host/oai?verb=ListSets
Look for doc-type:researchdata (spelling varies). Then:
GET https://host/oai?verb=ListRecords&metadataPrefix=oai_dc&set=doc-type:researchdata
Solr UI often supports a doctype facet (doctypefq=researchdata). Thesis-only OPUS hosts have no dataset set — skip them for a data crawl (they can still be valid catalog records).
Keep: doc-type:researchdata / ResearchData OAI or Solr facet. Drop: thesis-only OPUS hosts for a data crawl.
mediaTUM (mediatum)
Generator mediatum - a multimedia content repository. Harvest published documents, media, and research-data records from the public search or collection listing. Drop image derivatives, login-only admin, and the project site mediatum.github.io. One harvest scope per installation. Distinct from OPUS on the same university.
Keep: public documents, media, and research-data records. Drop: derivatives, admin, and the project site.
MyCoRe (mycore)
GET https://host/api/v2/objects
GET https://host/servlets/OAIDataProvider?verb=ListSets
Classification values are local (mir_types, state). Filter to data/Forschungsdaten classes after reading one object and ListSets. Unfiltered /api/v2/objects is the whole IR.
Keep: objects in data/Forschungsdaten classes. Drop: unfiltered /api/v2/objects as the whole IR.
PHAIDRA (phaidra)
GET https://host/api/search/select?q=*:*&rows=0
GET https://host/api/oai?verb=Identify
GET https://host/api/openapi
Add a type constraint once you see stored fields (often cmodel, dc_type, or object_type). Example patterns to try: cmodel:*Dataset*, dc_type:dataset. Drop image/book/thesis cmodels.
Type Dataset Solr select (cmodel:*Dataset* or dc_type:dataset) as rest.
Keep: Dataset cmodels / dc_type:dataset. Drop: image, book, and thesis cmodels.
DiVA Portal (divaportal)
Mixed IR. Publications dominate. Prefer a research-data filter on smash search or OAI.
GET https://host/smash/search.jsf
GET https://www.diva-portal.org/smash/oai?verb=Identify
Type /smash/search.jsf as index. Keep records typed as research data / dataset. Drop articles, theses, and reports. One harvest scope per {org}.diva-portal.org tenant.
Keep: research data / dataset records (smash filter or OAI). Drop: articles, theses, and reports.
WEKO3 (weko3)
Item type IDs are per instance. The registry probe uses type= on /api/records/ — that integer is not portable.
- Open the public search UI or API and list item types.
- Find the id for research data / 研究データ / Dataset.
- Crawl
/api/records/?type=ITEM_TYPE_ID&page=1&size=20(replaceITEM_TYPE_ID).
Without a resolved type id, you will ingest articles and reports. OAI Identify is portable when present; still filter ListRecords to a research-data set or dc:type.
GET https://host/api/records/?page=1&size=20
GET https://host/oai?verb=Identify
Keep: WEKO3 items whose type id is research data / Dataset. Drop: articles and reports (unfiltered /api/records/).
Elsevier Pure (pure)
The public portal lists /en/datasets/ (locale prefix varies: /de/datasets/, /da/datasets/). Publications live under /publications/ and /persons/.
Prefer the datasets channel:
GET https://host/sitemap/datasets.xml
GET https://host/en/datasets/?search=&format=rss
Locale prefix varies (/de/datasets/, /da/datasets/). Type those dataset RSS feeds as rss.
OAI: /ws/oai?verb=Identify then ListSets for a datasets set.
Pure Web Services (/ws/api/datasets) often need an API key. If you get 401, use the public portal/OAI/sitemap. Do not guess keys.
Drop: /publications/, activities, prizes, student theses unless typed as datasets.
Keep: Pure datasets channel (/datasets/ sitemap, RSS, or OAI datasets set).
Esploro (esploro)
Research outputs include datasets as one resource type.
The registry map checks /view/google/siteindex.xml for /dataset/ paths. Use that sitemap when it exists.
Otherwise use the public research search with a datasets facet (UI labels: Dataset, Research data). The SOAP/WADL probe /esplorows/rest/research/simpleSearch is a capability URL, not a full crawl.
Drop: articles, books, conference papers, ETDs in the same index.
GET https://host/view/google/siteindex.xml
Keep: Esploro records with a datasets / research-data facet.
Elsevier Digital Commons (elsevierdigitalcommons)
Collections mix articles and data series. OAI: /do/oai/?verb=ListSets or /oai?verb=ListSets on Elsevier Data Repository hosts. Harvest only sets whose names are data/datasets/statistics — not the whole IR.
Sitemap /sitemap/index can list every series; still skip photograph and journal series.
GET https://host/do/oai/?verb=ListSets
GET https://host/oai?verb=ListSets
Keep: OAI sets named data/datasets/statistics. Drop: photograph and journal series, and the unfiltered IR.
Figshare (figshare)
Institutional Figshare (not every figshare.com article). Item types are numeric.
item_type | Meaning |
|---|---|
| 3 | Dataset — keep |
| 4 | Fileset — keep (collection of files) |
| 9 / 18 | Code / software — not a dataset |
| 6, 8, 5, 7 | Paper, thesis, poster, presentation — drop |
GraphQL/search endpoints vary by tenant. Prefer the institution’s public API or sitemap entries under /articles/dataset/. Do not crawl figshare.com/articles globally.
GET https://host/articles/dataset/
Type /articles/dataset/ as index. Keep: institutional Figshare item_type 3 (dataset) and 4 (fileset). Drop: papers, theses, posters, presentations, and a global figshare.com crawl.
Renku (renku)
One harvest scope per deployment (renkulab.io or an institutional host).
GET https://host/api/data/search/query?q=water
Page the search. Keep: datasets and data connectors. Drop: compute sessions, user profiles, and individual project files.
Haplo (haplo)
Output types include publications and datasets. Use the public catalog/OAI and keep records typed as dataset / research data. Skip grant and HR objects. Skip haplo.com marketing hosts.
Keep: public catalog/OAI records typed as dataset or research data. Drop: grant, HR, and person objects; haplo.com marketing hosts.
GET https://host/oaiprovider?verb=Identify
GET https://host/oaiprovider?verb=ListSets
Worktribe (worktribe)
Public catalog/OAI (/oaiprovider?verb=Identify). Keep dataset / research data. Skip grant/HR objects and worktribe.com marketing.
Keep: public catalog/OAI records typed as dataset / research data. Drop: grant/HR objects and worktribe.com marketing.
GET https://host/oaiprovider?verb=Identify
GET https://host/oaiprovider?verb=ListRecords&metadataPrefix=oai_dc
Omega-PSIR (omegapsir)
CRIS with separate publications vs data modules when configured. Prefer URLs/APIs under a datasets/research-data listing. A global publication search is the wrong crawl.
Keep: records from a datasets / research-data module when configured. Drop: global publication search hits (articles, theses) as datasets.
GET https://host/oai?verb=Identify
VuFind (vufind)
Discovery layer over mixed IRs.
GET https://host/vufind/Search/Results?type=AllFields&filter[]=format%3A"Dataset"
GET https://host/api?openapi
Add a format/type facet (format:Dataset, document_type:dataset) before paging. Type the Dataset Search/Results listing as index. Attach /Search/Results to the catalog link; do not prefix /vufind again when the link already includes that mount. Type /api?openapi as openapi. Keep: facet-filtered dataset records. Drop: unfiltered library-catalog hits.
LibreCat (librecat)
Same facet-first harvest as VuFind when the public UI is LibreCat. Optional OAI Identify is /oai?verb=Identify at the repository origin (not under a /search UI mount).
Keep: facet-filtered dataset / research-data records (same grain as VuFind). Drop: publications and person records.
GET https://host/vufind/Search/Results?type=AllFields&filter[]=format%3A"Dataset"
GET https://host/oai?verb=Identify
Type the Dataset Search/Results listing as index (same grain as VuFind). OAI Identify is oaipmh20.
InstDB (instdb)
FairStack institutional research-data nodes. Harvest the public dataset/API list on the node (/api when present). Skip fairstack.cn marketing and per-file URLs.
Keep: dataset records from the node /api (or documented catalog list).
Drop: fairstack.cn marketing pages and per-file object URLs.
GET https://host/api
META-SHARE (metashare)
Language-resource nodes. Harvest the public resource catalog (corpora, lexica, tools) from the node search or documented export.
GET https://host/
Keep resource records. Drop a single corpus/tool landing page as a crawl seed and META-NET marketing pages. One harvest scope per node. Victoria MetaShare is GeoNetwork — use that recipe instead.
Keep: public resource catalog records (corpora, lexica, tools). Drop: a single corpus/tool landing as a crawl seed and META-NET marketing.
NYU Data Catalog (nyudatacatalog)
Medical-library dataset catalog (schema.org DataCatalog JSON-LD on listing pages). Harvest Dataset objects from JSON-LD or the public search listing. Drop expert/person pages. Drupal JSON:API only if a dataset bundle exists.
Keep: schema.org Dataset objects from JSON-LD or the public listing. Drop: expert/person pages.
DataLad (datalad)
Harvest the published catalog dataset list (catalog.json or the catalog site’s dataset pages), not git-annex keys.
Keep: dataset entries in the published DataLad catalog (catalog.json or equivalent dataset pages).
Drop: git-annex keys, annex object URLs, and Git commit pages.
GET https://host/catalog.json
GIN (gin)
GET https://host/api/v1/repos/search
Keep: public repositories that are datasets. Drop: git objects and private repos. Stop on 401.
HUBzero (hubzero)
Scientific gateway. Harvest public resources typed as datasets/databases. Drop tools, tickets, and login-only groups.
Keep: public resources typed as datasets/databases. Drop: tools, tickets, and login-only groups.
GET https://host/resources?sortby=date
LinkAhead (linkahead)
CaosDB REST (/api/v1/). Query Record types that are datasets/collections. Drop files and properties as extra datasets.
Keep: Record types that are datasets/collections. Drop: files and properties as extra datasets.
Type /api/v1/ as rest.
GET https://host/api/v1/
Fedora (fedora)
Use Fedora LDP /fcrepo/rest (or /rest) only when Fedora is the public catalog. Prefer Hyrax/Islandora/PHAIDRA/Archipelago recipes on the same host.
Keep: LDP containers that represent datasets or data collections when Fedora is the public catalog. Drop: bitstreams as separate datasets when a parent object exists; prefer Hyrax/Islandora/PHAIDRA/Archipelago on the same host.
GET https://host/fcrepo/rest
GET https://host/rest
DABAR (dabar)
Mixed IR (theses, publications, and some research data) on SRCE’s national stack.
GET https://host/oai/?verb=Identify(hub and some tenants use/oai/?verb=Identify; others/oai?verb=Identify).verb=ListSets— keep setSpecs that mean research data / datasets. Ignore the whole-repository set.- If no dataset set exists, harvest
oai_dcand keep records whosedc:typematches the keep list.
Do not crawl dabar.srce.hr/search?ns= as a separate catalog. Prefer the institutional hostname already in the registry. Drop ETDs and journal articles unless typed as datasets.
Keep: OAI sets that mean research data / datasets (or dc:type keep-list). Drop: ETDs and journal articles unless typed as datasets.
OpenScience.si repository (opensciencesi)
Mixed IR. OAI is usually:
GET https://host/oai/oai2.php?verb=Identify
Some tenants use /oai/?verb=Identify. ListSets then keep research-data / dataset sets. Otherwise filter oai_dc with the keep list. Skip the national aggregator www.openscience.si for dataset harvest — use the university tenants.
Keep: research-data / dataset OAI sets (or oai_dc keep-list). Drop: the national aggregator www.openscience.si as a dataset harvest.
Islandora (islandora)
Drupal+Fedora. Harvest Solr/REST with a Dataset content model — not every Drupal node. Prefer Islandora over raw Fedora.
Keep: Islandora objects with a Dataset (or equivalent) content model. Drop: every Drupal node, exhibit pages, and raw Fedora bitstreams.
GET https://host/solr/select?q=RELS_EXT_hasModel_uri_ms:*Dataset*&wt=json&rows=25
GET https://host/jsonapi
Type the Solr Dataset select as rest. Solr wt=json may be served as text/plain.
Optional OAI-PMH Identify (portable paths only; skip host-specific /api/oai2 and /oaiprovider/):
GET https://host/oai/request?verb=Identify
GET https://host/oai2?verb=Identify
GET https://host/oai?verb=Identify
Then ListSets / ListRecords with the OAI fallback.
Archipelago Commons (archipelago)
Drupal Strawberryfield ADOs. Mixed GLAM instances need a Dataset (or accession/isolate) filter; germplasm and culture-collection catalogs may harvest every ADO.
GET https://host/search?f[0]=descriptive_metadata_object_types:Dataset
GET https://host/rss.xml
GET https://host/jsonapi/node/digital_object
GET https://host/api/oai_pmh/oai?verb=Identify
Type Dataset search as index. Keep Dataset, accession, and isolate records. Drop Photograph, Book, Finding Aid, and WebPage exhibits. OAI-PMH /api/oai_pmh/oai?verb=Identify is optional and often restricted. Prefer Archipelago over raw drupal.
Keep: Dataset, accession, and isolate records. Drop: Photograph, Book, Finding Aid, and WebPage exhibits.
CONTENTdm (contentdm)
Only when the site was accepted as a dataset catalog (discovery-scientific.md). /digital/api/collections plus OAI; keep statistical/climate collections, skip photo exhibits.
Keep: collections that are statistical, climate, or research datasets. Drop: photo/manuscript exhibits and individual image records.
GET https://host/digital/api/collections
GET https://host/digital/oai/oai.php?verb=Identify
Type /digital/api/collections as contentdm:collections. OAI-PMH Identify stays oaipmh20.
Omeka S (omekas)
Only when accepted as a dataset catalog. /api/items filtered to Dataset / DataCatalog classes; skip exhibit images.
Keep: /api/items filtered to Dataset / DataCatalog classes.
Drop: exhibit images. Only when the site was accepted as a dataset catalog.
GET https://host/api/items?resource_class_label=Dataset
Type that items URL as omekas:items.
OSF (osf)
Harvest institution or named project catalogs only (https://api.osf.io/v2/). Keep nodes/registrations that are data. Do not crawl all of osf.io. Stop on 401.
Keep: institution or named project nodes/registrations that are data.
Drop: a crawl of all osf.io. Stop on 401.
GET https://api.osf.io/v2/nodes/?filter[parent]=null
Converis (converis)
Clarivate CRIS. Same publication-vs-data problem as Pure: harvest datasets, not publications or persons. Filter exports on software.id = 'converis'.
Prefer a public datasets / research-data listing or OAI setSpec for data. Stop on /ws API keys. Do not page an unfiltered publication search.
Keep: public datasets / research-data listing or OAI data setSpec. Drop: unfiltered publication search, persons, and /ws API keys.
Djehuty (djehuty)
4TU.ResearchData stack. Harvest the public dataset search (Invenio-like resource_type filter when exposed).
Keep: records typed as dataset / research data in the public search. Drop: publications, presentations, and login-only deposit forms.
GET https://host/api/records?q=metadata.resource_type.type:dataset&size=25
GET https://host/v2/articles
Type Invenio-like /api/records as inveniordmapi:records. Some tenants list articles at /v2/articles (rest).
RADAR (radar)
FIZ Karlsruhe research data repositories (RADAR Cloud and RADAR Local). Filter exports on software.id = 'radar'.
GET https://host/radar/api/datasets
GET https://host/oai/OAIHandler?verb=Identify
Already datasets (totalHits in the JSON). Page the API; keep dataset ids/DOIs. Skip a single /radar/de/dataset/ landing page as a seed and the FIZ marketing site. OAI is a fallback. Discovery: discovery-scientific.md.
Keep: RADAR dataset ids/DOIs from /radar/api/datasets or OAI. Drop: a single landing page as a seed and FIZ marketing.
Detection strips /radar/{lang}/home so /radar/api/datasets and /oai/OAIHandler attach at origin.
Redivis (redivis)
Dataset-native SaaS. The public OpenAPI spec does not need a token; listing datasets does. Filter exports on software.id = 'redivis'. Org name is the {org} subdomain (stanford.redivis.com → stanford).
GET https://host/api/v1/openapi.json
GET https://redivis.com/api/v1/organizations/{org}/datasets?maxResults=100
Page with pageToken. Keep dataset.list rows (kind / dataset name). Drop workflows, notebooks, members, and individual tables when a parent dataset exists. A Bearer token with the public scope is required for the list URL; stop on 401/403. Do not crawl redivis.com globally or a single /ORG/dataset-name landing page. Discovery: discovery-scientific.md.
Keep: dataset.list rows. Drop: workflows, notebooks, members, and individual tables when a parent dataset exists.
Yoda (yoda)
Utrecht / SURF research-data vault on iRODS. Filter exports on software.id = 'yoda'.
GET https://host/oai/oai?verb=Identify
Keep: published vault datasets (DataCite DOI landing pages or the public catalog API in endpoints[]). Drop: /research/ collaboration collections, iRODS tickets, and every file in a vault package. Stop on 401.
DLCM (dlcm)
swissuniversities OAIS stack. Filter exports on software.id = 'dlcm'. Prefer endpoints[] OAI-PMH on the access module.
GET https://access.host/oai-info/oai-provider/oai?verb=Identify
GET https://access.host/oai-info/oai-provider/oai?verb=ListRecords&metadataPrefix=oai_dc
Keep deposited datasets and their DOIs. Follow resumption tokens. Drop the Angular UI chrome, WordPress marketing pages (olos.swiss), and login-only OAI. Discovery: discovery-scientific.md.
Keep: deposited datasets and their DOIs (OAI on the access module). Drop: Angular UI chrome, WordPress marketing, and login-only OAI.
DaSCH Service Platform (dsp)
Humanities repository. Filter exports on software.id = 'dsp'. One harvest scope for app.dasch.swiss. Projects listed by the admin API are collections inside that catalog.
GET https://api.dasch.swiss/admin/projects
GET https://repository.dasch.swiss/dpe/oai?verb=Identify
GET https://repository.dasch.swiss/dpe/oai?verb=ListRecords&metadataPrefix=oai_dc
Keep: research-data projects and their OAI records. Drop: dasch.swiss organization pages, DSP-APP chrome, and user accounts. Discovery: discovery-scientific.md.
easydb (easydb)
Programmfabrik easydb 5 / fylr. Filter exports on software.id = 'easydb'. There is usually no public dataset-list API; /api/v1/session is session metadata, not a catalog dump.
Harvest the public object/search UI the catalog link points at (or a documented public search export if present). Keep collection objects that are datasets or media catalog records. Drop login-walled objects and session JSON. Stop on 401. Discovery: discovery-scientific.md.
Keep: collection objects that are datasets or media catalog records. Drop: login-walled objects and /api/v1/session JSON.
LabKey Server (labkey)
GET https://host/login/begin.view
Keep studies / published folders (Panorama Public libraries, Open Research Portal projects). Drop assay run rows and a single begin.view folder as a seed. Stop on 401.
Keep: studies / published folders. Drop: assay run rows and a single begin.view folder as a seed.
BioUML (biouml)
GET https://host/ (database home: table listing)
Keep databases / collections (GTRD ChIP-seq experiments, HOCOMOCO motif models, EpiFactors entries). Drop a single table row, motif model, or track download as a seed. Grain is the collection/table, not the record.
Keep: databases / collections. Drop: a single table row, motif model, or track download as a seed.
Synapse (synapse)
GET https://repo-prod.prod.sagebase.org/repo/v1/entity/synNNNN/children
Keep projects and tables/files that are cited as datasets. Drop every child file under a project when a parent dataset entity exists. Prefer the catalog link origin and endpoints[]. Stop on 401.
Keep: projects and tables/files cited as datasets. Drop: every child file under a project when a parent dataset entity exists.
Gen3 (gen3)
Data-commons portal. Prefer endpoints[] (Indexd, DRS, GraphQL). Defaults:
GET https://host/_status
GET https://host/index/ga4gh/drs/v1/service-info
Type DRS service-info as ga4gh:drs.
Keep studies / projects from the public GraphQL or portal catalog. Drop individual DRS objects, files, and Fence /user/login as harvest seeds. Stop on 401/403. Do not harvest NCI GDC/PDC/IDC under this recipe.
Keep: studies / projects from public GraphQL or the portal catalog. Drop: individual DRS objects, files, and Fence login as harvest seeds.
XNAT (xnat)
GET https://host/data/projects
GET https://host/xnat/data/projects
Keep projects (and experiment collections when the user asked). Drop individual imaging sessions and DICOM files when a parent project exists. Stop on 401.
Keep: projects (and experiment collections when asked). Drop: individual imaging sessions and DICOM files when a parent project exists.
Type /data/projects as xnat:projects.
Shanoir (shanoir)
GET https://host/shanoir-ng/welcome
Keep studies / datasets listed on the public instance. Drop individual imaging examinations, DICOM files, and the Inria project homepage. Stop on 401/403. One harvest scope per public Shanoir instance (Neurinfo, OFSEP, …).
Keep: public studies / datasets. Drop: individual imaging examinations, DICOM files, and the Inria project homepage.
LORIS (loris)
GET https://host/
Keep published instruments / imaging collections / datasets on the public portal. Drop candidate pages, visit forms, and demo.loris.ca. Stop on 401/403. One harvest scope per public LORIS instance.
Keep: published instruments / imaging collections / datasets. Drop: candidate pages, visit forms, and demo.loris.ca.
OMERO (omero)
GET https://host/api/v0/m/projects/
GET https://host/webclient/
Keep projects / screens / studies (IDR annotations). Drop individual images and wells. Some public archives return 404 on /api/v0/m/ — fall back to the documented webclient catalog. Stop on 401.
Type /api/v0/m/projects/ as omero:projects. Fall back /webclient/ as omero:webclient.
Keep: projects / screens / studies. Drop: individual images and wells.
Kadi4Mat (kadi4mat)
GET https://host/api/records
GET https://host/api/collections
Keep records and collections. Drop individual file blobs when a parent record exists. Stop on 401.
Type /api/records as kadi4mat:records and /api/collections as kadi4mat:collections.
Keep: records and collections. Drop: individual file blobs when a parent record exists.
TR32DB (tr32db)
/site/index.php Cologne CRC databases. Harvest the public metadata / dataset search if unauthenticated. Do not scrape file blobs or require project login. One harvest scope per CRC database (TR32, CRC1211, TRR228). Distinct from CRC806DB.
GET https://host/site/index.php
Keep: public metadata / dataset search. Drop: file blobs and project-login walls.
e!DAL (edal)
GET https://host/
Keep versioned DOI datasets. Drop a single landing page as a crawl seed. Prefer the documented e!DAL API in endpoints[]. Stop on 401.
Keep: versioned DOI datasets. Drop: a single landing page as a crawl seed.
NOMAD (nomad)
GET https://host/prod/v1/api/v1/info
GET https://host/prod/v1/api/v1/entries
Keep uploads / entries that are published datasets. Drop individual calculation files and parser logs. One Oasis or the central archive = one harvest scope. Stop on 401.
Type /prod/v1/api/v1/info and /prod/v1/api/v1/entries as rest.
Keep: published uploads / entries. Drop: individual calculation files and parser logs.
High-Throughput Toolkit (httk)
OPTIMADE where the site publishes it. The base is /optimade/{db}/v1/ on the catalog host, or https://optimade.{host}/v1/ on a sibling host.
GET https://host/optimade/{db}/v1/info
GET https://host/optimade/{db}/v1/structures
GET https://optimade.{host}/v1/info
GET https://optimade.{host}/v1/structures
Keep OPTIMADE structures (and related entry types such as references) as datasets. Drop raw calculation inputs, charge-density blobs, and the httk.org documentation site. One public database is one harvest scope.
Type the OPTIMADE /v1/info URL as rest. Follow page_limit and page_offset from the response meta links.
Keep: OPTIMADE structures and related entry types. Drop: raw calculation files and the toolkit documentation site.
httkweb catalogs with no OPTIMADE base (ADAQ, the hard-coating alloys database) are harvested from the public materials or defect search, not from calculation files.
Pathogens Portal Node (nodepathogensportal)
Toolbox nodes publish a dataset listing at /datasets/. The Swedish original lists datasets from its own data section on www.pathogens.se.
GET https://host/datasets/
Keep listed datasets and data highlights. Drop dashboards, news, events, training pages, and the EMBL-EBI central portal. One national node is one harvest scope.
Keep: listed datasets and data highlights. Drop: dashboards, news, events, and training pages.
dLibra (dlibra)
Polish digital library. Most installs expose Identify at /dlibra/oai-pmh-repository.xml?verb=Identify (catalog links that already end in /dlibra are stripped before that path is attached).
GET https://host/dlibra/oai-pmh-repository.xml?verb=Identify
Keep: OAI-PMH records with a dataset / dane set or dc:type filter (harvest-protocols.md). Drop: manuscript/photo libraries that were never accepted as dataset catalogs, and unfiltered ListRecords.
ARPHA Platform (arpha)
Pensoft publishing platform tenants (preprint server, journal hosts). One harvest scope per tenant root. OAI-PMH 2.0 with oai_dc and mods at /oai:
GET https://preprints.arphahub.com/oai?verb=ListRecords&metadataPrefix=oai_dc
Keep: preprints and data-bearing article records (occurrence/checklist/treatment data articles with DOIs). Drop: journal news, issue tables of contents, and full HTML article bodies; Pensoft journal data portals are gbifplatform — harvest those via harvest-biodiversity.md, not ARPHA OAI.
Dataset-native platforms (short)
Little publication noise. Still skip non-dataset objects.
| Platform | List | Notes |
|---|---|---|
| Dataverse | above | type=dataset only |
SciCat (scicat) | harvest-earthdata.md | Facility datasets; stop on 401 |
RADAR (radar) | above | Already datasets; skip marketing and single landings |
Yoda (yoda) | above | Published datasets only; skip the authenticated vault |
LabKey (labkey) | above | Studies / published folders |
Synapse (synapse) | above | Projects and dataset entities, not every file |
XNAT (xnat) | above | Projects, not sessions |
Shanoir (shanoir) | above | Studies, not imaging sessions |
LORIS (loris) | above | Published collections, not candidate visits |
OMERO (omero) | above | Projects/screens, not images |
Kadi4Mat (kadi4mat) | above | Records and collections |
e!DAL (edal) | above | DOI datasets |
NOMAD (nomad) | above | Published entries/uploads |
High-Throughput Toolkit (httk) | above | OPTIMADE structures; httkweb search when no API |
Pathogens Portal Node (nodepathogensportal) | above | /datasets/ listings |
Domain stacks (IPT, THREDDS, Breedbase, ESGF, …): harvest-scientific-domain.md. Omeka S and CONTENTdm: sections above.
Pagination checklist
- Read
total/nHits/page.totalPages/ OAIresumptionTokenfrom the first response. - Cap page size; do not request
size=10000on Solr-backed IRs. - Deduplicate on DOI, handle, or native id plus catalog
uid(harvest-identifiers.md). Emit output records. - Re-run with
from=(OAI) orupdatedsort for incremental harvests when the API supports it (harvest-incremental.md).
FLAT (flat)
Start at the registered repository and use the advertised OAI-PMH endpoint. For Lund:
GET https://host/flat/oai2?verb=Identify
Typical catalog links already end in /flat/. Cleanup strips that mount so Identify is not doubled (/flat/flat/).
Enumerate metadata formats and sets before ListRecords; prefer CMDI when offered, otherwise a supported descriptive format. Follow resumption tokens and preserve repository identifiers and collection membership.
Keep deposited language-resource collections, corpora and dataset metadata. Exclude navigation nodes, user profiles and individual media files as independent datasets. Public metadata does not imply that restricted audio or video is downloadable. The FLAT source documentation describes its Fedora/Islandora components; administrative Fedora endpoints are not public harvesting seeds.
Keep: deposited language-resource collections, corpora, and dataset metadata. Drop: navigation nodes, user profiles, and individual media files as independent datasets.
openEQUELLA (openequella)
Start at the registered institution's repository and its advertised OAI-PMH or REST search
interface. For RADAR: GET https://radar.brookes.ac.uk/radar/oai?verb=Identify.
Enumerate metadata formats and sets, then use ListRecords with resumption tokens.
Apereo's product description
documents OAI and REST interfaces; authentication and route prefixes vary by institution.
Keep dataset/research-resource metadata and attached resource links. Exclude teaching objects, publication-only records and administrative collections when harvesting research datasets. An item can have several versions and files; preserve its stable identifier and version without counting every attached file as a new dataset.
Keep: dataset / research-resource metadata and attached resource links. Drop: teaching objects, publication-only records, and every attached file as a new dataset.
Aubrey (aubrey)
Start at the registered collection, such as
GET https://digital.library.unt.edu/explore/collections/UNTDRD/, and follow its API link.
Official API guidance documents collection-scoped
interfaces and OAI-PMH formats untl and oai_dc. Resolve the exact collection scope from
that help page instead of harvesting all historical materials in the library.
Keep deposited datasets and their ARK identifiers, collection relationships and resource links. Follow OAI resumption tokens; exclude page images, IIIF tiles, navigation pages and non-data historical collections from dataset output. The metadata service is publicly documented; resource reuse rights still vary by item.
Keep: deposited datasets and their ARK identifiers. Drop: page images, IIIF tiles, navigation pages, and non-data historical collections.
Dialnet CRIS (dialnetcris)
Start at the registered institutional portal. La Rioja exposes
GET https://investigacion.unirioja.es/oai/openaire?verb=Identify.
Enumerate the actual metadata formats and available sets before harvesting records.
The provider's product page
describes an additional REST export service; do not assume that this service is public on
every tenant or invent its URL.
Keep explicitly typed datasets and their metadata/resource links. Drop researcher profiles, projects, indicators and publication-only entries from a dataset harvest. A CRIS record may link to a deposit in another repository: preserve that relationship instead of counting the same dataset twice. The presence of a CRIS platform does not prove that every tenant contains datasets; return an empty dataset result if the available records are only publications.
Keep: explicitly typed datasets and their metadata/resource links. Drop: researcher profiles, projects, indicators, and publication-only entries.
ReDBox (redbox)
Institutional ReDBox / RDMP portals. Prefer the tenant's public list or OAI-PMH endpoint used
to feed Research Data Australia. Paths often look like /default/rdmp/.
GET https://host/default/rdmp/
Keep: public research data collection / dataset metadata records and their download or landing URLs. Drop: research data management plans (RDMP forms), person/Mint lookup records, login-walled drafts, and Research Data Australia itself (national aggregator, not a ReDBox tenant).
Craft CMS (craftcms)
Filter exports on software.id = 'craftcms'. One harvest scope per institution.
No standard data API; catalog pages are CMS-published HTML. Keep the public data catalog / download pages and the dataset records or files they list. Grain is the dataset or data collection, not each news or outreach page. Drop site chrome (news, people, events) and any embedded third-party apps as separate records. Stop on 401/403.
Keep: public data catalog pages and listed datasets/files. Drop: news, people, events chrome.
SBS Digital Collection (sbsdigitalcollection)
University digital-collection installs from Simply Bright System. Items mix theses, articles, books, archives, and research outputs in one index.
Collection search is HTML, not a public dataset JSON API. Paths look like /Search/index/{collection} or /frontend/Search/index/{collection}. DigiVerse robots.txt disallows /api/ and those paths return 404. Do not invent an API URL.
Keep: items in a collection whose scope is research data or datasets. Drop: theses, dissertations, articles, books, exam papers, newspapers, archival scans, and each file page inside an item.
ScienceDB (sciencedb)
Generalist repository at https://scidb.cn. No public bulk JSON API confirmed; dataset discovery is via the site search HTML, and each dataset resolves through its DOI (10.57760/sciencedb.*) to a landing page with metadata and file downloads.
GET https://www.scidb.cn/en/search
Keep: dataset landing pages (one DOI = one dataset). Drop: news, help/about pages, journal partner pages, and per-file download URLs as separate records.