Skip to main content

Harvesting biodiversity and genomics catalogs

IPT, Symbiota, Biodiv, BIMS, Living Atlases, Ensembl, PlutoF, InterMine, JGI, cBioPortal, and BirdMap Africa publish datasets, collections, studies, or genome databases. Occurrence rows, gene records, mutation tables, and map clicks are the wrong grain.

Overview: harvest.md. Finding portals: discovery-scientific.md. GET only. Stop on 401/403. Prefer endpoints[].

What to keep​

KeepDrop
IPT Darwin Core archiveOccurrence rows inside the archive
Symbiota published dataset (RSS) or collection (collid) if askedImages, checklists, single occurrences
Biodiv dataset, datatable, or document collectionSpecies pages, single observations, user profiles
BIMS source reference or spatial layerOccurrence rows, taxon pages, user accounts
ALA collection / data resource/ws/occurrences/search hits
GBIF dataset (api.gbif.org)Occurrence search; publisher orgs as datasets
Ensembl species / genome databaseEvery gene, variation, or REST ping
PlutoF dataset / DOI recordOccurrences, sequences, UNITE taxon pages
InterMine experiment / dataset listGene report pages
JGI genome / transcriptome projectGene pages, BLAST hits, IMG/GOLD
cBioPortal studyMutation/CNA rows, patient samples
GRIN-Global accession catalog exportIndividual accession HTML pages
BirdMap Africa project catalog / API species listsSurvey cards, pentad clicks, BirdLasser
GeoNature-atlas species / taxon sheetsIndividual observations, TaxHub media files
iDigBio datasets / collectionsOccurrence search hits (use IPT for Darwin Core archives)
iNaturalist projects / export datasetsPer-observation API rows

GBIF IPT (ipt)​

GET https://host/inventory/dataset
GET https://host/rss.do
GET https://host/dcat

Each inventory/RSS item is one dataset. Prefer the IPT root on the catalog link. Skip harvesting gbif.org when you only needed publisher IPTs already in the registry.

Keep: IPT Darwin Core archives (inventory/RSS/DCAT). Drop: occurrence rows inside the archive.

Symbiota (symbiota)​

Filter exports on software.id = 'symbiota'.

GET https://host/collections/index.php
GET https://host/collections/datasets/rsshandler.php

Keep Darwin Core datasets (RSS). Collection-level harvest only if the user wants one record per collid. One portal = one harvest scope. Login-only: stop. Directory: symbiota.org/symbiota-portals.

Keep: published Darwin Core datasets (RSS). Drop: images, checklists, and single occurrences.

Biodiv (biodiv)​

Filter exports on software.id = 'biodiv'. Public list pages:

GET https://host/dataset/list
GET https://host/datatable/list
GET https://host/document/list

Keep: published datasets, datatables, and document collections. Drop: species pages, single observations, and user profiles. One portal = one harvest scope. Stop on 401/403.

BIMS (bims)​

Filter exports on software.id = 'bims'. Portal list: bims.kartoza.com.

GET https://host/source-references/
GET https://host/api/module-summary/
GET https://host/api/layer/

Keep: source references and published spatial layers. Drop: occurrence rows, taxon pages, and user accounts. One portal = one harvest scope. Stop on 401/403.

Atlas of Living Australia (ala)​

GET https://host/ws/registry/collections

Harvest collections (data resources). Species autocomplete and occurrence search are not dataset lists. Same pattern on other Living Atlases.

Keep: ALA collections / data resources. Drop: /ws/occurrences/search hits.

BirdMap Africa (birdmap)​

African Bird Atlas country portals ({project}.birdmap.africa). Harvest the project catalog and public API (https://api.birdmap.africa/{project}/v2/ when listed in endpoints[]). Keep pentad coverage / species-list datasets. Drop individual survey cards, pentad map clicks, and BirdLasser app traffic. One country project = one harvest scope.

Keep: project catalog / API species-list datasets. Drop: survey cards, pentad clicks, and BirdLasser.

GET https://api.birdmap.africa/sabap2/v2/

GBIF platform (gbifplatform)​

GET https://api.gbif.org/v1/dataset?limit=100&offset=0

Use this only when the registry record is GBIF (or a national GBIF portal whose API is GBIF). Filter with publishingCountry / publishingOrg when the catalog is a country node. Prefer harvesting member IPTs from this registry for publisher-level ids. Do not page /v1/occurrence/search.

Keep: GBIF datasets. Drop: occurrence search and publisher orgs as datasets.

Ensembl (ensembl)​

GET https://host/info/ping
GET https://host/info/species

REST base is often https://rest.ensembl.org or https://host/rest. Harvest species / assembly databases on that taxon portal (Fungi, Protists, Metazoa, …). Do not harvest every gene. Do not clone ensembl.org if you only needed an existing registry row.

Keep: Ensembl species / genome databases. Drop: every gene, variation, or REST ping.

SEANOE / IFREMER Catalog (ifremercatalog)​

Marine-science dataset repository (seanoe.org). Prefer endpoints[] (OAI-PMH Identify is already recorded).

GET https://www.seanoe.org/oai/OAIHandler?verb=Identify
GET https://www.seanoe.org/oai/OAIHandler?verb=ListRecords&metadataPrefix=oai_dc

Keep datasets (DataCite/OAI type Dataset). Drop publications mixed into the same OAI set without a type filter. Do not harvest every NetCDF file under a parent dataset. Skip cloning seanoe.org if you only needed the existing registry row.

Keep: SEANOE datasets. Drop: publications without a type filter and every NetCDF file under a parent dataset.

software.idHarvestSkip
ipt, symbiota, alasections aboveoccurrences
gbifplatformGBIF dataset API with a country/org filteroccurrence API
ensemblspecies list on that portalgene endpoints
breedbaseBrAPI trials/studiesplots, samples, marker calls
tripalanalyses / downloadable datasetsgene pages, BLAST hits
veupathdbexperiment / isolate / genome datasetsgene records, strategy rows
ifremercatalogSEANOE OAI/dataset listpublication mix; file-level NetCDF
plutofPlutoF dataset/DOI APIoccurrences, UNITE sequence pages
intermineexperiments / dataset listsgene reports, /begin.do crawls
jgiGenome Portal projectsgene pages, IMG, GOLD, data.jgi.doe.gov
cbioportal/api/studiesmutation/CNA rows
checklistbankChecklistBank dataset APItaxon / name-usage pages
gringlobalaccession catalog exportseach accession HTML page
birdmapcountry portal / api.birdmap.africa/{project}/v2/survey cards, pentad clicks
geonatureGeoNature-atlas species sheetsobservations; GeoNature back-office
idigbioPortal datasets/collections; IPT is iptoccurrence search hits
inaturalisthub projects / exportsper-observation rows

Institutional IRs that also hold Darwin Core: use harvest-scientific.md type filters, not occurrence APIs. Domain harvest recipes: harvest-scientific-domain.md.