Harvesting biodiversity and genomics catalogs
IPT, Symbiota, Biodiv, BIMS, Living Atlases, Ensembl, PlutoF, InterMine, JGI, cBioPortal, and BirdMap Africa publish datasets, collections, studies, or genome databases. Occurrence rows, gene records, mutation tables, and map clicks are the wrong grain.
Overview: harvest.md. Finding portals: discovery-scientific.md. GET only. Stop on 401/403. Prefer endpoints[].
What to keep
| Keep | Drop |
|---|---|
| IPT Darwin Core archive | Occurrence rows inside the archive |
Symbiota published dataset (RSS) or collection (collid) if asked | Images, checklists, single occurrences |
| Biodiv dataset, datatable, or document collection | Species pages, single observations, user profiles |
| BIMS source reference or spatial layer | Occurrence rows, taxon pages, user accounts |
| ALA collection / data resource | /ws/occurrences/search hits |
GBIF dataset (api.gbif.org) | Occurrence search; publisher orgs as datasets |
| Ensembl species / genome database | Every gene, variation, or REST ping |
| PlutoF dataset / DOI record | Occurrences, sequences, UNITE taxon pages |
| InterMine experiment / dataset list | Gene report pages |
| JGI genome / transcriptome project | Gene pages, BLAST hits, IMG/GOLD |
| cBioPortal study | Mutation/CNA rows, patient samples |
| GRIN-Global accession catalog export | Individual accession HTML pages |
| BirdMap Africa project catalog / API species lists | Survey cards, pentad clicks, BirdLasser |
| GeoNature-atlas species / taxon sheets | Individual observations, TaxHub media files |
| iDigBio datasets / collections | Occurrence search hits (use IPT for Darwin Core archives) |
| iNaturalist projects / export datasets | Per-observation API rows |
GBIF IPT (ipt)
GET https://host/inventory/dataset
GET https://host/rss.do
GET https://host/dcat
Each inventory/RSS item is one dataset. Prefer the IPT root on the catalog link. Skip harvesting gbif.org when you only needed publisher IPTs already in the registry.
Keep: IPT Darwin Core archives (inventory/RSS/DCAT). Drop: occurrence rows inside the archive.
Symbiota (symbiota)
Filter exports on software.id = 'symbiota'.
GET https://host/collections/index.php
GET https://host/collections/datasets/rsshandler.php
Keep Darwin Core datasets (RSS). Collection-level harvest only if the user wants one record per collid. One portal = one harvest scope. Login-only: stop. Directory: symbiota.org/symbiota-portals.
Keep: published Darwin Core datasets (RSS). Drop: images, checklists, and single occurrences.
Biodiv (biodiv)
Filter exports on software.id = 'biodiv'. Public list pages:
GET https://host/dataset/list
GET https://host/datatable/list
GET https://host/document/list
Keep: published datasets, datatables, and document collections. Drop: species pages, single observations, and user profiles. One portal = one harvest scope. Stop on 401/403.
BIMS (bims)
Filter exports on software.id = 'bims'. Portal list: bims.kartoza.com.
GET https://host/source-references/
GET https://host/api/module-summary/
GET https://host/api/layer/
Keep: source references and published spatial layers. Drop: occurrence rows, taxon pages, and user accounts. One portal = one harvest scope. Stop on 401/403.
Atlas of Living Australia (ala)
GET https://host/ws/registry/collections
Harvest collections (data resources). Species autocomplete and occurrence search are not dataset lists. Same pattern on other Living Atlases.
Keep: ALA collections / data resources. Drop: /ws/occurrences/search hits.
BirdMap Africa (birdmap)
African Bird Atlas country portals ({project}.birdmap.africa). Harvest the project catalog and public API (https://api.birdmap.africa/{project}/v2/ when listed in endpoints[]). Keep pentad coverage / species-list datasets. Drop individual survey cards, pentad map clicks, and BirdLasser app traffic. One country project = one harvest scope.
Keep: project catalog / API species-list datasets. Drop: survey cards, pentad clicks, and BirdLasser.
GET https://api.birdmap.africa/sabap2/v2/
GBIF platform (gbifplatform)
GET https://api.gbif.org/v1/dataset?limit=100&offset=0
Use this only when the registry record is GBIF (or a national GBIF portal whose API is GBIF). Filter with publishingCountry / publishingOrg when the catalog is a country node. Prefer harvesting member IPTs from this registry for publisher-level ids. Do not page /v1/occurrence/search.
Keep: GBIF datasets. Drop: occurrence search and publisher orgs as datasets.
Ensembl (ensembl)
GET https://host/info/ping
GET https://host/info/species
REST base is often https://rest.ensembl.org or https://host/rest. Harvest species / assembly databases on that taxon portal (Fungi, Protists, Metazoa, …). Do not harvest every gene. Do not clone ensembl.org if you only needed an existing registry row.
Keep: Ensembl species / genome databases. Drop: every gene, variation, or REST ping.
SEANOE / IFREMER Catalog (ifremercatalog)
Marine-science dataset repository (seanoe.org). Prefer endpoints[] (OAI-PMH Identify is already recorded).
GET https://www.seanoe.org/oai/OAIHandler?verb=Identify
GET https://www.seanoe.org/oai/OAIHandler?verb=ListRecords&metadataPrefix=oai_dc
Keep datasets (DataCite/OAI type Dataset). Drop publications mixed into the same OAI set without a type filter. Do not harvest every NetCDF file under a parent dataset. Skip cloning seanoe.org if you only needed the existing registry row.
Keep: SEANOE datasets. Drop: publications without a type filter and every NetCDF file under a parent dataset.
Related scientific IDs
software.id | Harvest | Skip |
|---|---|---|
ipt, symbiota, ala | sections above | occurrences |
gbifplatform | GBIF dataset API with a country/org filter | occurrence API |
ensembl | species list on that portal | gene endpoints |
breedbase | BrAPI trials/studies | plots, samples, marker calls |
tripal | analyses / downloadable datasets | gene pages, BLAST hits |
veupathdb | experiment / isolate / genome datasets | gene records, strategy rows |
ifremercatalog | SEANOE OAI/dataset list | publication mix; file-level NetCDF |
plutof | PlutoF dataset/DOI API | occurrences, UNITE sequence pages |
intermine | experiments / dataset lists | gene reports, /begin.do crawls |
jgi | Genome Portal projects | gene pages, IMG, GOLD, data.jgi.doe.gov |
cbioportal | /api/studies | mutation/CNA rows |
checklistbank | ChecklistBank dataset API | taxon / name-usage pages |
gringlobal | accession catalog exports | each accession HTML page |
birdmap | country portal / api.birdmap.africa/{project}/v2/ | survey cards, pentad clicks |
geonature | GeoNature-atlas species sheets | observations; GeoNature back-office |
idigbio | Portal datasets/collections; IPT is ipt | occurrence search hits |
inaturalist | hub projects / exports | per-observation rows |
Institutional IRs that also hold Darwin Core: use harvest-scientific.md type filters, not occurrence APIs. Domain harvest recipes: harvest-scientific-domain.md.