Skip to main content

Incremental harvesting

A full harvest lists every public dataset. An incremental harvest lists only records that appeared or changed since the last successful run. Use this page after you have a first snapshot from a type or protocol guide.

Store checkpoints outside this repository (your index, object store, or harvest DB). Do not write dataset YAML here.

What to persist per catalog​

FieldWhy
Catalog uid / idJoin key back to this registry
Native dataset id (and DOI/handle if present)Dedup — harvest-identifiers.md
software.id and the filter you usedReplay
Last successful timestamp or tokenNext incremental
Skip countsSpot a broken filter

Do not invent cdi######## ids for datasets.

Prefer server-side “since”​

Protocol / APIIncremental hook
OAI-PMHfrom= / until= (ISO date) on ListRecords
CKANfq=metadata_modified:[SINCE TO *] or sort metadata_modified desc and stop
DSpace 7lastModified / discover query with a date facet
InvenioRDMq=updated:>=SINCE (ISO)
Dataversesort=dateSort / fq=dateSort:[SINCE TO *]
CSWModified / RevisionDate PropertyIsGreaterThan in GetRecords
STACdatetime=SINCE/.. on /search (collections still preferred for grain)
Socrataupdated_at on /api/catalog/v1
OpenDataSoftmodified on explore API
ERDDAPCompare info/index.json datasetID set; some servers expose dataTimestamp
ArcGIS Hubmodified on search
PxWebNo standard “since” — re-walk tables; diff table ids
SDMX dataflow listRe-list /dataflow; diff ids. Do not incremental-page observation cubes
World Bank / GHO indicatorsRe-list indicator APIs; diff indicator ids
openEORe-list /collections; diff collection ids. Do not incremental-page /jobs
Breedbase BrAPIRe-list /brapi/v2/studies (and trials); diff study ids
ESGF esg-searchfrom/to on Solr when documented; else re-query and diff dataset_id
PxStat ReadCollectionRe-list matrices; diff table codes. No standard since
FENIX groupsanddomainsRe-list domains; diff domain codes. Do not incremental-page observation cubes
Sentinel Hub STACRe-list /collections; diff collection ids. Do not incremental-page items
ResourceContractsRe-list /contract/resources; diff contract ids
LabKey published foldersRe-list studies/folders; no standard since
SynapseRe-list project children; diff entity ids
XNAT /data/projectsRe-list projects; diff project ids
OMERO /api/v0/m/projects/Re-list projects/screens; diff ids
Kadi4Mat /api/recordsRe-list records; diff record ids
e!DAL DOI catalogRe-list DOI datasets; diff dataset ids
NOMAD entriesRe-list /prod/v1/api/v1/entries; diff entry ids
InterMineRe-list template/dataset queries; diff ids
GRIN-GlobalRe-list accession catalog exports; diff accession ids
PlutoF /v1/Re-list datasets; diff DOI/ids. Do not incremental-page occurrences
cBioPortal /api/studiesRe-list studies; diff study ids
ESA TAP tablesRe-list TAP tables; do not incremental-page observations
ODWeb /odweb/Re-list dataset pages; no standard since
IMF NSDPRe-list SDMX category links on the country page

If the API has no date filter, harvest identifiers only (cheap list), then GET metadata for ids not in your store. Do not re-download every observation cube.

Tokens and paging​

  • OAI resumptionToken is for one ListRecords session. Do not save it across days; save from= instead.
  • STAC / OGC API: follow rel=next. Do not invent page numbers.
  • CKAN start is an offset; if the catalog mutates mid-crawl, prefer metadata_modified sort.
  • Honor Retry-After and back off on 429. Cap page size (10–100).

First run vs later runs​

  1. First run: apply the dataset-type filter from the platform guide, page to completion, store ids + checkpoint time (use the server’s clock from Date / OAI responseDate when possible).
  2. Later runs: same filter plus from= / updated>=. Union new ids; update changed ids; do not delete missing ids unless the user asked for a tombstone pass.
  3. Tombstones (optional): a rare full harvest to mark disappeared datasets. IRs often keep withdrawn records as unavailable — treat that as a status, not a new dataset.

Failures​

  • 401 / 403: stop. Do not rotate keys.
  • Empty incremental: inspect one unfiltered sample; the clock format may be wrong (YYYY-MM-DD vs full ISO).
  • Huge “everything changed”: the server may ignore from=. Fall back to a full harvest once, then fix the filter.