Skip to main content

Dataset identifiers from a harvest

This registry’s uid (cdi########) identifies catalogs. Harvested datasets need their own ids, always paired with the catalog uid. Incremental crawls: harvest-incremental.md. Protocol grain: harvest-protocols.md.

Do not mint cdi######## or temp######## for datasets. Do not write dataset YAML into data/entities/.

Store this tuple​

FieldRule
Catalog uidFrom this registry (DuckDB/Parquet)
Catalog id / linkReplay and debug
software.idWhich recipe you used
Native idThe catalog’s primary key (required)
Persistent idDOI, handle, or ARK when present (optional but preferred for dedup)
Landing URLCanonical dataset page, not a session or tile URL
Type filterThe query/set you applied

Prefer persistent ids when present​

Order when several exist:

  1. DOI (10.…) — Dataverse persistentId, DataCite, Figshare, many IRs
  2. Handle (hdl:… or https://hdl.handle.net/…) — DSpace, some EPrints
  3. ARK / PURL documented by the catalog
  4. Software-native id (CKAN id UUID, Invenio record id, ERDDAP datasetID, STAC collection id, CSW fileIdentifier, IPT dataset UUID, PxWeb table path)

Use the package/dataset id, not a file/resource/distribution id. ipums, dhis2, openaire, yoda, and radar are published software.id values — filter exports on those ids.

Native ids by platform (typical)​

software.idNative idNot an id
ckan / dkan / datapressid (UUID) and/or nameResource id
ekanDrupal dataset UUID (/jsonapi/dataset/dataset)Resource UUID
opendatasoftdataset_idAttachment filename
socrataid (four-four)Column id
dataversepersistentId (DOI)File id
dspace / dspacecrisHandle or UUIDBitstream id
inveniordmRecord id / DOIFile key
geonetworkISO fileIdentifierThumbnail URL
stacserverCollection idItem id (unless item grain)
openeoCollection idProcess id, job id
qgisserver / mapserverWMS layer nameGetMap URL
mviewerConfig layer idTile URL
isogeoMetadata record idWorkgroup id
arcgisserverService URL + nameExtent query
erddapdatasetIDTable row
iptDataset UUID / keyOccurrence id
pxwebTable path (type: t)Folder path
pxstatMatrix / table codeWidget embed, demo tables
fairdatapointDataset IRIDistribution IRI
ipumsSeries / sample idExtract job id
dhis2dataSet / indicator idorgUnit, analytics cell
openaireGraph dataset product id / DOIPublication id
yodaVault dataset DOIiRODS path in /research/
radarRADAR dataset id / DOILanding-page URL only
geocortexEssentials site idHtml5Viewer tile URL
vertigisstudiowebTenant host / ?app= GUIDExtra ?app= GUID on the same tenant; Html5Viewer tiles
mapgisigserverMap document / service name/igs/manager
hygmapgisaplicacion= + layer nameMap tile URL / mapa.jsp session
breedbasestudy / trial idplot or marker-call id
esgfdataset_id / master_idfile id
symbiotadataset RSS id / collidoccurrence id
fenixFENIX domain / dataset codeObservation cube cell
tabnet.def table nameCGI session / TabWin .TAB
sparkmapMap layer / assessment idSaved user map
g3wsuitegroup/project + WMS layer name/admin path
sentinelhubSTAC collection idProcess job id, granule/item
geusmapmapname + WMS layer nameMap tile URL
resourcecontractsContract idPer-clause annotation
gxopendataTenant dataset idApply-gateway request
rdfrepositoryLicense / workspace idLogin-only company filing
converisDataset / research-data idPublication or person id
landfolioPublic cadastre layer / license idArcGIS REST service already harvested as arcgisserver
labkeyStudy / published folder idAssay run id
synapseProject or dataset entity id (syn########)Child file entity under a dataset
xnatProject idImaging session / DICOM file
omeroProject / screen / study idImage or well id
kadi4matRecord or collection idFile blob id
edalDOI dataset idSingle landing-page URL as a seed
nomadEntry / upload idCalculation file or parser log
intermineExperiment / list / template-result idGene report
gringlobalAccession catalog export idIndividual accession HTML page
plutofDataset / DOI idOccurrence, sequence, UNITE taxon
jgiGenome / transcriptome project idGene page, BLAST hit, IMG/GOLD
cbioportalStudy idMutation / CNA / patient sample row
esasciencearchiveTAP table / observation-catalog idFITS file or cutout
odweb/odweb/ dataset idParent government homepage
imfnsdpSDMX category / series code on the country pageDSBB directory, Knoema wrapper

Normalize DOI to 10.prefix/suffix (lowercase). Strip https://doi.org/ and doi:. Handles: keep the handle string, not only the UI URL.

Deduplication​

  • Inside one catalog: native id (plus version if the API versions datasets).
  • Across catalogs: DOI/handle first. The same Dataverse study harvested from DataONE and from the MN should collapse if you asked for a global index.
  • Aggregators: prefer the source catalog uid when Idra/DCAT names a publisher URL that is already in this registry (harvest-other.md).

A landing URL with query tokens, /latest/, or map bbox is a locator, not a durable id.

Versions and replacements​

Keep version / metadata_modified for incremental updates. A new DOI is a new dataset; a new CKAN revision_id on the same id is an update. Withdrawn IR items: keep the id, set status — do not delete unless the user asked for tombstones (harvest-incremental.md).