Dataset identifiers from a harvest
This registry’s uid (cdi########) identifies catalogs. Harvested datasets need their own ids, always paired with the catalog uid. Incremental crawls: harvest-incremental.md. Protocol grain: harvest-protocols.md.
Do not mint cdi######## or temp######## for datasets. Do not write dataset YAML into data/entities/.
Store this tuple
| Field | Rule |
|---|---|
Catalog uid | From this registry (DuckDB/Parquet) |
Catalog id / link | Replay and debug |
software.id | Which recipe you used |
| Native id | The catalog’s primary key (required) |
| Persistent id | DOI, handle, or ARK when present (optional but preferred for dedup) |
| Landing URL | Canonical dataset page, not a session or tile URL |
| Type filter | The query/set you applied |
Prefer persistent ids when present
Order when several exist:
- DOI (
10.…) — DataversepersistentId, DataCite, Figshare, many IRs - Handle (
hdl:…orhttps://hdl.handle.net/…) — DSpace, some EPrints - ARK / PURL documented by the catalog
- Software-native id (CKAN
idUUID, Invenio record id, ERDDAPdatasetID, STAC collection id, CSWfileIdentifier, IPT dataset UUID, PxWeb table path)
Use the package/dataset id, not a file/resource/distribution id. ipums, dhis2, openaire, yoda, and radar are published software.id values — filter exports on those ids.
Native ids by platform (typical)
software.id | Native id | Not an id |
|---|---|---|
ckan / dkan / datapress | id (UUID) and/or name | Resource id |
ekan | Drupal dataset UUID (/jsonapi/dataset/dataset) | Resource UUID |
opendatasoft | dataset_id | Attachment filename |
socrata | id (four-four) | Column id |
dataverse | persistentId (DOI) | File id |
dspace / dspacecris | Handle or UUID | Bitstream id |
inveniordm | Record id / DOI | File key |
geonetwork | ISO fileIdentifier | Thumbnail URL |
stacserver | Collection id | Item id (unless item grain) |
openeo | Collection id | Process id, job id |
qgisserver / mapserver | WMS layer name | GetMap URL |
mviewer | Config layer id | Tile URL |
isogeo | Metadata record id | Workgroup id |
arcgisserver | Service URL + name | Extent query |
erddap | datasetID | Table row |
ipt | Dataset UUID / key | Occurrence id |
pxweb | Table path (type: t) | Folder path |
pxstat | Matrix / table code | Widget embed, demo tables |
fairdatapoint | Dataset IRI | Distribution IRI |
ipums | Series / sample id | Extract job id |
dhis2 | dataSet / indicator id | orgUnit, analytics cell |
openaire | Graph dataset product id / DOI | Publication id |
yoda | Vault dataset DOI | iRODS path in /research/ |
radar | RADAR dataset id / DOI | Landing-page URL only |
geocortex | Essentials site id | Html5Viewer tile URL |
vertigisstudioweb | Tenant host / ?app= GUID | Extra ?app= GUID on the same tenant; Html5Viewer tiles |
mapgisigserver | Map document / service name | /igs/manager |
hygmapgis | aplicacion= + layer name | Map tile URL / mapa.jsp session |
breedbase | study / trial id | plot or marker-call id |
esgf | dataset_id / master_id | file id |
symbiota | dataset RSS id / collid | occurrence id |
fenix | FENIX domain / dataset code | Observation cube cell |
tabnet | .def table name | CGI session / TabWin .TAB |
sparkmap | Map layer / assessment id | Saved user map |
g3wsuite | group/project + WMS layer name | /admin path |
sentinelhub | STAC collection id | Process job id, granule/item |
geusmap | mapname + WMS layer name | Map tile URL |
resourcecontracts | Contract id | Per-clause annotation |
gxopendata | Tenant dataset id | Apply-gateway request |
rdfrepository | License / workspace id | Login-only company filing |
converis | Dataset / research-data id | Publication or person id |
landfolio | Public cadastre layer / license id | ArcGIS REST service already harvested as arcgisserver |
labkey | Study / published folder id | Assay run id |
synapse | Project or dataset entity id (syn########) | Child file entity under a dataset |
xnat | Project id | Imaging session / DICOM file |
omero | Project / screen / study id | Image or well id |
kadi4mat | Record or collection id | File blob id |
edal | DOI dataset id | Single landing-page URL as a seed |
nomad | Entry / upload id | Calculation file or parser log |
intermine | Experiment / list / template-result id | Gene report |
gringlobal | Accession catalog export id | Individual accession HTML page |
plutof | Dataset / DOI id | Occurrence, sequence, UNITE taxon |
jgi | Genome / transcriptome project id | Gene page, BLAST hit, IMG/GOLD |
cbioportal | Study id | Mutation / CNA / patient sample row |
esasciencearchive | TAP table / observation-catalog id | FITS file or cutout |
odweb | /odweb/ dataset id | Parent government homepage |
imfnsdp | SDMX category / series code on the country page | DSBB directory, Knoema wrapper |
Normalize DOI to 10.prefix/suffix (lowercase). Strip https://doi.org/ and doi:. Handles: keep the handle string, not only the UI URL.
Deduplication
- Inside one catalog: native id (plus version if the API versions datasets).
- Across catalogs: DOI/handle first. The same Dataverse study harvested from DataONE and from the MN should collapse if you asked for a global index.
- Aggregators: prefer the source catalog
uidwhen Idra/DCAT names a publisher URL that is already in this registry (harvest-other.md).
A landing URL with query tokens, /latest/, or map bbox is a locator, not a durable id.
Versions and replacements
Keep version / metadata_modified for incremental updates. A new DOI is a new dataset; a new CKAN revision_id on the same id is an update. Withdrawn IR items: keep the id, set status — do not delete unless the user asked for tombstones (harvest-incremental.md).