Skip to main content

AI consumer guide

Consumption contract for LLM agents, enrichment pipelines, and programmatic integrators. Installation and contribution: README.md.

Scope​

In scope​

Metadata about catalogs (portals, geoportals, repositories, and related infrastructure):

  • Identity: id, uid, name, link, status
  • Classification: catalog_type, software, tags, topics, content_types
  • Geography: owner.location, coverage[] (country, subregion, macroregion, level)
  • Access: access_mode, api, api_status, endpoints[], rights
  • Crosswalks: identifiers[] (wikidata, re3data, fairsharing, …)
  • Optional enrichment: _re3data, trust_score, properties.is_national (official national catalog of that type only — data-model.md)

Out of scope​

Do not expect:

  • Dataset-level records (the contents of each catalog)
  • A production HTTP search API or MCP server in this repository (use dateno-api for search)
  • Uniform last_verified_at timestamps on every record
  • Balanced geographic coverage (US records are over-represented)

Preferred access paths​

MethodPath
DuckDB (preferred in-repo)data/datasets/datasets.duckdb table catalogs
Parquetdata/datasets/full.parquet
JSONL entitiesdata/datasets/catalogs.jsonl.zst
JSONL + scheduleddata/datasets/full.jsonl
Software definitionsdata/datasets/software.jsonl and DuckDB table software
YAML sourcedata/entities/**/*.yaml — authoring only

Prefer DuckDB or Parquet over parsing thousands of YAML files.

Join keys​

EntityPrimary keyAlso useful
Cataloguid (cdi########)id (filename stem), link
Softwaresoftware.idjoin to data/software/**/{id}.yaml
Country coveragecoverage[].location.country.idISO alpha-2; some records use World / numeric M49
Owner countryowner.location.country.idshould match path country
External registryidentifiers[].id + identifiers[].valuewikidata, re3data, fairsharing

id is unique among files; uid is the stable identifier across exports. Do not join on name.

Nested fields in DuckDB / Parquet​

Lists are LIST and objects are STRUCT. Filter with field access and list functions:

SELECT id, name, link
FROM catalogs
WHERE software.id = 'ckan'
AND list_contains(
list_transform(coverage, x -> x.location.country.id),
'FR'
);
SELECT id, software.id AS software_id
FROM catalogs
WHERE software.id = 'geonetwork';

Column inventory: exports.md. Identifier types: vocabularies.md.

Status and access​

FieldValuesNotes
statusactive, inactive, scheduled, deprecatedcurated, not a live probe
access_modelist; prefer open / restrictedschema also allows limited, public, protected, closed, private
apibooleanif true, api_status should be set
api_statusactive, inactive, uncertain

Versioning​

Exports are regenerated with python scripts/builder.py build. There is no per-table _meta identity like internacia-db; treat git commit / GitHub release as the snapshot version. Record counts are listed in README.md.

Known limitations​

  • Geographic bias: see DATASHEET.md
  • Many records lack description, endpoints, or topics
  • Scheduled entries (when present) are unverified; current queue size is in exports.md
  • DuckDB/Parquet lag source YAML until the next build. Published snapshot vs working tree: exports.md
  • software may be custom / unknown when the platform is undetected