Skip to main content

Getting started

dataportals-registry is a reference-data registry of open data portals, geoportals, scientific repositories, and related data infrastructure. Source records are YAML; consumers should prefer the exported datasets. High-volume platforms include CKAN, GeoNetwork, Dataverse, ArcGIS, openEO, mviewer, and DHIS2 — full map: software-index.md. Code is MIT; data and documentation are CC BY 4.0.

Working tree (30 September 2026): 43,047 verified catalog entities, 49 scheduled YAML records, and 784 software definitions across 225 country/territory folders. Last published snapshot is v1.23.0 (43,047 catalogs, 49 scheduled, 784 software). Record-count contract: exports.md.

Fastest path (analytics)​

Query the DuckDB export. Nested objects are STRUCT / LIST types, so filter with field access:

duckdb data/datasets/datasets.duckdb \
-c "SELECT id, name, link FROM catalogs WHERE software.id = 'ckan' LIMIT 10;"
import duckdb

con = duckdb.connect("data/datasets/datasets.duckdb")
con.execute(
"""
SELECT id, name, link
FROM catalogs
WHERE catalog_type = 'Open data portal'
AND list_contains(
list_transform(coverage, x -> x.location.country.id),
'US'
)
LIMIT 10
"""
).fetchall()

Parquet is interchangeable:

import duckdb

con = duckdb.connect()
con.execute("SELECT count(*) FROM 'data/datasets/full.parquet'").fetchone()

Fastest path (spreadsheet / JSONL)​

  • data/datasets/catalogs.jsonl.zst — one JSON object per verified catalog (compressed JSONL)
  • data/datasets/full.parquet — same records, analytics-friendly

Decompress .zst files with unzstd file.zst. Filter by type or software with DuckDB (query-examples.md).

Authoring path (YAML)​

Edit source YAML only when adding or correcting a catalog:

  1. Place the file at data/entities/{COUNTRY}/{Federal|SUBREGION}/{type}/{id}.yaml (see directory-layout.md)
  2. Match id to the filename (lowercase letters and digits only)
  3. Run python scripts/builder.py assign if uid is missing
  4. Validate with python scripts/builder.py validate-yaml --id {id}

Do not hand-edit data/datasets/.

Citation​

See CITATION.cff and DATASHEET.md.

dataportals-registry: A global registry of open data portals and catalogs
(Common Data Index, 2026). CC-BY-4.0.
https://github.com/datenoio/dataportals-registry

Next steps​

GoalDoc
Scope and when not to use this repowhen-to-use.md
Pipeline diagramarchitecture.md
CLIcli.md
Find catalogs not yet registereddiscovery.md
Google, Censys, and other search toolsdiscovery-search-tools.md
Configure search tools in Cursor / ChatGPTdiscovery-agent-tools.md
Open data / geo / scientific / metadata / indicators / other typesdiscovery-opendata.md, discovery-geoportals.md (SDI, viewers), discovery-scientific.md (domain), discovery-metadata.md, discovery-indicators.md, discovery-other.md
Harvest datasets from catalog APIsharvest.md, harvest-scientific.md (domain), harvest-opendata.md, harvest-geoportals.md, harvest-indicators.md, harvest-metadata.md, harvest-other.md, harvest-protocols.md, harvest-incremental.md, harvest-earthdata.md, harvest-biodiversity.md, harvest-viewers.md, harvest-identifiers.md, harvest-output.md
Endpoint detection / URL livenessapidetect.md, liveness.md
Field referencedata-model.md
Vocabularies (levels, identifiers, endpoints)vocabularies.md
Catalog typescatalog-types.md
Software IDs → discovery/harvest recipesoftware-index.md
Software IDs and new platform YAMLsoftware-taxonomy.md
Join keys and DuckDB columnsai-consumers.md
Quality issue codesquality-rules.md
Verified SQLquery-examples.md
Agent query workflowagents/query.md
Agent discovery workflowagents/discover.md
Agent harvest workflowagents/harvest.md
Add or edit YAMLagents/contribute.md
What to hunt or fix nextagents/improve.md