# dataportals-registry — agent index > Registry of open data portals, geoportals, scientific repositories, and related data infrastructure worldwide. Reference-data repository (YAML source of truth + exported datasets). Production search APIs live in dateno-api; this repository does not host a query API or MCP server. ## Start here - Documentation site: https://datenoio.github.io/dataportals-registry/ - [docs/getting-started.md](docs/getting-started.md): DuckDB / Parquet / YAML entry points - [docs/when-to-use.md](docs/when-to-use.md): scope and related Dateno repositories - [docs/ai-consumers.md](docs/ai-consumers.md): consumption contract, join keys, DuckDB LIST/STRUCT nested fields - [docs/data-model.md](docs/data-model.md), [docs/vocabularies.md](docs/vocabularies.md), [docs/catalog-types.md](docs/catalog-types.md), [docs/software-taxonomy.md](docs/software-taxonomy.md), [docs/software-index.md](docs/software-index.md) - [docs/exports.md](docs/exports.md), [docs/query-examples.md](docs/query-examples.md), [docs/cli.md](docs/cli.md) - [docs/discovery.md](docs/discovery.md): find catalogs not yet in the registry (human). A discovery session starts with `python scripts/hunt.py prior --target TARGET` and then follows each `next:` line ([docs/agents/discover.md](docs/agents/discover.md)). Hunt patterns: [docs/discovery.md#hunt-patterns](docs/discovery.md#hunt-patterns). Next hunt when no target was named: `python scripts/hunt.py next`. If `datasets.duckdb` is locked, `hunt.py dedupe` uses `full.parquet`. - [docs/software-index.md](docs/software-index.md): every `software.id` → discovery/harvest recipe - [docs/harvest.md](docs/harvest.md): crawl datasets from catalog APIs; filter datasets from publications in scientific IRs ([harvest-scientific.md](docs/harvest-scientific.md) ([domain](docs/harvest-scientific-domain.md)), [harvest-opendata.md](docs/harvest-opendata.md), [harvest-geoportals.md](docs/harvest-geoportals.md), [harvest-indicators.md](docs/harvest-indicators.md), [harvest-metadata.md](docs/harvest-metadata.md), [harvest-other.md](docs/harvest-other.md), [harvest-protocols.md](docs/harvest-protocols.md), [harvest-incremental.md](docs/harvest-incremental.md), [harvest-earthdata.md](docs/harvest-earthdata.md), [harvest-biodiversity.md](docs/harvest-biodiversity.md), [harvest-viewers.md](docs/harvest-viewers.md), [harvest-identifiers.md](docs/harvest-identifiers.md), [harvest-output.md](docs/harvest-output.md)) - [docs/discovery-search-tools.md](docs/discovery-search-tools.md): Google, Censys, Shodan, FOFA (Censys alternative), GitHub forks and code search, URLScan, crt.sh - [docs/discovery-agent-tools.md](docs/discovery-agent-tools.md): configure Cursor, ChatGPT, Claude, MCP, and search APIs for discovery - [docs/discovery-opendata.md](docs/discovery-opendata.md), [docs/discovery-geoportals.md](docs/discovery-geoportals.md) ([SDI](docs/discovery-geoportals-sdi.md), [viewers](docs/discovery-geoportals-viewers.md)), [docs/discovery-scientific.md](docs/discovery-scientific.md) ([domain](docs/discovery-scientific-domain.md)), [docs/discovery-metadata.md](docs/discovery-metadata.md), [docs/discovery-indicators.md](docs/discovery-indicators.md), [docs/discovery-other.md](docs/discovery-other.md): per-platform discovery queries - [docs/apidetect.md](docs/apidetect.md), [docs/liveness.md](docs/liveness.md): endpoint probes and URL reachability - [docs/agents/query.md](docs/agents/query.md), [docs/agents/discover.md](docs/agents/discover.md), [docs/agents/harvest.md](docs/agents/harvest.md), [docs/agents/contribute.md](docs/agents/contribute.md), [docs/agents/improve.md](docs/agents/improve.md), [docs/agents/openspec-quickstart.md](docs/agents/openspec-quickstart.md) - [AGENTS.md](AGENTS.md): project scope, schema, CLI commands, contribution workflow for AI assistants - [CONTRIBUTING.md](CONTRIBUTING.md): human contributor guide - [DATASHEET.md](DATASHEET.md): dataset purpose, bias, limitations, recommended uses - [CITATION.cff](CITATION.cff): academic citation metadata - [SECURITY.md](SECURITY.md): vulnerability reporting - [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md): contributor covenant - [openspec/ROADMAP.md](openspec/ROADMAP.md): OpenSpec change proposals (implementation waves; archive when a change is complete) ## Data layout - `data/entities/`: verified catalog records (one YAML per catalog, organized by country and catalog type). Source count 30 September 2026: **43,047** files in **225** country/territory folders (published v1.23.0: 43,047). - `data/scheduled/`: unverified records pending promotion — [docs/scheduled.md](docs/scheduled.md) (**49** YAML as of 30 September 2026; **49** in `scheduled.jsonl`) - `data/software/`: software/platform definitions (**784** YAML records plus `types.yaml`; `software.jsonl` rebuilt at **784**) - `data/schemes/catalog.json`: Cerberus validation schema (CI source of truth) - `data/schemes/catalog.schema.json`: JSON Schema with descriptions for external tooling - `data/schemes/catalog.context.jsonld`: DCAT/schema.org JSON-LD mappings - `data/reference/`: vocabularies (countries, subregions, owner types, endpoint types, …) ## Exports (generated) Run `python scripts/builder.py build` to regenerate. - `data/datasets/catalogs.jsonl.zst`: entity records only - `data/datasets/full.jsonl` (+ `.zst`, `.parquet`, `datasets.duckdb`): entities + scheduled - `data/datasets/software.jsonl`: software definitions - `data/datasets/catalogs.jsonld`: optional JSON-LD export (`python scripts/builder.py build --jsonld`) Filter by catalog type or `software.id` in DuckDB / Parquet; there are no `bytype/` or `bysoftware/` slice directories. ## Quality and monitoring - [docs/metadata-quality.md](docs/metadata-quality.md), [docs/quality-rules.md](docs/quality-rules.md), [docs/trust-score.md](docs/trust-score.md) - `dataquality/full_report.jsonl`: machine-readable quality issues (join on `uid`) - `dataquality/full_report.txt`: human-readable summary - `dataquality/primary_priority.jsonl`: CRITICAL + IMPORTANT issues - `dataquality/baseline_counts.json`: CI regression baseline for priority counts - `dataquality/liveness_report.jsonl`: URL reachability probes (weekly workflow artifact; apply with `check_liveness.py --apply-dead`) - `dataquality/hunts.jsonl`: discovery session completeness log (software/country/IGO/list URL) ```bash python scripts/builder.py validate-yaml python scripts/builder.py analyze-quality python scripts/update_quality_baseline.py python scripts/check_liveness.py --sample 10 ``` ## Pipelines - [docs/re3data.md](docs/re3data.md): Re3Data enrichment - [docs/ckan-sync.md](docs/ckan-sync.md): CKAN ecosystem sync - [docs/openaire-sync.md](docs/openaire-sync.md): OpenAIRE Graph data-source harvest list - [docs/harvest.md](docs/harvest.md): crawl datasets from catalog APIs (not a crawler in this repo; production harvest: reaper). Type guides: scientific, opendata, geoportals, indicators, metadata, other, plus protocols, incremental, earthdata, biodiversity, viewers, identifiers, and output. - [docs/apidetect.md](docs/apidetect.md): fill `endpoints[]` for known `software.id` maps - [docs/liveness.md](docs/liveness.md): weekly `link` probes (`dataquality/liveness_report.jsonl`) - [docs/enrichment.md](docs/enrichment.md): quality-fix scripts and legacy enrich CLIs - [docs/architecture.md](docs/architecture.md), [docs/directory-layout.md](docs/directory-layout.md) - [docs/releasing.md](docs/releasing.md): GitHub release checklist and docs GitHub Pages deploy - Tests: [tests/README.md](tests/README.md) ## Scope boundary In-scope: YAML records, validation, enrichment, quality analysis, dataset exports. Out-of-scope: production query API/MCP runtime ([dateno-api](https://github.com/datenoio/dateno-api), [reaper](https://github.com/datenoio/reaper)). ## License - Code: MIT - Data: CC-BY 4.0