Documentation contents
Overview
- Getting startedDuckDB, Parquet, and YAML paths into the registry.
- When to useWhat this registry is for, and what lives elsewhere.
- ArchitectureSource YAML, validation, enrichment, and exports.
- Directory layoutCountry, type, and filename conventions.
- CLIbuilder.py commands for build, validate, and quality.
- Scheduled entriesPromote unverified YAML from scheduled to entities.
- ReleasesTag, changelog, and GitHub release checklist.
Discovery
- Discover catalogsFind portals not yet in the registry, then add them.
- Search enginesGoogle, Censys, Shodan, FOFA, URLScan, and crt.sh recipes.
- Agents and LLM clientsConfigure Cursor, ChatGPT, Claude, MCP, and search APIs.
- Open data portalsCKAN, OpenDataSoft, Socrata, Idra, Piveau, Our Open Data, DataPress.
- GeoportalsOverview, then SDI stacks and regional viewers.
- Geoportals — SDIGeoNetwork, ArcGIS, STAC, openEO, Lizmap, QGIS Server, mviewer.
- Geoportals — viewersWagmap, EWMAPA, Tianditu, Masterportal, GeoMapFish, MapGIS.
- Scientific repositoriesDataverse, DSpace, Invenio, and other institutional IRs.
- Scientific — domainIPT, THREDDS, ERDDAP, Breedbase, Tripal, MassBank, ESGF.
- Metadata catalogsFAIR Data Point, Aristotle MDR, Fusion Registry, Metadata Browser.
- Indicators and microdataPxWeb, OpenSDG, Knoema, SDMX-RI, NADA, REDATAM, Mica.
- Search, ML, API, marketplacesIdra, OpenAIRE gateways, OpenML, API directories, data markets.
Harvest
- Harvest datasetsCrawl catalog APIs; filter datasets from publications.
- Scientific repository APIsDSpace, Invenio, EPrints, Pure, Esploro, and mixed IRs.
- Domain scientific APIsIPT, THREDDS, Breedbase, Tripal, VEuPathDB, MassBank, ESGF.
- Open data portal APIsCKAN packages vs resources; OpenDataSoft; Socrata.
- Geoportal APIsCSW, GeoNode, ArcGIS, STAC, OGC API collections.
- Indicators and microdata APIsPxWeb tables, SDMX dataflows, NADA studies, OpenSDG.
- Metadata catalog APIsFAIR Data Point DCAT, Aristotle MDR, Fusion Registry.
- Search, ML, and other APIsAggregators, OpenML, marketplaces, custom catalogs.
- Harvest by protocolOAI-PMH, CSW, DCAT, STAC, SDMX, OGC, ArcGIS REST.
- Incremental harvestsfrom=, updated filters, tokens, checkpoints.
- Earth-observation APIsTHREDDS, ERDDAP, STAC collections, Copernicus.
- Biodiversity and genomicsIPT, Symbiota, ALA, GBIF datasets, Ensembl.
- Map viewersLayer lists from QWC2, Masterportal, Lizmap — not tiles.
- Dataset identifiersNative id + catalog uid; DOI/handle; no cdi ids for datasets.
- Harvest outputJSON record shape, skip counts, empty-harvest checklist.
Data contracts
- AI consumer guideJoin keys, scope, and how to consume exports.
- Data modelRequired and recommended catalog fields.
- Catalog typesOpen data, geo, scientific, and related types.
- Software taxonomyPlatform IDs, categories, and subtypes.
- Software indexEvery software.id linked to discovery and harvest recipes.
- VocabulariesGeographic levels, identifiers, endpoints, topics.
- ExportsJSONL, Parquet, and DuckDB exports.
- Metadata qualityRecommended fields and quality-report tracks.
- Quality issue typesEvery analyze-quality code, track, and fix hint.
- Trust score0–100 scoring components and interpretation.
Pipelines
- Re3Data enrichmentFill _re3data from re3data.org identifiers.
- CKAN ecosystem syncImport CKAN sites from ecosystem.ckan.org.
- API endpoint detectionFill endpoints[] for known software.id URL maps.
- URL livenessWeekly link probes; report-only JSONL artifact.
- Enrichment and quality fixesQuality-fix scripts, infer_endpoints, and legacy enrich CLIs.