Skip to main content

Architecture

Layers​

  1. Source YAML — one file per catalog or software definition. Edit these; never hand-edit data/datasets/.
  2. Reference vocabularies — allowed values under data/reference/ (owner types, catalog types, software IDs, access modes, status).
  3. Validation — Cerberus schema (data/schemes/catalog.json), JSON Schema (catalog.schema.json), and quality rules.
  4. Build — flattens YAML into JSONL, compresses with zstd, writes Parquet and DuckDB. Shared URL helpers live in scripts/url_utils.py; frozen one-shot scripts sit in scripts/archive/.
  5. Consumers — DuckDB/Parquet preferred (nested fields are native STRUCT / LIST); JSONL for line-oriented tools; YAML only when authoring.

Enrichment and monitoring​

PipelineScript / workflowOutput
Re3Data metadatascripts/re3data_enrichment.py_re3data on matching entities
CKAN ecosystem syncscripts/sync_ckan_ecosystem.pynew scheduled or entity YAML
OpenAIRE Graph data sourcesscripts/extract_openaire_portals.pyharvest list + scheduled YAML
API endpoint probescripts/apidetect.pyendpoints[] on known software.id maps
Quality analysispython scripts/builder.py analyze-qualitydataquality/
URL liveness.github/workflows/liveness.yml + check_liveness.py --apply-deaddataquality/liveness_report.jsonl; optional status: inactive
Hunt completenessappend-only dataquality/hunts.jsonlskip exhausted software/IGO/list hunts
Integrity regressiontests/test_quality_regression.pyfails CI if CRITICAL/IMPORTANT counts grow

Scope boundary​

In-scope: YAML records, schema/validation, enrichment, quality analysis, dataset exports.

Out-of-scope: production query APIs and MCP servers (dateno-api and related Dateno services).