Architecture
Layers
- Source YAML — one file per catalog or software definition. Edit these; never hand-edit
data/datasets/. - Reference vocabularies — allowed values under
data/reference/(owner types, catalog types, software IDs, access modes, status). - Validation — Cerberus schema (
data/schemes/catalog.json), JSON Schema (catalog.schema.json), and quality rules. - Build — flattens YAML into JSONL, compresses with zstd, writes Parquet and DuckDB. Shared URL helpers live in
scripts/url_utils.py; frozen one-shot scripts sit inscripts/archive/. - Consumers — DuckDB/Parquet preferred (nested fields are native
STRUCT/LIST); JSONL for line-oriented tools; YAML only when authoring.
Enrichment and monitoring
| Pipeline | Script / workflow | Output |
|---|---|---|
| Re3Data metadata | scripts/re3data_enrichment.py | _re3data on matching entities |
| CKAN ecosystem sync | scripts/sync_ckan_ecosystem.py | new scheduled or entity YAML |
| OpenAIRE Graph data sources | scripts/extract_openaire_portals.py | harvest list + scheduled YAML |
| API endpoint probe | scripts/apidetect.py | endpoints[] on known software.id maps |
| Quality analysis | python scripts/builder.py analyze-quality | dataquality/ |
| URL liveness | .github/workflows/liveness.yml + check_liveness.py --apply-dead | dataquality/liveness_report.jsonl; optional status: inactive |
| Hunt completeness | append-only dataquality/hunts.jsonl | skip exhausted software/IGO/list hunts |
| Integrity regression | tests/test_quality_regression.py | fails CI if CRITICAL/IMPORTANT counts grow |
Scope boundary
In-scope: YAML records, schema/validation, enrichment, quality analysis, dataset exports.
Out-of-scope: production query APIs and MCP servers (dateno-api and related Dateno services).
Related
- directory-layout.md
- discovery.md
- harvest.md (dataset crawl recipes; production harvest is reaper)
- exports.md
- cli.md
- metadata-quality.md
- re3data.md
- ckan-sync.md
- apidetect.md
- liveness.md
- enrichment.md