Agent guide: contributing catalog records
Platform-neutral workflow for adding or editing catalog YAML. Full human guide: CONTRIBUTING.md.
Before editing
- Search exports (and
data/scheduled/) so you do not duplicatelink/id. Finding candidates: discover.md. - Read directory-layout.md and data-model.md.
- Consumers querying data should use ai-consumers.md — do not parse YAML unless authoring.
Source layout
| Path | Rule |
|---|---|
data/entities/{CC}/{Federal|SUB}/{type}/{id}.yaml | Verified records; filename = id |
data/scheduled/ | Unverified; promote later (scheduled.md) |
data/software/ | Platform definitions (software-taxonomy.md) |
data/datasets/ | Generated only — never hand-edit |
New catalog checklist
-
Prefer CLI:
python scripts/builder.py add-single "https://example.com/data" \--software ckan \--catalog-type "Open data portal" \--name "Example Data Portal" \--country US \--scheduledUseful options:
--subregion PT-11(routes to{CC}/{SUB}/, sets level 30),--owner-type "Local government"(validated, synonyms canonicalized),--is-national,--id(override for path-based tenants),--no-detect(skip the apidetect probe). Language is auto-filled from the country; records are schema-validated before writing. For several finds at once, write a JSONL manifest and runpython scripts/builder.py add-batch manifest.jsonl— it dedupes against exports, validates rows, and assigns UIDs itself.After a live GET, promote in the same session (
python scripts/promote_scheduled.py --id {id} [--probe]—--idpromotes only your finds,--probere-checks liveness and keepsdeadrecords in the queue; subregion records route to{CC}/{SUBREGION}/{type}/automatically), thenassign(not needed afteradd-batch) andvalidate-yaml --id. Ifsoftware.idis inCATALOGS_URLMAP, runpython scripts/apidetect.py detect-software {id} --action insert --max-endpoints 1. Append one line todataquality/hunts.jsonl(or usepython scripts/hunt.py log). To triage the whole queue without moving anything:python scripts/promote_scheduled.py review-scheduled. -
Filename /
id: lowercase letters and digits only -
Required fields:
id,uid,name,link,catalog_type,access_mode,status,software,owner, pluscoverage(enforced by theMISSING_COVERAGEquality rule rather than the schema) -
Do not invent
uid. Runpython scripts/builder.py assign -
owner.typefromdata/reference/owner_types.yaml -
software.idfromdata/software/(orcustom) -
Path country must match
owner.location.country.id/ coverage country -
Resolve a country name or a bloc with Internacia (
pip install internacia):python -c "from internacia import InternaciaClient; c=InternaciaClient(); print(c.search.fuzzy('QUERY', limit=5))"Country
idis an alpha-2 withcode_status == official_iso3166_1, plusXK. Writecountry.namefromCOUNTRIES(scripts/constants.py) ordata/reference/countries.csv. Blocs (EU, ASEAN, Africa, treaties) are Internacia intblocks; registry path roots stayPATH_COUNTRY_ALLOWLIST. ISO 3166-2 subdivisions stay onpycountryordata/reference/subregions/. -
Regional/local owners:
owner.location.level30 and a subregion folder (US-CA/, …) -
properties.is_national: trueonly for the country’s official catalog of that type (national open-data portal, NSDI/geoportal, or NSO statistical product). Do not set it because the owner is a central/federal agency, the file is underFederal/, or the host is.gov. Agency, thematic, scientific, and subnational catalogs getis_national: falseor omit the key. Full rule: data-model.md
After editing
python scripts/builder.py validate-yaml --id {id}
python scripts/builder.py assign
pytest tests/test_yaml.py -q
Validate a single file with --file path/to/file.yaml. Run full validate-yaml before a large PR.
Updating existing records
Prefer the CLI over hand-edits and throwaway scripts:
- One field on one record:
python scripts/builder.py set-field --id {id} --path properties.is_national --value true(--appendfor list fields liketags; values parse as YAML scalars). - Many records: write a JSONL manifest of
{"id": ..., "set": {...}}/{"id": ..., "merge": {...}}rows and runpython scripts/builder.py enrich-batch updates.jsonl(dry-run shows per-record diffs), then--writeto apply.
Both refuse uid/id changes and re-validate the record against the schema before saving; invalid updates leave files unchanged.
Do not
- Commit generated
data/datasets/dumps unless the change is an intentional rebuild - Put secrets in YAML
- Implement schema or pipeline changes without OpenSpec — see openspec-quickstart.md
Quality
If analyze-quality flags the record, follow metadata-quality.md and quality-rules.md. Integrity-track issues (invalid enums, duplicates, path mismatches) must be fixed; enrichment-track gaps are optional in the same PR.
New shared platforms: software-taxonomy.md. After adding a data/software/ YAML, add fingerprints to the matching discovery guide.
What to work on next (software-first hunts, country-shape gaps, record reviews): improve.md.