Skip to main content

Agent guide: contributing catalog records

Platform-neutral workflow for adding or editing catalog YAML. Full human guide: CONTRIBUTING.md.

Before editing​

  1. Search exports (and data/scheduled/) so you do not duplicate link / id. Finding candidates: discover.md.
  2. Read directory-layout.md and data-model.md.
  3. Consumers querying data should use ai-consumers.md — do not parse YAML unless authoring.

Source layout​

PathRule
data/entities/{CC}/{Federal|SUB}/{type}/{id}.yamlVerified records; filename = id
data/scheduled/Unverified; promote later (scheduled.md)
data/software/Platform definitions (software-taxonomy.md)
data/datasets/Generated only — never hand-edit

New catalog checklist​

  • Prefer CLI:

    python scripts/builder.py add-single "https://example.com/data" \
    --software ckan \
    --catalog-type "Open data portal" \
    --name "Example Data Portal" \
    --country US \
    --scheduled

    Useful options: --subregion PT-11 (routes to {CC}/{SUB}/, sets level 30), --owner-type "Local government" (validated, synonyms canonicalized), --is-national, --id (override for path-based tenants), --no-detect (skip the apidetect probe). Language is auto-filled from the country; records are schema-validated before writing. For several finds at once, write a JSONL manifest and run python scripts/builder.py add-batch manifest.jsonl — it dedupes against exports, validates rows, and assigns UIDs itself.

    After a live GET, promote in the same session (python scripts/promote_scheduled.py --id {id} [--probe] — --id promotes only your finds, --probe re-checks liveness and keeps dead records in the queue; subregion records route to {CC}/{SUBREGION}/{type}/ automatically), then assign (not needed after add-batch) and validate-yaml --id. If software.id is in CATALOGS_URLMAP, run python scripts/apidetect.py detect-software {id} --action insert --max-endpoints 1. Append one line to dataquality/hunts.jsonl (or use python scripts/hunt.py log). To triage the whole queue without moving anything: python scripts/promote_scheduled.py review-scheduled.

  • Filename / id: lowercase letters and digits only

  • Required fields: id, uid, name, link, catalog_type, access_mode, status, software, owner, plus coverage (enforced by the MISSING_COVERAGE quality rule rather than the schema)

  • Do not invent uid. Run python scripts/builder.py assign

  • owner.type from data/reference/owner_types.yaml

  • software.id from data/software/ (or custom)

  • Path country must match owner.location.country.id / coverage country

  • Resolve a country name or a bloc with Internacia (pip install internacia):

    python -c "from internacia import InternaciaClient; c=InternaciaClient(); print(c.search.fuzzy('QUERY', limit=5))"

    Country id is an alpha-2 with code_status == official_iso3166_1, plus XK. Write country.name from COUNTRIES (scripts/constants.py) or data/reference/countries.csv. Blocs (EU, ASEAN, Africa, treaties) are Internacia intblocks; registry path roots stay PATH_COUNTRY_ALLOWLIST. ISO 3166-2 subdivisions stay on pycountry or data/reference/subregions/.

  • Regional/local owners: owner.location.level 30 and a subregion folder (US-CA/, …)

  • properties.is_national: true only for the country’s official catalog of that type (national open-data portal, NSDI/geoportal, or NSO statistical product). Do not set it because the owner is a central/federal agency, the file is under Federal/, or the host is .gov. Agency, thematic, scientific, and subnational catalogs get is_national: false or omit the key. Full rule: data-model.md

After editing​

python scripts/builder.py validate-yaml --id {id}
python scripts/builder.py assign
pytest tests/test_yaml.py -q

Validate a single file with --file path/to/file.yaml. Run full validate-yaml before a large PR.

Updating existing records​

Prefer the CLI over hand-edits and throwaway scripts:

  • One field on one record: python scripts/builder.py set-field --id {id} --path properties.is_national --value true (--append for list fields like tags; values parse as YAML scalars).
  • Many records: write a JSONL manifest of {"id": ..., "set": {...}} / {"id": ..., "merge": {...}} rows and run python scripts/builder.py enrich-batch updates.jsonl (dry-run shows per-record diffs), then --write to apply.

Both refuse uid/id changes and re-validate the record against the schema before saving; invalid updates leave files unchanged.

Do not​

  • Commit generated data/datasets/ dumps unless the change is an intentional rebuild
  • Put secrets in YAML
  • Implement schema or pipeline changes without OpenSpec — see openspec-quickstart.md

Quality​

If analyze-quality flags the record, follow metadata-quality.md and quality-rules.md. Integrity-track issues (invalid enums, duplicates, path mismatches) must be fixed; enrichment-track gaps are optional in the same PR.

New shared platforms: software-taxonomy.md. After adding a data/software/ YAML, add fingerprints to the matching discovery guide.

What to work on next (software-first hunts, country-shape gaps, record reviews): improve.md.