Skip to main content

Improve the registry (session playbook)

How to grow coverage and quality, based on ~3,900 Cursor sessions (November 2025–8 September 2026) and the releases they produced. This is the what to work on next guide. Mechanics live in discover.md, contribute.md, scheduled.md, and metadata-quality.md. Hunt-pattern table: discovery.md.

Working tree (30 September 2026): 43,047 verified catalogs, 49 scheduled YAML, 784 software IDs, 225 country folders. Published snapshot: v1.23.0 (43,047 catalogs, 49 scheduled, 784 software). Rebuild exports when YAML and data/datasets/ diverge.

What those sessions actually did​

First-user-message mix across 3,027 indexed chats (through v1.18.0):

Session typeShareTypical prompt
Add a single URL50%Add https://… / Adding URL to entities
Country or topic gap hunt18%Missing India data catalogs
Software taxonomy11%Popular geoportal software / add a software.id
Country record review7%Review records at data/entities/XX and fix them
Metadata / endpoint repair3%Update {id} metadata
Software-instance hunt1%Which DSpace catalogs are missing?
Docs, changelog, release1%Update README and CHANGELOG

| 31 August–8 September 2026 (~547 chats, covering v1.19.0 and v1.20.0) inverted that mix: ~82% were software-instance or subnational municipal-GIS hunts (Which {software} catalogs are missing?, Which {country} cities and counties geoportals are missing?). Custom-software reviews on 1–2 and 7 September extracted dozens of new software.id values, then instance hunts filled them. A 0-missing result is a valid done state — it documents completeness.

8–16 September 2026 (~390 chats, covering v1.21.0 and the unreleased +1,172 wave): country/IGO hunts 63%, software-instance 9%, municipal GIS 8%, country record reviews 0%. Volume came from IGO named-directory passes (WHO, FAO, IMF, UNESCO, WMO, …) plus Africa country-shape and a few new municipal GIS products (Map2Web, LDP SIT, Earthlight, Tailormap). Agents often stopped at scheduled YAML instead of promoting in the same session.

Volume came from a small number of session types, not from the 1,500 one-off URL adds:

Release / cycleNet catalogsSoftware IDsWhat drove it
v1.16.0+1,002+24Platform instance lists (DSpace, Figshare, Nordic viewers, TabNet, FENIX, …)
v1.17.0+2,608+15OpenAIRE Graph harvest (2,409 promoted; 664 dropped) + IMF NSDP + domain science
v1.18.0+2,243+20Viewer products (e-mapa, Kortasjá, Swing, Hajk), DSpace IRs, MappingSupport ArcGIS, India/Nigeria depth
v1.19.0+4,823+84Municipal GIS viewers (Experience Builder, Mapotip, GisMaster, GISPLAN, IntraMaps, SonicWeb, …), harvest-source dumps, university IRs, named directories
v1.20.0+5,450+57Polish e-mapa.net, Czech/Slovak municipal GIS (GISPLAN, GisOnline, Mapotip, GEPRO, mOBEC), Italian GisMaster, Swiss GeoMapFish, Brazilian CTMGEO, Japanese WagMap, new software IDs (WebEWID, DaCHS, Argenmap, …)
v1.21.0+1,904+63Custom-catalog OSS IDs (Géoclip, GeoNature, FROST-Server, MINERVA, Datasette, TerriSTORY, InstantAtlas, OPENDATAENTE), harvest/apidetect recipe fill, scheduled queue emptied
v1.22.0+3,997+179Municipal GIS and open-data SaaS (Map2Web, Municipium, BODIK ODCS, LDP SIT, GFMaplet, GajaMatrix), custom-catalog retags, hunt toolkit and batch ingestion
v1.23.0+1,880+119Gosweb Russia mega-hunt (408 regional portals), global indicators and trade portals, MapaInversiones tenants, TLD sweeps (.io/.ai/.int), custom-software retags, gazettes scoped out

Lesson: one bounded vendor tenant list or hostname pattern (e-mapa.net, {city}.gisplan.sk, GisMaster IdCliente=) outperforms Google for every city in the country. After v1.18.0 the high-yield prompts were harvest sources, university IRs, country indicators, and named directories. After v1.19.0 they shifted again: define the municipal GIS product, then hunt its tenants. Country hunts still matter when the hole is shape (no dataset-bearing scientific IRs, no native NSO table DB, UN members with zero open-data YAML), not count. Do not repeat “which France/Spain/Texas cities geoportals” now that those markets are dense.

Operating loop​

Run these in order. Skipping a step is how duplicates, dead hosts, and custom sprawl get in.

1. Scope DuckDB gap query or a vendor list (not a YAML walk)
2. Discover software-first, then country-shape. Probe candidate hosts only
3. Stage add-single --scheduled (temp UIDs)
4. Probe live GET; 401/403 stop; dead → drop or inactive
5. Promote promote_scheduled.py in the **same session**; assign; validate-yaml
6. Enrich detect-software for mapped IDs; country review: owner, endpoints, is_national, type path
7. Log append one line to dataquality/hunts.jsonl (0 missing is still a complete row)
8. Release analyze-quality, build, README/CHANGELOG/AGENTS.md/llms.txt, software-index

If datasets.duckdb is locked (another session writing), query data/datasets/full.parquet instead — it is read-only and enough for hostname duplicate checks. Do not walk YAML because DuckDB failed.

Reuse a prior hunt transcript or dataquality/hunts.jsonl for the same software.id, IGO, list URL, or country before starting a second pass. If that hunt already exhausted the vendor list, report completeness and stop unless a new list URL exists.

Hunt log (one JSON object per line):

{"date":"2026-09-16","kind":"software-instance","target":"tailormap","list_url":"https://www.tailormap.com/","added":11,"skipped_dupes":4,"status":"complete","notes":"vendor list exhausted"}

kind is one of: analysis, country-gap, country-geoportals, country-indicators, country-opendata, country-review, country-scientific, country-shape, custom-review, igo-organization, municipal-gis, named-directory, national-harvest-sources, process, scheduled-review, single-add, software-instance, university-ir. status is complete, partial, or blocked. The same set is HUNT_KINDS in scripts/hunt.py.

Do not commit .tmp_aq/ probe scripts or JSON. Rebuild data/datasets/ only when the user asked for a release or the YAML/export counts have diverged.

What moved the needle (do more of this)​

1. Software-first discovery​

Highest catalogs-per-hour. Pattern that worked:

  1. Confirm the product is shared → add data/software/{category}/{id}.yaml (software-taxonomy.md).
  2. Add fingerprints + harvest recipe in the same change (software-index.md is CI-guarded).
  3. Take the vendor/government list, not a web-wide crawl: OpenAIRE Graph, re3data, OpenDOAR, Dataverse installations, CKAN ecosystem, IMF DSBB, IHSN/NADA, DHIS2 country list, MappingSupport, national harvest source APIs (data.gov.* harvest endpoints), named directories (ODIS, CoreTrustSeal, WIS2 GDC, STAC Index, GeoNode gallery).
  4. Duplicate-check exports on hostname. Probe /api/… fingerprints from discover.md.
  5. Stage, then promote only hosts that respond.

v1.16–v1.17 software definitions (Trimble Locus, Spatial Suite, G3W-SUITE, IMF NSDP, InterMine, GRIN-Global, LabKey, …) each unlocked a clean instance batch. v1.19–v1.20 municipal GIS IDs (e-mapa, GISPLAN, GisMaster, IntraMaps, SonicWeb, WebEWID, …) did the same at country scale: product first, then the vendor tenant list, not “Google every city”. Retag existing software.id: custom rows onto the new id in the same PR.

After v1.20.0, prefer instance hunts for IDs that still have pending vendor lists (changelog “instance hunts pending”) over re-running products already hunted on 3–5 September 2026. A Censys or FOFA title/body pass is worth it when Google and the gallery are exhausted and the product has a distinctive HTML fingerprint — not as a replacement for the vendor list. Use FOFA when Censys search is unavailable, or for East Asian hosts.

Branded products with a first-party product page or vendor deployment list may get a software.id even with one or two registry rows, then the instance hunt. Unnamed one-off .gov map roots stay custom. Rule: software-taxonomy.md.

2. Graph and harvest-source dumps​

OpenAIRE (scripts/extract_openaire_portals.py) added thousands of scientific IRs in one cycle — and also produced 664 rejects (dead, parked, journals, software forges, staging). CKAN ecosystem is now ~10 unmatched sites; do not re-run it as a volume hunt.

Accept only public catalog UIs / harvest APIs. Journals, forges, single publications, and login walls are out of scope (discover.md).

National harvest source lists (data.gouv.fr, data.gov.uk, dados.gov.pt, data.europa.eu, dataportal.se EntryStore, data.go.id, govdata.de, datos.gob.es, data.go.kr, data.gov.ru, opendata.swiss, dane.gov.pl, search.open.canada.ca) beat Google for municipal open data in countries that already have a national CKAN. data.go.id yielded 112 origin catalogs in one pass; dane.gov.pl’s 7,477 institutions were almost all XML dataset feeds — only the six CKAN harvests were catalogs. Probe the origin UI, not the harvest-source row.

3. Country shape hunts (not more US ArcGIS)​

The registry is dense in the US (~29% of rows) and Western Europe, and geoportal-heavy (~half of YAML). The remaining high-yield holes are:

  • Populous countries with few catalogs relative to known ecosystems (India remaining states/cities after the first depth pass).
  • Countries that already have CKAN/geo but almost no dataset-bearing scientific IRs (filter OpenDOAR/DSpace to Dataset type; publication-only IRs are out of scope).
  • Native NSO table databases still missing after the 29–30 August indicators wave (Africa, some Pacific) vs IMF NSDP proxies already present.
  • DHIS2 / NADA / REDATAM remaining countries (short vendor lists).
  • Types that barely exist: microdata, metadata, API catalogs, ML catalogs.
  • Named directories not yet exhausted: national harvest leftovers, ODIS/WIS2/CLARIN follow-ups.

Deprioritize: another US county ArcGIS Server, another French/Spanish/Czech/Slovak/Polish/Italian commune geoportal, Open Data Inception rows that are PDFs or election pages, tiny territories that already have an IMF NSDP stub, university IR hunts for Monaco/Liechtenstein/Kiribati-class stubs, guessed HCI/Virtual LMI county hostnames, repeating a software-instance hunt from the last two weeks unless a new gallery appeared.

Exception: a bounded, high-precision list in a saturated geography is still worth it (MappingSupport live REST roots, Cadcorp council WebMaps, SeaSketch public /app tenants, GDi Visios county viewers, Instant Apps Filter Gallery hosts). Unscoped “missing US catalogs” or “all 400 Iranian counties” is not — public catalogs are the few cities that publish a REST/GeoServer/GeoNode UI, not every administrative unit.

4. Country record reviews​

Review records at data/entities/{CC} and fix them was run across ~216 folders. That is how owners, HTTPS, dead hosts, harvest URLs, topics, and is_national got cleaned. It is still the right quality pass when a country was filled quickly from OpenAIRE or a vendor list.

Probe the live site before editing. Do not invent /csw or /api/3 paths. Inactive catalogs must not keep api: true or live harvest endpoints.

5. is_national as official product, not federal owner​

A dedicated classifier (scripts/national_catalog.py, scripts/fix_is_national_flags.py) unset the flag on 988 agency/thematic/scientific/subnational records. Rule: true only for that country’s official catalog of that type (national open-data portal, NSDI/geoportal, or NSO product). File path Federal/ and .gov are not enough. Full rule: data-model.md.

6. Scheduled as a staging area, then empty it​

Successful cycles grew data/scheduled/, live-checked URLs, promoted hundreds, and dropped the rest (duplicates, HTTP 502, empty dashboards, hijacked domains). Prefer --scheduled on discovery; promote only after a GET. Script: scheduled.md. Keep the queue near zero between releases so catalogs.jsonl.zst and YAML counts stay explainable.

What wasted time (do less of this)​

Anti-patternWhat happened
Walk data/entities/**/*.yaml to searchSlow, misses scheduled, duplicates slip through. Use DuckDB / Parquet.
Treat DuckDB lock as a blockerParallel hunts lock datasets.duckdb. Query full.parquet.
Guess harvest endpoints from software docsQuality rules then fire SOFTWARE_EXPECTED_ENDPOINTS_MISSING; inactive sites get fake APIs. Probe, then write.
Treat every federal/agency catalog as national988 false is_national: true flags.
Add OpenAIRE / re3data rows without accept/rejectJournals, forges, parked domains, staging hosts.
Login-walled CRIS (hosted Symplectic, SSO Pure)Stop on 401/403; do not add.
“Missing catalogs” for Nauru / San Marino-class stubsAlready have indicator pages; yield is one PDF.
University IR hunt for Monaco / Liechtenstein / KiribatiNo dataset-bearing IR; waste a session.
Add every OpenDOAR / DSpace hostPublication-only IRs (Kazakhstan: 10 later removed). Require Dataset type or a research-data community.
Treat national harvest rows as catalogsdane.gov.pl: 7,477 institutions, almost all XML feeds; opendata.swiss geocat/I14Y slices. Probe the origin UI.
Guess HCI / Virtual LMI / Cancer-Rates county hostsTimeouts and login loops. Use the vendor tenant list, not DNS guesses.
Register every PISO / SeaSketch / GISApp copyOne catalog per public product; skip marketplace demos and REST adaptors.
Google every city after a tenant list existsIran/Texas-class sweeps. Use {country} + the product hostname pattern.
Repeat a software hunt from this weekSept 3–5 already covered most IDs. 0 missing is done.
Directory hubs and survey platformsIPUMS Health Surveys (lists collections), SurveySolutions (data collection, not a catalog), Oskari RPC embeds of Suomi.fi.
Mix ArcGIS product IDs on one orgHub vs Experience Builder vs Web AppBuilder vs Instant Apps vs Dashboards vs StoryMaps vs Server — one public catalog UI unless they are distinct products.
Leave YAML ahead of exportsREADME/docs quote stale counts; CI quality baseline drifts. Rebuild before release.
Commit .tmp_aq/ probesScratch only.
New software.id for an unnamed .gov rootKeep custom until a named product (vendor page or ≥3 independent installs) exists.
Implement query APIs / MCP in this repoOut of scope. Reference data only.

Priority queue​

Re-check counts in DuckDB (and YAML if exports lag) before starting. India/Nigeria/DHIS2/DSpace, the 29–30 August country-indicators wave, the university-IR country wave, and the 31 August–5 September municipal-GIS / software-instance wave already landed. After v1.20.0, PL/CZ/SK/IT/CH/JP/BR municipal GIS is dense — do not start another commune sweep there.

RankHuntWhyHow
1Dataset-bearing scientific IRsShape hole: KH, BA, PS, EG, LA still ~0–2 scientific at ≥20 total catalogsOpenDOAR + re3data + OpenAIRE, filter DSpace Dataset type / Dataverse; skip microstates
2Named directoriesBounded lists still convert: ODIS, CoreTrustSeal leftovers, STAC Index, WIS2 GDC, GeoNode galleryOne list URL per session (discovery.md); log the URL in hunts.jsonl
3IGO / World catalogsSep 15 IGO wave worked; UN family still has leftover portalsWhich {agency} data catalogs are missing? (discover.md)
4Africa national + capital open dataStill 0 OD YAML: EG, AO, CM, SD, HT, NE, YE, ZW, and othersNational CKAN/DKAN/uData, then capital city. Skip more DHIS2 if already added
5Custom retag on scientific + indicatorscustom is 40% of scientific and 53% of indicatorsHostname/path clusters with a named product; one-off .gov roots stay custom
6Native NSO / health / education leftoversAfrica and some subnational explorers remainPxWeb / .Stat / STATcube / DHIS2 / TabNet; skip IMF NSDP already present
7National harvest-source leftoversOther national portals still have unmatched harvest URLsHarvest/organisations API → probe origin UI
8Thin typesMicrodata 298, metadata 137, API 51, ML 71, marketplace 46IHSN leftover, government API directories, FAIR Data Point
9India remaining depthPopulation × empty states/cities after the first passesState SDI, city CKAN, dataset-bearing university IRs. Duplicate-check *.data.gov.in
10Country reviews after bulk addsReviews stopped in late August; World/NL/GB/IT/RO/AT gained thousands of rowsReview records at data/entities/{CC}
—More US/EU commune ArcGIS or e-mapa-class GISGeo is 58% of the registry; PL/CZ/IT/NL/RO Map2Web filledOnly if a named gallery remains unmatched

Recipes​

Software-instance hunt​

Which {software.name} catalogs are missing?

Agent steps:

  1. Read the software YAML and software-index.md row.
  2. SELECT link, owner.location.country.id FROM catalogs WHERE software.id = '{id}' on datasets.duckdb (or full.parquet if DuckDB is locked).
  3. Fetch the vendor list / gallery / crt.sh hostname pattern (not a scanner). For open-source products deployed by forking, that list is GitHub forks plus a code search that excludes forks (discovery-search-tools.md). Skip if a hunt in the last two weeks already exhausted that list.
  4. Match on hostname; probe fingerprints; add-single --scheduled.
  5. If the vendor list is exhausted and probes found nothing new, stop and report completeness (0 missing is done). Append dataquality/hunts.jsonl.
  6. If ≥3 custom rows are clearly this product, or a first-party product page names it, retag them and add the software definition first.
  7. Optional second pass: Censys html_title / body fingerprint when Google and the gallery are silent, or the FOFA equivalent (title= / body= / country=) when Censys is not configured.
  8. Promote in the same session after a live GET. Run detect-software {id} --action insert --max-endpoints 1 when the software is in CATALOGS_URLMAP.

IGO / international organization hunt​

Which {agency} data catalogs are missing?

Agent steps:

  1. Query exports for World/ plus hostnames (who.int, fao.org, …).
  2. Take the agency’s own data/publications/geoportal directory — not a web-wide crawl.
  3. Accept: catalog UIs, statistical databases, geoportals, SDMX registries. Place under data/entities/World/ unless a single-country owner is clear.
  4. Reject: publication article hubs, one-file Excel “databases”, login walls, marketing homepages, duplicate language mirrors.
  5. Promote live finds in the same session. Log the agency in dataquality/hunts.jsonl.

Country-shape hunt​

Missing {country} data catalogs

Agent steps:

  1. Count YAML by type for that ISO folder (opendata / geo / scientific / indicators / microdata).
  2. Hunt the missing type, not the type already in the hundreds.
  3. Sources: national harvest API, re3data country facet, OpenDOAR, NSO site, university IR lists, local-language open-data terms (datos abiertos, data terbuka, mở dữ liệu).
  4. Place local owners in {CC}/{ISO-3166-2}/{type}/ with owner.location.level 30.
  5. For geoportals, use the municipal GIS product list for that country (e-mapa, GISPLAN, GisMaster, IntraMaps, SonicWeb, …). Do not Google every city name.

Subnational municipal GIS hunt​

Which {country} cities and counties have geoportals that are missing?

Agent steps:

  1. Count existing geo/ YAML for that ISO folder. If geo is already the majority type, stop unless a named product list remains unmatched.
  2. Identify the dominant viewer product(s) from software-index.md / prior country hunts.
  3. Hunt that product's tenant list or hostname pattern — not every administrative unit.
  4. Accept live public viewers. Reject REST adaptors of an existing Hub, marketplace demos, login staff GIS, and “all 400 counties” guesses (Iran: only cities with a public ArcGIS/GeoServer/GeoNode UI).

National harvest-source hunt​

Which data sources harvested by {national portal} are missing?

Agent steps: discover.md. Full accept/reject: discovery-opendata.md.

Country university IR hunt​

There are a lot of {country} universities and research organizations that could have scientific data repositories that are not yet listed. Which of them are missing?

Agent steps: discover.md. Require Dataset type: discovery-scientific.md.

Country indicators hunt​

Which {country} indicators catalogs are missing?

Agent steps: discover.md. Skip IMF NSDP already present: discovery-indicators.md.

Named directory hunt​

Which data catalogs from {list URL} are missing?

One bounded URL. Duplicate-check hostname. Probe live. Skip preservation-only systems and org homepages. Lists: discovery.md.

Country review​

Review records at data/entities/{CC} and fix them

Checklist that reviews actually used:

  • Live link (HTTPS, no trailing slash unless required); status: inactive if dead/parked
  • software.id matches a probe, not the old guess
  • catalog_type matches the directory (scientific/ vs opendata/)
  • Owner name + owner.link from the site, not a generic ministry string
  • Coverage country matches path; quote UN M49 macroregion ids ('155') so they stay strings
  • Harvest endpoints only after a 200 on that path; set api to match
  • properties.is_national only for the official product of that type
  • assign + validate-yaml --id on touched files

Quality batch​

Integrity-track issues (invalid enums, duplicates, path mismatches) block CI. Enrichment-track gaps (MISSING_TAGS, expected endpoints) can wait. Workflow: metadata-quality.md, quality-rules.md. After OpenAIRE-scale adds, run analyze-quality before the next hunt or the baseline explodes.

Software definition multiplier​

Do not hunt instances of a product that has no software.id. Add the definition + discovery/harvest headings + docs_software_coverage in one change, then the instance hunt. CI fails unless both guides have a unique {#id} heading (tests/test_docs_software_coverage.py).

DuckDB checks before a hunt​

-- Type mix for a country (working-tree export; rebuild if YAML is ahead)
SELECT catalog_type, count(*) n
FROM catalogs
WHERE owner.location.country.id = 'IN'
GROUP BY 1 ORDER BY n DESC;

-- Custom share (retag candidates)
SELECT catalog_type,
count(*) n,
sum(CASE WHEN software.id = 'custom' THEN 1 ELSE 0 END) AS custom_n
FROM catalogs
GROUP BY 1;

-- National open-data flag coverage
SELECT owner.location.country.id AS cc, count(*)
FROM catalogs
WHERE catalog_type = 'Open data portal'
AND properties.is_national = true
GROUP BY 1;

If YAML and catalogs.jsonl.zst disagree, say so and prefer YAML counts (data/entities/**/*.yaml) for “what is already added,” exports for hostname duplicate checks until the next build.

Session prompts that work​

Copy these; they match the loops above.

  • Which {software} catalogs are missing? — instance hunt (0 missing is done)
  • Which {agency} data catalogs are missing? — IGO / World hunt
  • Which data sources harvested by {national portal} are missing? — harvest-source hunt
  • There are a lot of {country} universities… Which scientific repositories are missing? — IR hunt (dataset-bearing only)
  • Which {country} indicators catalogs are missing? — then skip IMF NSDP / national StatBank already registered
  • Which catalogs from {list URL} are missing? — named directory
  • Missing {country} data catalogs — then follow the type-shape table, do not add more geo if geo is already 70%
  • Which {country} cities and counties have geoportals that are missing? — subnational; use the municipal GIS product tenant list, not every city name
  • Review custom {type} catalogs and identify new software definitions
  • Review records at data/entities/{CC} and fix them — quality
  • Review scheduled and promote if they are ok, otherwise remove them
  • Find popular {type} software not yet in software records — taxonomy
  • Update README and CHANGELOG — after a batch, with rebuilt exports

Avoid: Find all missing catalogs in the world, Search the internet for ArcGIS, Mark every Federal catalog as national, university IR hunts for microstates, guessed county HCI hostnames, Which {saturated country} cities geoportals are missing? when that country's municipal GIS product is already tenant-complete, repeating Which {software} catalogs are missing? for an ID hunted in the last two weeks.

Done when​

A discovery session is done when every accepted URL has YAML, UID, and validate-yaml --id, skipped duplicates are listed with their existing id, and exhausted vendor lists are reported as complete (including 0 missing).

A country review is done when validate-yaml passes for that folder and live probes match status / api / endpoints.

A release session is done when YAML count = export count, scheduled is 0 (or explained), quality baseline is refreshed, and README / CHANGELOG / llms.txt quote the same numbers.