Harvesting datasets from open data portals
CKAN, OpenDataSoft, Socrata, and similar portals already treat datasets (packages, views) as the primary object. ResourceContracts and RDF Online Repository list contracts / licenses; Guangxi tenants and ODWeb catalogs list open datasets. You rarely need the publication-type filters used for scientific repositories. You still must avoid harvesting the wrong object (resources, showcases, harvest sources, individual files, login-only filings).
Overview: harvest.md. Finding portals: discovery-opendata.md. Shared DCAT/CKAN grain: harvest-protocols.md.
Replace https://host with the catalog origin. GET only. Stop on 401/403. Prefer URLs already in endpoints[].
What to keep
| Keep | Drop |
|---|---|
| Dataset / package / view / explore dataset | A file resource (CSV) as if it were a separate catalog record |
| ResourceContracts contract / RDF Online Repository license | Vendor About page, per-clause chrome, login-only company filing |
| Guangxi tenant open dataset | Login-only apply / API-gateway flows |
ODWeb /odweb/ dataset | Parent government homepage, login-only apply flows |
OpenGov report / story on {org}.opengov.com | Vendor marketing; Highcharts series as separate catalogs |
| The dataset landing API object | Harvest source metadata (the remote catalog CKAN is pulling from) |
| Showcase, article, blog, app gallery, idea box | |
| Organization or group objects without a dataset list |
One dataset with five CSV resources is one dataset.
CKAN (ckan)
Search (preferred):
GET https://host/api/3/action/package_search?q=&rows=100&start=0
Use start += rows until count is reached. Each results[] element is a dataset.
Optional filter when the site mixes types:
GET https://host/api/3/action/package_search?fq=dataset_type:dataset&rows=100
Drop dataset_type:showcase (ckanext-showcase), harvest objects, and type:harvest.
package_list returns names only and is painful on large sites — prefer package_search.
DataPress (datapress) is CKAN plus CMS — harvest package_search, not CMS pages. Do not also harvest the same host as ckan.
OpenAIRE Graph/CONNECT gateways are data search engines, not open-data CMSs. Recipe: harvest-other.md.
Keep: package_search packages (results[]). Drop: dataset_type:showcase, harvest source objects, type:harvest, and individual resources.
Optional dumps when present in endpoints[]: OAI-PMH Identify (/oai?verb=Identify) and SPARQL (/sparql, ckanext-dcat / ckanext-sparql). Prefer package_search for dataset harvest; SPARQL is a catalog dump, not a substitute for packages.
Andino (andino)
Argentine CKAN distribution (datos.gob.ar) — harvest exactly like CKAN: GET https://host/api/3/action/package_search?q=&rows=100&start=0. Andino also exposes DCAT dumps at /data.json, /catalog.xml, and /catalog.jsonld, and a time-series API (/series/api/series/) — the series API is indicator observations, not datasets; do not harvest it as packages.
Keep: package_search packages. Drop: showcases, harvest objects, series API observations.
BODIK ODCS (bodikodcs)
Shared CKAN at data.bodik.jp. Harvest one municipality at a time.
GET https://data.bodik.jp/api/3/action/package_search?q=organization:{lgcode}&rows=100
Keep: packages for that organization. Drop: other municipalities' packages, the WordPress marketing pages, and non-municipality organization ids. The aggregate package_search without an organization filter belongs only on the shared data.bodik.jp catalog.
jig.jp Open Data Platform (jigodp)
One shared CKAN. Municipalities are organizations, not separate catalogs.
GET https://ckan.odp.jig.jp/api/3/action/package_search?rows=100
Keep: packages from that catalog. Drop: the odp.jig.jp marketing site. Do not register each organization as its own catalog.
DKAN (dkan)
Same Action API as CKAN when enabled; also /api/1/search. Confirm JSON "success": true. If only Drupal JSON:API is public, harvest /jsonapi/dataset/dataset (DKAN 2 dataset entity) and prefer dkan when the product is DKAN.
Keep: packages from CKAN-compatible Action API or /api/1/search, or DKAN 2 JSON:API dataset entities.
Drop: Drupal nodes that are not datasets. Prefer dkan over drupal when the product is DKAN.
GET https://host/api/3/action/package_search?rows=25
GET https://host/api/1/search
GET https://host/jsonapi/dataset/dataset
GET https://host/api/3
EKAN (ekan)
Drupal 9/10 DKAN-1 successor. Harvest Drupal JSON:API dataset entities when public; otherwise the DCAT dump.
Keep: /jsonapi/dataset/dataset entities, or /data.json / /catalog.xml dataset records.
Drop: Drupal page / article nodes, login forms, vendor demo and staging hosts. Prefer ekan over dkan or drupal when ekan_theme or /profiles/contrib/ekan/ is present. Do not harvest GetDKAN /api/1/search on these sites — that route is not EKAN.
GET https://host/jsonapi/dataset/dataset
GET https://host/data.json
GET https://host/catalog.xml
OpenDataSoft (opendatasoft)
Already a dataset catalog:
GET https://host/api/explore/v2.1/catalog/datasets?limit=100&offset=0
Follow links / offset until total_count. Do not harvest the vendor academy or www.opendatasoft.com.
Optional DCAT dump (same dcat:Dataset grain as harvest-protocols.md):
GET https://host/api/v2/catalog/exports/dcat
GET https://host/api/explore/v2.1/catalog/exports/dcat
GET https://host/data.json
/data.json is Project Open Data (dcatus11) on some tenants. Prefer the explore catalog API when both exist.
Keep: explore datasets. Drop: the vendor academy, www.opendatasoft.com, and individual records as extra catalogs.
Socrata (socrata)
Views include charts, maps, files, and stories.
GET https://host/api/catalog/v1?only=datasets&limit=100&offset=0
GET https://host/opensearch.xml
OpenSearch description is /opensearch.xml. Detection types it opensearch.
Legacy: /api/views.json mixes types — filter viewType / displayType to tabular datasets, or use the catalog API only=datasets. Drop only=stories, only=filters.
Keep: catalog API views with only=datasets. Drop: only=stories, only=filters, charts-as-datasets, and mixed /api/views.json without a type filter.
uData (udata)
GET https://host/api/1/datasets/?page_size=100&page=1
Do not page /api/1/reuses/ or /api/1/posts/ as datasets.
Keep: /api/1/datasets/ dataset objects. Drop: /api/1/reuses/ and /api/1/posts/.
PortalJS (portaljs)
Datopian catalog shells. Harvest dataset pages and, when it answers, the CKAN package API on the host the portal uses.
GET https://host/search
GET https://host/@{org}/{slug}
GET https://host/api/3/action/package_search?rows=100
The package API is sometimes on a linked host (api., admin., or ckan.). Keep: dataset pages under /@{org}/{slug} and package_search results. Drop: portaljs.com marketing, arc.portaljs.com sign-in, and GitHub example repos. One portal host = one harvest scope.
Magda (magda)
Catalog search API from endpoints[]. Keep datasets, not portal chrome.
Keep: datasets from the catalog search API in endpoints[].
Drop: portal chrome, organisations, and distributions as extra datasets.
GET https://host/api/v0/search/datasets
JKAN (jkan)
Often static JSON in the repo.
GET https://host/data.json
GET https://host/datasets.json
Keep: published dataset entries. Drop: GitHub issues and the JKAN docs site.
KUKAN (kukan)
CKAN-compatible read API plus the native package list. Keep datasets, not admin or auth routes.
GET https://host/api/3/action/package_search?rows=100
GET https://host/api/v1/packages
Keep: package/dataset records and their resources. Drop: /api/auth, /api/mcp, /api/v1/admin, and the kukan.dev product site. If /api/3/action/status_show returns ckan_version, harvest it as CKAN instead.
Datasette (datasette)
Harvest the published instance root. Each SQLite database or listed table is a dataset; canned queries and SQL result rows are not extra catalogs.
GET https://host/-/databases.json
GET https://host/{database}.json
GET https://host/{database}/{table}.json?_shape=array&_size=1
Keep: published databases/tables and their JSON/CSV exports. Drop: /-/ debug pages, canned-query result rows, and datasette.io marketing.
Datadex (datadex)
Static portals. The dataset page lists Parquet and CSV tables. Files are often on data.{host}.
GET https://host/datasets
GET https://data.host/{table}.parquet
GET https://data.host/{table}.csv
Keep: tables named on that dataset page. Drop: datadex.datonic.io, GitHub source repositories, and Hugging Face dataset pages already covered by the Hugging Face Datasets catalog. Some instances (Datania) only mirror tables onto Hugging Face; harvest the portal’s own file list when it has one, and do not open a second harvest of the Hugging Face copy.
Junar (junar)
Dataset API on the tenant, not junar.com marketing.
Keep: tenant dataset API records. Drop: junar.com marketing and individual visualization embeds.
GET https://host/api/v2/datasets
EntryScape (entryscape)
DCAT-AP. Harvest Dataset / dcat:Dataset only. Public DCAT/search API on the tenant.
Keep: dcat:Dataset from the tenant DCAT/search API.
Drop: catalog chrome, documents that are not datasets, and EntryStore admin.
GET https://host/store/search?type=dcat:Dataset
Felles datakatalog (fellesdatakatalog)
Use the public search service. Pages are one-based:
POST https://search.api.fellesdatakatalog.digdir.no/search/datasets
Content-Type: application/json
{"pagination":{"size":1000,"page":1}}
Keep one dataset per hits[] element. Increment pagination.page through
page.totalPages. Keep distributions nested under their dataset; do not emit them
as separate datasets.
For Transportportal, add "profile":"TRANSPORT" to the request body. That profile
is a transport-focused subset of data.norge.no, so deduplicate by canonical dataset
URI if both portals are harvested into one collection. Do not copy the Digdir search
or SPARQL hosts onto tenant catalog links; harvest uses those absolute URLs from
endpoints[].
Keep: dataset hits[] and their distributions. Drop: aggregations, search
suggestions, concepts, information models, services, events, and admin/registration
objects unless the user explicitly requested those resource types.
Piveau (piveau)
DCAT-AP search; skip the hub UI chrome. Keep dcat:Dataset only.
Keep: dcat:Dataset from DCAT-AP search.
Drop: hub UI chrome and harvested copies of catalogs already in this registry.
GET https://host/api/hub/search
GET https://host/sparql
GET https://host/api/sparql
Type /api/hub/search as piveau:search. SPARQL stays sparql.
SPARQL is /sparql or /api/sparql on some hubs.
Idra (idra)
Federation of other catalogs (catalog_type is often Data search engine). Harvesting Idra duplicates member catalogs — prefer harvesting the source portals from this registry unless you need the federation view.
Keep: federation dataset records only if the user asked for the hub view. Drop: member catalogs already in this registry (harvest those origins instead).
GET https://host/Idra/api/v1/catalogues
Type /Idra/api/v1/catalogues as idra:catalogues.
ArcGIS Hub (arcgishub) as open data
GET https://host/api/search/v1
GET https://host/api/feed/dcat-us/1.1.json
GET https://host/data.json
Keep dataset / feature layer items that are public data. Drop StoryMaps, sites, and applications unless you have a separate apps index.
Keep: public dataset / feature layer items. Drop: StoryMaps, Hub sites, and applications unless you have a separate apps index.
Data Fair (datafair)
Koumoul portals. Typical list:
GET https://host/data-fair/api/v1/datasets
Page the JSON dataset collection. Type the catalog API as datafairapi. Drop applications and remote-service catalog chrome. Paths vary — use endpoints[] when present. Do not invent a catalog-level DCAT dump path.
Keep: Data Fair datasets. Drop: applications and remote-service catalog chrome.
Datawheel (datawheel)
Front-end data/economic-complexity sites. There is often no common /api. Harvest a documented JSON/CSV catalog if the portal publishes one; otherwise stop rather than scraping every visualization.
Keep: a documented JSON/CSV catalog if the portal publishes one. Drop: every visualization tile. Stop when there is no catalog API.
GET https://host/api
TriplyDB (triplydb)
GET https://host/_api/facets/datasets
GET https://host/opensearch.xml
OpenSearch description is /opensearch.xml.
Keep datasets, not every named graph or SPARQL binding. One instance, not one graph per harvest record. See harvest-protocols.md.
Keep: TriplyDB datasets. Drop: named graphs and SPARQL bindings as extra datasets.
LKOD (lkod)
Czech local DCAT-AP-CZ. Harvest dcat:Dataset from the municipal LKOD UI/API. Slovak clones often expose Turtle at /opendata/set/lkod:
GET https://host/opendata/set/lkod
Typical catalog links already are /opendata/set/lkod. Cleanup strips /opendata so that dump path attaches at origin and is not doubled.
Harvest that graph, not POMOSAM disclosure pages on the same city. Do not also harvest NKOD or data.slovensko.sk for the same datasets.
Keep: dcat:Dataset from the municipal LKOD UI/API or /opendata/set/lkod. Drop: POMOSAM disclosure pages, NKOD, and data.slovensko.sk copies.
OGD Platform India (ogdindia)
Ministry/state tenants on data.gov.in. Catalog APIs often require a registered key — stop on 401.
Keep: the public CKAN-style or HTML catalog list when it is available without a key. Drop: the national hub when you only needed a tenant, and every dataset behind a developer-key Open API.
data eye (dataeye)
Japanese municipal SaaS. Some tenants speak CKAN-compatible metadata.
GET https://host/api/3/action/status_show
GET https://host/api/3/action/package_search?rows=0
GET https://host/ckan_api/package_search
GET https://host/ckan_api/package_list
Keep: tenant datasets from package_search or the public catalog JSON. Drop: idea-box posts and the vendor homepage. One tenant = one scope (%.dataeye.jp). If status_show 404s, harvest the HTML catalog list only.
LinkData (linkdata)
Single hub at linkdata.org. Harvest published data works, not apps or ideas.
GET https://linkdata.org/
Keep: data works (table title + CSV or RDF download). Drop: App.LinkData applications, Idea.LinkData posts, and Knowledge Connector pages. One hub = one harvest scope. Do not open a second record for a municipality site that only links to a work.
Seoul Open Data Plaza (seoulopendataplaza)
/openinf/ JSP catalogs. Open API developer space often needs a key.
GET https://host/openinf/
Keep: the public dataset listing / sitemap when unauthenticated access exists. Drop: developer-key Open API calls and the Seoul homepage. One gu tenant = one scope.
oPortal (oportal)
Inspur /oportal/ catalogs. There is no verified anonymous default API on every tenant.
GET https://host/oportal/
Keep: /oportal/ dataset listing or DCAT if public. Drop: the application gallery and login-only 数据开放 admin.
Epoint Big Data Open Platform (epointopendata)
Public catalog pages under /extranet/openportal/. No verified anonymous list API on every tenant.
GET https://host/extranet/openportal/pages/default/index.html
Keep: the public dataset or catalog listing for that government tenant. Drop: login-only申请 workflows, the API application gallery, and /oportal/ Inspur tenants.
Taiji Digital Public Data Open Platform (tykyopendata)
Taiji Digital shells. No verified anonymous list API on every tenant. Shenzhen serves the catalog UI at the site root.
GET https://host/
Keep: the public dataset or catalog listing for that government. Drop: login-only 数据申请 workflows, the developer center, and Inspur /oportal/ or Epoint /extranet/openportal/ tenants.
data.world (dataworld)
Public hub data.world. Harvest datasets from the public catalog/API. Do not crawl private organization spaces.
GET https://api.data.world/v0/datasets/search?q=*
Stop on 401. Keep: public datasets. Drop: per-user files as crawl seeds. One registered hub.
IUDX Catalogue (iudx)
IUDX catalogue server. Harvest items from the catalogue API.
GET https://central-catalogue.iudx.org.in/iudx/cat/v1/search
Keep: catalogue items. Drop: each resource ID as a crawl seed. One harvest scope per public tenant.
Liferay (liferay)
Spanish RISP / datos abiertos modules on Liferay. Harvest only the dataset list (often a JSON/CSV/XML table or /documents/ open-data folder).
Keep: rows that are datasets (title + landing or file URL). Drop: /web/guest/ CMS homepages, news, and generic document libraries.
If the list is HTML-only with no machine table, stop. Do not crawl every Liferay page.
POMOSAM (pomosam)
Slovak municipal disclosure platforms.
Keep: the open-data / dataset module. Drop: procurement and contracts-only pages unless that module is the data catalog. If the list is HTML-only with no machine table, stop.
SEU-e (seue)
seu-e.cat e-office tenants.
Keep: dades obertes listings. Drop: the rest of the electronic office (procedures, notifications, sede chrome). If the list is HTML-only with no machine table, stop.
ATM Maggioli (atmmaggioli)
Spanish sede electrónica tenants. Harvest the dataset list on /transparencia/datos/catalogo only.
GET https://host/transparencia/datos/catalogo
Keep: rows that are datasets (title + landing or file URL). Drop: the rest of the e-office, transparency obligation pages, and guessed CKAN/OpenDataSoft/Socrata paths (they are HTML). If the list is HTML-only with no machine table, stop. One municipality tenant = one harvest scope.
Gobierto Datos (gobierto)
Populate Gobierto dataset hub. Harvest the /datos catalog, not the budget or contracts modules.
GET https://host/datos
Keep: dataset pages under /datos/{slug} (title, metadata, file or SQL API link). Drop: /presupuestos, contracts, agendas, planes, gobierto.es marketing, and presupuestos.gobierto.es. One municipality host = one harvest scope. Do not set ckan.
Viavansi Open Government (viavansi)
WordPress theme viavansi-open-government or viavansi-ogov-current. Harvest the public dataset list on the catalog page (title and file or DCAT link). One harvest scope per institution. Do not set ckan unless the public UI is the CKAN catalog. Do not set wordpress when this theme is present.
Keep: dataset rows on the catalog page. Drop: theme assets, the WordPress admin, and login walls.
Municipium Portale Opendata (municipium)
Maggioli Municipium tenants (Italian comuni). Harvest the catalog page on the tenant host, not the shared API host.
GET https://host/it/page/catalogo
Keep: dataset entries listed on the catalog page (title + landing page). Drop: informative pages (/it/page/informazioni-*), the DATI.GOV national catalog links, siti tematici navigation, and the vendor demo opendata.municipiumapp.it. The {tenant}-opendata-api.cloud.municipiumapp.it backend serves assets and has no documented public dataset API — do not guess endpoints there. Do not harvest opendata.maggioli.cloud organization slices here; that hub is the shared CKAN catalog. One comune tenant = one harvest scope.
OPENDATAENTE (opendataente)
Italian Actainfo municipal SaaS. Harvest dcat:Dataset from the tenant DCAT-AP_IT catalog, not the React theme tiles.
GET https://host/backend/api/catalog/
Keep: dcatapit:Dataset / dcat:Dataset in that RDF. Drop: dcat:Distribution files, the vendor sites opendataente.it / opendataente.cloud, ActaLogin hosts, and dati.gov.it copies of the same datasets. One tenant = one harvest scope.
DataPortal.AI (dataportalai)
Sister StatPortal Open Data / DataPortal.AI portals. Harvest the catalog index and dataset pages, not the vendor site. Fingerprints: discovery-opendata.md.
StatPortal Open Data / SPOD (Drupal):
GET https://host/catalogo-opendata
GET https://host/catalog
GET https://host/opendata/{slug}
The catalog path is /catalogo-opendata (Pisa) or /catalog (Veneto). Dataset records live under /opendata/{slug}.
DataPortal.AI SPA (dati.lavoro.gov.it): the HTML shell is the same document for /, /catalogo-opendata/, and /content/. /config/config.json sets baseURL to /api/core/ and names /api/odata/. Harvest a dataset list from those APIs only after a GET returns dataset records.
Keep: dataset entries (title + /opendata/{slug} landing page + file link). Drop: /content/ CMS pages, ?t=Tabella / ?t=Grafico / ?t=Scarica views of the same dataset, sister.it and almawave.com marketing, the Istat browser statportal.it, Sister Data Browser (title Data Browser), StatKit / ASTATDATA, and dati.gov.it copies of the same datasets. One administration portal = one harvest scope. Do not guess /api/3. A host whose /api/3/action/status_show returns ckan_version is CKAN (dati.unionevallesavio.it).
ComunWeb (comunweb)
Trentino ComunWeb sites. Harvest the dataset class, not the rest of the municipal CMS.
GET https://host/api/opendata/v1/content/class/opendata_dataset
GET https://host/exportas/csv/opendata_dataset
Keep: nodes whose classIdentifier is opendata_dataset (title objectName, landing page fullUrl). Drop: other classes from /api/opendata/v1/content/classList (albo, news, banners, comments), comunweb.it marketing, and login pages. A ckan_ prefix on objectRemoteId does not make the host CKAN. One ente host = one harvest scope.
PA-Online Open Data (paonline)
Technical Design RDF catalogs on pa-online.it.
GET https://www.pa-online.it/OpenData/{istat}/METADATO.rdf
The file name is sometimes the comune name instead of METADATO.rdf. Keep: dcatapit:Dataset records in that RDF. Drop: hosting.pa-online.it sportello pages, GisMasterWebS SUAP and payment screens, and GisMaster map viewers. One ISTAT directory = one harvest scope.
OpenGov (opengov)
US {org}.opengov.com financial transparency.
GET https://host/transparency
GET https://host/data/
Keep: named reports / stories on /transparency or /data/. Drop: every Highcharts breakdown row and www.opengov.com marketing. One government tenant = one harvest scope. Distinct from the city’s CKAN/Socrata/ArcGIS Hub open-data portal.
OneGov Election Day (onegov)
Swiss election and vote portals. Filter exports on software.id = 'onegov'. Distinct from Tyler OpenGov (opengov).
GET https://host/catalog.rdf
GET https://host/json
GET https://host/archive/{year}/json
GET https://host/vote/{id}/data-json
GET https://host/election/{id}/data-json
Keep: each election or vote as a dataset, with the JSON/CSV/Excel download linked from the catalog. Drop: the portal chrome, map tiles, and PDF-only result sheets when a machine-readable file exists. One portal host = one harvest scope. Discovery: discovery-opendata.md.
Ísland.is (islandis)
Icelandic government content catalogs published as Ísland.is applications (Opin gögn; the Stjórnartíðindi official gazette is out of scope — documents, not datasets). Filter exports on software.id = 'islandis'.
GET https://island.is/s/stafraent-island/opingogn
GET https://api.island.is/graphql
The GraphQL gateway at api.island.is is authentication-gated for most queries; public catalog pages are server-rendered and list documents as HTML. Keep: each listed dataset as a record, with the linked JSON/CSV download. Drop: portal navigation, service pages, gazette issues, and ministry profiles. One island.is catalog page = one harvest scope. Discovery: discovery-opendata.md.
Drupal (drupal)
Only when the public product is a dataset catalog (not a news CMS).
GET https://host/jsonapi/node/dataset
GET https://host/jsonapi/node/open_data
GET https://host/jsonapi/node/ckan_dataset
GET https://host/data.json
Bundle names vary (dataset, open_data, ckan_dataset). Inspect /jsonapi once for dataset-like bundles. Keep: those dataset nodes. Drop: article, page, media, and user accounts.
DKAN on Drupal: use the DKAN Action API when enabled — prefer dkan as software.id. EKAN: use the EKAN JSON:API dataset list — prefer ekan.
WordPress (wordpress)
GET https://host/wp-json/
GET https://host/wp-json/wp/v2/dataset
Joomla (joomla)
Joomla CMS sites have no standard dataset list API. Harvest the catalog or map-gallery pages as HTML lists; keep each linked dataset or map as one record and drop article chrome, menus, and login pages.
OutSystems (outsystems)
OutSystems apps expose no standard dataset API. Harvest the catalog HTML or the per-app REST endpoints the operator documents. Keep: dataset, indicator, and legislation pages. Drop: login/session chrome and reactive app-shell routes.
Plone (plone)
Plone has no standard dataset API. Harvest the catalog HTML. Keep: linked microdata files, study pages, and dataset downloads. Drop: news, navigation, and login chrome. gov.br agency sites share one Plone installation; record each catalog page, not the ministry home.
/wp-json/ only for a datasets custom post type (/wp-json/wp/v2/dataset or the type the catalog documents). Keep: dataset posts. Drop: /wp/v2/posts, media, and ordinary WordPress homepages.
Government Site Builder (governmentsitebuilder)
GSB has no standard dataset API. Harvest the catalog pages as HTML lists. Keep: linked datasets, statistics tables, and download files. Drop: news, navigation, and ordinary agency pages. GSB sites often embed Apache Solr search forms; treat those as HTML search UI, not an API.
Bitrix (bitrix)
1C-Bitrix government portals. Harvest only a published open-data / dataset module (JSON/CSV/DCAT list). Skip the rest of the CMS, news, and /bitrix/admin/.
Keep: records from a published open-data / dataset module (JSON, CSV, or DCAT list).
Drop: CMS news, /bitrix/admin/, and the rest of the government site.
GET https://host/opendata/
GET https://host/opendata/opendata.json
Type /opendata/ as bitrix:catalog and /opendata/opendata.json as opendata:json when present. Do not append /opendata/ onto catalog links that already are the open-data page. Cleanup strips /opendata (keeping any locale prefix such as /ru) so that JSON and HTML paths attach at origin. Catalog pages without /opendata (.php / .aspx) fall back to origin so /opendata/ is not concatenated onto the filename.
Gosweb (gosweb)
Filter exports on software.id = 'gosweb'. One harvest scope per municipal or agency site.
Harvest the open-data page /ofitsialno/statistika/otkrytye-dannye/ and the machine-readable register (list.csv, meta.csv) linked from it. Grain is the dataset passport, not each CSV column or the rest of the municipal website.
Keep: dataset passports and their download files. Drop: news, service pages, and other sections of the Gosweb site.
GET https://host/ofitsialno/statistika/otkrytye-dannye/
GET https://host/ofitsialno/statistika/otkrytye-dannye/list.csv
DataPress (datapress)
CKAN plus CMS. Harvest package_search as in CKAN. Do not harvest CMS pages or double-count the host as ckan.
Keep: CKAN packages via package_search (same grain as CKAN).
Drop: CMS pages. Do not double-count the host as ckan.
GET https://host/api/3/action/package_search?rows=25
GET https://host/api/3/action/package_list
Optional SPARQL dump when present in endpoints[] (/sparql, same grain as CKAN): a catalog dump, not a substitute for packages.
Our Open Data (ouropendata)
Japanese Our Open Data / SHIRASAGI catalog.
GET https://host/api/package_list
Keep: datasets from /api/package_list or the catalog home / numeric dataset list. Drop: idea-box posts. Type /api/package_list as ouropendata:packages. Do not treat it as CKAN.
Gipuzkoa Irekia (gipuzkoairekia)
Tenant DCAT. Keep datasets, not the rest of the Irekia CMS.
Subdomain tenants ({town}.gipuzkoairekia.eus/es/datu-irekien-katalogoa) expose dumps at the origin (/catalog.xml, /catalog.rdf, /catalog.jsonld, /api/feed/dcat). Catalogs on www.gipuzkoairekia.eus/es/web/{tenant}/datu-irekien-katalogoa keep the tenant path ({link}/catalogo.rdf). Do not copy the provincial hub origin dump onto those www tenants.
Keep: tenant DCAT datasets. Drop: the rest of the Irekia CMS. Do not treat guessed CKAN / Socrata / OpenDataSoft paths on the Irekia HTML shell as harvest APIs.
GET https://host/catalogo.rdf
GET https://host/catalog.xml
GET https://host/catalog.rdf
GET https://host/catalog.jsonld
GET https://host/api/feed/dcat
Open Data Euskadi (opendataeuskadi)
Filter exports on software.id = 'opendataeuskadi'. One harvest scope: the Basque Government hub.
GET https://opendata.euskadi.eus/catalogo-datos/
The catalog is paginated HTML on the euskadi.eus stack; dataset pages carry download resources (CSV/JSON/XML/RDF) and the portal federates DCAT metadata to datos.gob.es. Keep datasets from the catalog listing. Drop euskadi.eus CMS chrome, news, and the geoEuskadi / Udalmap products (separate catalogs).
Keep: catalog datasets and their download resources. Drop: CMS chrome, news, geoEuskadi and Udalmap content.
MODA (modaopendata)
Tenant catalog API, not a second national data.gov.tw clone.
Keep: tenant catalog API datasets. Drop: a second copy of national data.gov.tw.
GET https://host/api/v2/rest/dataset/od{limit}
Taiwan Government Website Open Data (twgovopendata)
ASP.NET dataset list on a county or city government host. Not the MODA Nuxt catalog.
GET https://host/OpenDataList.aspx
GET https://host/opendata/OpenDataList.aspx
Keep: one dataset per OpenDataDetail.aspx or OpenDataContent.aspx page, including its OpenDataFileHit.ashx files. Drop: the county homepage, news, and accessibility chrome. A Default.aspx open-data section without those pages stays custom.
PublishMyData (publishmydata)
Linked-data publishing (Swirrl).
GET https://host/data.json
Keep: the DCAT dataset list or SPARQL that the catalog documents — named datasets. Drop: every triple in the graph (harvest-protocols.md).
data.gov.my (datagovmy)
Malaysia national / agency tenants on the data.gov.my stack.
GET https://api.data.gov.my/data-catalogue?id=fuelprice
id is required. That response is observations for one catalogue dataset — treat each id as one dataset analog; do not page rows as datasets. There is no verified anonymous GET that lists all catalogue ids; collect ids from the public catalogue UI or developer.data.gov.my. Drop dashboards and documentation pages. Agency sites (open.dosm.gov.my, data.moh.gov.my) are separate registry rows — harvest each tenant once.
Keep: each catalogue id as one dataset analog. Drop: observation rows, dashboards, and documentation pages.
JDOP (jdop)
Zhejiang public-data open platform (/jdop_front/, /dopServer/).
GET https://host/jdop_front/
Keep: the tenant dataset catalog API when public. Drop: the Zhejiang government homepage and login-only 数据开放 admin.
Open Data Registry (opendatareg)
AWS Open Data Registry style catalogs (catalog.json / YAML dataset files, optional STAC). Keep dataset entries. Drop bucket listings and every STAC item. Skip cloning registry.opendata.aws if you only needed the existing registry row.
GET https://host/catalog.json
Keep: dataset entries. Drop: bucket listings and every STAC item.
D4Science (d4science)
Keep: public catalogue items on the VRE (gCat / documented dataset list). Drop: workspace files, private VREs, and d4science.org marketing. Stop on 401.
Semantic MediaWiki (smw)
GET https://host/w/api.php?action=ask&query=[[Category:Dataset]]
Keep pages typed as Dataset (or the site’s equivalent category). Drop ordinary wiki articles. api.php without a dataset query is not a harvest.
Probe /w/api.php?action=ask&query=[[Category:Dataset]]&format=json (or /api.php without /w/) as smw:ask.
Keep: pages typed as Dataset (or the site’s equivalent category). Drop: ordinary wiki articles and unfiltered api.php.
Strapi (strapi)
GET https://host/api/datasets
Public dataset content-type REST only (/api/datasets or the type the catalog documents). Keep: dataset entries. Drop: posts, users, and admin.
Tablion (tablion)
Aristotle’s data-portal product.
Keep: the public dataset API. Drop: every MDR object (concepts, classifications, quality statements) unless the user asked for metadata catalogs (harvest-metadata.md).
Copernicus CDS (copernicuscds): harvest-earthdata.md. Discovery fingerprints: discovery-opendata.md.
RDF Online Repository (rdfrepository)
Public license/workspace tables on *.revenuedev.org. Filter exports on software.id = 'rdfrepository'.
GET https://host/
Keep: the published dataset / license list if unauthenticated. Drop: login-only company filing modules. One harvest scope per country tenant. Distinct from W3C RDF.
ResourceContracts (resourcecontracts)
GET https://host/contract/resources
Keep contracts (documents). Drop the vendor About page and per-clause annotation chrome unless that is the catalog. One hub or country tenant = one harvest scope.
Keep: ResourceContracts contracts. Drop: vendor About page and per-clause chrome.
OpenSpending (openspending)
GET https://openspending.org/
Keep budget / fiscal datasets (Fiscal Data Packages). Drop individual treemap visualizations and the OKF About page. One harvest scope per OpenSpending hub or independent deployment.
Keep: budget / fiscal datasets. Drop: treemap visualizations and the OKF About page.
ODWeb (odweb)
Public dataset list under /odweb/. Filter exports on software.id = 'odweb'.
Keep open datasets for that tenant. Drop login-only apply flows and the parent government homepage. One /odweb/ host = one harvest scope. Not CKAN; not oPortal; not JDOP.
GET https://host/odweb/
Keep: ODWeb /odweb/ datasets. Drop: login-only apply flows and the parent government homepage.
Guangxi Public Data Open Platform (gxopendata)
Public dataset / directory list on data.gxzf.gov.cn or {city}.data.gxzf.gov.cn. Filter exports on software.id = 'gxopendata'.
Keep open datasets for that tenant. Drop login-only apply/API-gateway flows. One tenant = one harvest scope. Not CKAN.
Keep: Guangxi tenant open datasets. Drop: login-only apply/API-gateway flows.
OpenGDC (opengdc)
Dutch municipal catalog. Harvest datasets only:
GET https://host/api/datasets
GET https://host/openapi.json
JSON:API collection (data[].type == "dataset", meta.total). Page with the API’s start/rows (or follow links). Keep: dataset objects. Drop: /api/documents and /api/dossiers (Woo files), CMS pages, and idea boxes. One municipality tenant = one harvest scope. Some hosts return 403 from a WAF — stop; do not scrape HTML as a substitute.
Portals without a dataset API
Liferay, POMOSAM, ATM Maggioli, oPortal, OGD India, Seoul plaza, Drupal, and WordPress are covered above when a list exists. If there is still no machine-readable catalog, stop. Generic DCAT paths: /catalog.xml, /data.json (harvest-protocols.md).
GIS Open Data Portal (gisopendataportal)
Start at /api/opendata/set/catalog/lkod: the DCAT-AP-SK JSON-LD catalog contains a
dataset array of metadata URLs under /api/opendata/set/{uuid}. Keep one record per
dataset URL, retaining the source IRI, publisher, description, and distributions. Resolve
its declared JSON-LD context; do not assume English JSON property names. Follow dataset
metadata to CSV downloads. Drop site navigation and individual spatial features from the
catalog inventory.
GET https://host/api/opendata/set/catalog/lkod
Keep: one DCAT-AP-SK dataset per metadata URL. Drop: site navigation and individual spatial features. Feature retrieval is a separate operation at
/api/open-api/features?limit=1&page=1, with filter.featureClass.id or
filter.featureClass.code; coordinates may use EPSG:5514. GraphQL is documented at
/api/open-api. Both municipal catalog responses were verified on 2026-09-07.
Use Tvrdošín or Nové Mesto for the deployment's API contract. Tvrdošín's catalog currently copies Nové Mesto's title: retain provenance and do not infer publisher identity from that title alone.
Esri UK Data Observatory (esridataobservatory)
Start at the deployment's linked Data Explorer page (for example, the
Suffolk explorer). Read its published
application configuration to resolve the ArcGIS data catalog and backing service URLs;
there is no assumed universal /api/3/action or DCAT endpoint. The
vendor embedding guide
documents dataCatalogExplorer.launch with an ArcGIS application ID. Keep discoverable
source datasets or indicator tables and their metadata; drop WordPress posts, ward-profile
pages, rendered charts, and per-area observation rows from the catalog inventory.
Keep: discoverable source datasets or indicator tables. Drop: WordPress posts, ward-profile pages, rendered charts, and per-area observation rows. Deduplicate repeated references to the same source item/table across reports. Use only publicly accessible services, retaining source IDs and licensing information. API URLs and access vary by deployment; a single standalone harvest endpoint was not verified in this review.
RUDI (rudi)
Use the deployment's documented API.
Portal metadata search uses /konsult/v1/datasets/metadatas; the documented anonymous
session requires a token from /authenticate. Follow the published public-access flow
and preserve access restrictions. Producer nodes instead expose /api/v1/resources
and /api/v1/resources/{id}, as documented in the
node catalog source. Do not assume
node routes exist on the portal host. Keep one metadata record per global_id, follow
pagination and producer provenance, and distinguish catalog visibility from permission to
retrieve the underlying data.
Keep: one RUDI metadata record per global_id. Drop: guessed node /api/v1/resources on the portal host, and login-only search. No anonymous API access was verified on the registered
Rennes portal in this review; its existing access fields are preserved.
SIMAI Open Data Portal (simaiopendata)
Start from the deployment's /datasets/ directory and follow dataset detail links.
Keep one dataset passport with its publisher, description and linked downloads; exclude
category pages, organization indexes, news and individual table rows.
GET https://host/datasets/
Keep: one SIMAI dataset passport per /datasets/ detail page. Drop: category pages, organization indexes, news, and individual table rows. Consult the
product manual. No stable
public metadata API was verified; do not infer one from the underlying Bitrix CMS.
The vendor demo contains demonstration data
and must remain distinguishable from production holdings.
Aid Management Platform (amp)
Country AMP portal (often /portal/). Harvest the public activity / project list if unauthenticated. One harvest scope per country installation, not per report or chart.
Keep: public aid activities / projects. Drop: news, login walls, and individual PDF reports.
Het Dataloket (dataloket)
Dutch Analyze data-asset catalog. Harvest public search hits, not the Vue shell or the embedded Kibana dashboard.
GET https://host/api/search
GET https://host/api/search?offset=10
Keep: each Results item with access Openbaar (contentUID, title, description, contentTypeName, path). Asset types include Dataset, Kaart, Dashboard, Rapportage, and Document. Drop: FAQ, /datavraag, /toegang, the portal Kibana embed, and rows that are not public. offset=0 and offset=10 each returned 10 rows on Venlo, Tilburg, and Limburg (2026-09-26); offset=20 returned an empty Results list while Total stayed larger. Do not assume further pages. One tenant host = one harvest scope. Do not harvest dataportaal-viewer.prvlimburg.nl or ckan.dataplatform.nl as Dataloket.
Contrataciones Abiertas (contratacionesabiertas)
INAI EDCA-MX dashboard. Read globals.site.url and globals.site.port from /contratacionesabiertas/static/javascripts/common.js.
GET {url}:{port}/edca/fiscalYears
GET {url}:{port}/edca/recordPackages/{year}
fiscalYears[].year with status: true are the years to request. Each recordPackages[] item is one OCDS record package (one contracting process). The datos abiertos page also documents /edca/contractingprocess/{year} and /edcapi/project/; use those only when they return JSON. On 27 September 2026 the same paths on port 443 of the dashboard host returned 404, while Universidad Veracruzana :8080 and Yucatán captura.contratacionesabiertas.inaipyucatan.org.mx returned record packages. INFO CDMX port 3000 timed out. If the capture host does not respond, stop.
Keep: OCDS record packages. Drop: fiscalYears admin rows, the capture UI, Excel sheet splits of one process, and /contratacionesabiertas/implementa. Do not harvest Peru OECE or a CKAN portal on the same institution as this dashboard.
Centurion (centurion)
eBdesk Centurion catalog. Harvest dataset rows from PostGraphile when the tenant exposes it.
POST https://host/api/v1/graphql
{"query":"{ metadata_datasets(first: 100, offset: 0) { totalCount nodes { id title reference_code status } } }"}
Page with offset or the after cursor. Keep: metadata_datasets nodes (id, title, reference_code, status). Drop: users, roles, auth, surveys, publications, and geospatial layers as extra datasets. Kalimantan Selatan data., opendata., and satupeta. return the same totalCount; harvest one host. satudata.kalselprov.go.id requires a token. Polri and Sumatera Utara returned HTTP 405 on this path (September 2026); do not invent another list URL for those shells.
CreatorCMS (creatorcms)
Hunan municipal data-open module. HTML listing only; no public JSON API confirmed.
GET https://host/webapp/{city}/dataPublic/index.jsp
GET https://host/webapp/{city}/dataPublic/dataDetail.jsp?id={n}
Keep: each dataDetail.jsp?id= record as one dataset. Drop: the parent government homepage, news/articles served by the same CMS, and login-only admin paths.
Anhui open-data-web (ahopendataweb)
Anhui public-data platform. Catalog HTML under /open-data-web/ (.do actions) or /dataopen-web/; no public JSON API confirmed.
GET https://host/open-data-web/index/index.do
Keep: open dataset entries in the catalog list. Drop: login-only 数据申请 apply flows, the parent government homepage, and app-gallery entries.
openportal (openportal)
Municipal open-data portal under /extranet/openportal/pages/.... HTML listing; no public JSON API confirmed.
GET https://host/extranet/openportal/pages/default/index.html
Keep: catalog dataset entries. Drop: the parent government homepage and login-only apply flows.
Oraș Digital (orasdigital)
Romanian city open-data SaaS ({city}.oras.digital). REST JSON API documented per deployment; grab the Postman collection for the method list.
GET https://{city}.oras.digital/api/
GET https://{city}.oras.digital/assets/api/api-postman-collection.json
Keep: dataset and resource records from /api/ (CSV/XLS/XLSX/HTML/API resources). Drop: {city}.digital city-app pages, terms/cookie pages, and user-account endpoints.
Bon Maximus e-Procurement (bonmaximus)
Nigerian state e-procurement portals (Bon Maximus Companies). No documented public API — harvest the OCDS-oriented publication and award pages from the HTML portal.
GET https://host/
GET https://host/publication.php
Keep: tender/award records and OCDS publication pages (title, buyer, award value, date, linked OCDS JSON when offered). Drop: vendor pages (bonmaximus.com), login/registration flows, and state government homepages. One state BPP portal = one harvest scope. Do not set ckan or budeshi.
Budeshi (budeshi)
PPDC open contracting platform deployments (e.g. Kaduna www.ocds.kdsg.gov.ng). Documented REST API serving OCDS releases.
GET https://host/api
GET https://host/ocds-api
Keep: OCDS releases from /api and /ocds-api (JSON packages; flatten releases to tender/award records). Drop: portal chrome, project marketing pages, and the platform home budeshi.ng. One deployment = one harvest scope. Do not set bonmaximus on Budeshi-hosted portals.
MapaInversiones (mapainversiones)
IDB public-investment transparency deployments. No documented public dataset API — harvest the national "Datos Abiertos" download section (CSV/Excel/OCDS files) and the public investment project list as the dataset grain.
Keep: open-data downloads (projects, budgets, OCDS procurement files). Drop: map chrome, news, and citizen-participation pages. One national deployment = one harvest scope.