Skip to main content

Exports

Generated artifacts live in data/datasets/. Rebuild with python scripts/builder.py build. Never hand-edit this directory.

Primary dumps​

FileContents
catalogs.jsonl.zstVerified entities only (uncompressed JSONL is not kept; it exceeds GitHub size limits)
scheduled.jsonl (+ .zst)Unverified scheduled records (may be empty)
full.jsonl (+ .zst)Entities + scheduled
software.jsonl (+ .zst)Software / platform definitions
full.parquetAnalytics table of full.jsonl
datasets.duckdbTables catalogs and software
catalogs.jsonldOptional; build --jsonld

Record counts​

Do not mix these three numbers. Exports lag YAML until python scripts/builder.py build.

LayerDateCatalogsScheduledSoftwareCountries
Published GitHub snapshotv1.23.0, 30 September 202643,04749784225
Working-tree exportslast build in this tree (30 September 2026)43,047 (catalogs.jsonl.zst)49784225
Current source YAML30 September 202643,047 (data/entities/)49 (data/scheduled/)784225

Working-tree exports match entity YAML (43,047 catalogs, 49 scheduled, 784 software). full.jsonl is 43,096 (entities plus the scheduled records). Canonical software IDs: data/reference/software_ids.yaml.

Filter by catalog type or software in DuckDB / Parquet (see query-examples.md); there are no pre-sliced bytype/ or bysoftware/ dumps.

Compression​

.zst files are zstandard. Decompress with unzstd file.zst or stream them in Python via zstandard. Verified catalog records are published only as catalogs.jsonl.zst; python scripts/builder.py build writes uncompressed catalogs.jsonl as a temporary merge input and then deletes it.

DuckDB columns​

datasets.duckdb table catalogs (same nested types as full.parquet). Lists are DuckDB LIST (VARCHAR[] or STRUCT[]); objects are STRUCT. Query with field access (software.id) and list functions (list_contains, unnest, list_transform) rather than LIKE on JSON text. Heterogeneous Re3Data leaves may remain JSON. api is BOOLEAN.

ColumnJSONL typeDuckDB
id, uid, name, link, catalog_type, status, api_status, description, catalog_exportstringVARCHAR
apibooleanBOOLEAN
access_mode, content_types, tagslist of stringsVARCHAR[]
langs, topics, identifiers, endpoints, coveragelist of objectsSTRUCT[]
software, owner, rights, propertiesobjectSTRUCT
_re3dataobjectSTRUCT (some nested leaves JSON)

Table software keeps Yes/No flags (has_api, has_bulk) as VARCHAR. Nested datatypes, metadata_support, owner, license, pid_support, and rights_management are STRUCT; export_formats and capabilities are VARCHAR[].

trust_score is optional on YAML and may be absent from a given DuckDB build if no records in the snapshot have the field.

JSON-LD / DCAT​

data/schemes/catalog.context.jsonld maps fields to DCAT-AP, Dublin Core, schema.org, and the cdi: namespace. Emit a framed dump with:

python scripts/builder.py build --jsonld
Catalog fieldJSON-LD term
(type)dcat:DataCatalog
namedct:title
descriptiondct:description
linkdcat:landingPage
ownerdct:publisher
rightsdct:rights
access_modedct:accessRights
identifierdct:identifier
idcdi:id
uidcdi:uid
catalog_typecdi:catalogType
statuscdi:status
softwarecdi:software
coveragecdi:coverage
endpointscdi:endpoints
identifierscdi:identifiers
apicdi:hasApi
api_statuscdi:apiStatus
tagscdi:tags
topicscdi:topics
langscdi:langs
content_typescdi:contentTypes
trust_scorecdi:trustScore
propertiescdi:properties
catalog_exportcdi:catalogExport
_re3datacdi:re3dataEnrichment

cdi: is https://commondata.io/ns/dataportals-registry#.