package
Generates, extends, and validates Frictionless Data Package descriptors (datapackage.json) from one or more data files. Supports optional package metadata, schema inference, and AI-powered metadata generation with --autodoc.
# Create datapackage.json for a single file
undatum package create data.csv --output datapackage.json
undatum package create workbook.xlsx --table Sheet2 --output datapackage.json
undatum package create nested.jsonl --flatten-nested --output datapackage.json
# Create a package directory with data file copies
undatum package create data.csv --package-dir out/package
# Zip the materialized package directory
undatum package create data.csv --package-dir out/package --zip out/package.zip
# Add another resource to an existing package
undatum package add-resource out/package/datapackage.json new.csv
# Validate a package descriptor
undatum package validate out/package/datapackage.json
# Provide metadata and enable AI metadata generation
undatum package create data.csv --title "Sales data" --keywords sales,finance \
--autodoc --ai-provider openai --ai-model gpt-4o-mini
Subcommands:
create— generate a new descriptor (default workflow)add-resource— append resources to an existing descriptorvalidate— validate descriptor structure (full checks withpip install undatum[frictionless])
Metadata options:
--name,--title,--description,--keywords--licenses(semicolon-separated entries, e.g.name=MIT;name=ODC-PDDL-1.0)--sources(semicolon-separated entries, e.g.title=World Bank,path=https://...)--contributors(semicolon-separated entries, e.g.title=Jane Doe,email=jane@example.com)--version- Package version string
Features:
- Frictionless profile: Emits
profile: tabular-data-packagewith resourceformat/mediatype - Schema inference: Automatically infers field types, descriptions, and uniqueness constraints
- Multiple resources: Package multiple files as separate resources
- Remote URIs: Support for HTTP/HTTPS URLs as resource paths
- Package directory: Bundle
datapackage.jsonwith data file copies - AI metadata: Use
--autodocto generate metadata with AI assistance (single-pass, no duplicate LLM calls) - Streaming-safe: Processes large datasets without loading everything into memory
- Python SDK:
Dataset.read("data.csv").package(output="datapackage.json")
Additional options:
--package-dir: Create a package directory with data file copies--zip: Create a ZIP archive of the package directory (requires--package-dir)--autodoc: Enable AI-powered metadata generation (reusesdoccommand logic)--engine: Processing engine (autoorduckdb)--delimiter,--encoding,--tagname,--start-line,--start-page: Passed through to analysis and sampling--limit: Maximum objects to analyze for schema inference (default: 10000)--sample-size: Number of sample records for metadata inference (default: 10)
Reference
undatum package create
undatum package create [OPTIONS] INPUT_FILES...
| Argument | Description |
|---|---|
INPUT_FILES... | Input file(s) to package. (required) |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Output datapackage.json path. | |
--package-dir TEXT | Package directory to materialize. | |
--name TEXT | Package name (slug). | |
--title TEXT | Package title. | |
--description TEXT | Package description. | |
--keywords TEXT | Comma-separated keywords. | |
--licenses TEXT | Licenses (semicolon-separated entries, e.g. 'name=MIT;name=ODC-PDDL-1.0'). | |
--sources TEXT | Sources (semicolon-separated entries, e.g. 'title=World Bank,path=https://...'). | |
--contributors TEXT | Contributors (semicolon-separated entries, e.g. 'title=Jane Doe,email=jane@example.com'). | |
--version TEXT | Package version string. | |
--sample-size INTEGER | Number of sample records to include in metadata inference. | 10 |
-n, --limit INTEGER | Maximum number of objects to analyze for schema inference. | 10000 |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | |
--tagname TEXT | XML tag name that contains individual records. | |
--start-line INTEGER | Line number (0-based) to start reading from. | 0 |
--start-page INTEGER | Page number (0-based) to start from for Excel files. | 0 |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl', 'xml'). | |
--autodoc / --no-autodoc | Enable AI-powered metadata generation. | --no-autodoc |
--lang TEXT | Language for AI-generated metadata (default: 'English'). | English |
--ai-provider TEXT | AI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml). | |
--ai-model TEXT | Model name to use (provider-specific). | |
--ai-base-url TEXT | Base URL for AI API (optional). | |
--pii-mask-samples / --no-pii-mask-samples | Mask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio. | |
--zip TEXT | Create a ZIP archive of the package directory (requires --package-dir). | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
Deprecated spellings (removed in 2.0): --objects-limit → --limit.
undatum package add-resource
undatum package add-resource [OPTIONS] PACKAGE_FILE INPUT_FILES...
| Argument | Description |
|---|---|
PACKAGE_FILE | Existing datapackage.json to extend. (required) |
INPUT_FILES... | Input file(s) to add. (required) |
| Option | Description | Default |
|---|---|---|
--package-dir TEXT | Package directory containing data files (defaults to descriptor dir). | |
--sample-size INTEGER | Sample size for metadata inference. | 10 |
-n, --limit INTEGER | Maximum objects to analyze. | 10000 |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding. | |
--tagname TEXT | XML record tag name. | |
--start-line INTEGER | Line number (0-based) to start reading from. | 0 |
--start-page INTEGER | Excel start page (0-based). | 0 |
-F, --format-in TEXT | Override input format. | |
--autodoc / --no-autodoc | Enable AI-powered metadata generation. | --no-autodoc |
--lang TEXT | Language for AI metadata. | English |
--ai-provider TEXT | AI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml). | |
--ai-model TEXT | AI model name. | |
--ai-base-url TEXT | AI API base URL. | |
--pii-mask-samples / --no-pii-mask-samples | Mask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
Deprecated spellings (removed in 2.0): --objects-limit → --limit.
undatum package validate
undatum package validate [OPTIONS] PACKAGE_FILE
| Argument | Description |
|---|---|
PACKAGE_FILE | Path to datapackage.json. (required) |
| Option | Description | Default |
|---|---|---|
--limit-rows INTEGER | Limit rows validated per resource. | |
--check-data / --no-check-data | Validate resource data in addition to metadata. | --check-data |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
See also shared options.