Skip to main content

package

Generates, extends, and validates Frictionless Data Package descriptors (datapackage.json) from one or more data files. Supports optional package metadata, schema inference, and AI-powered metadata generation with --autodoc.

# Create datapackage.json for a single file
undatum package create data.csv --output datapackage.json
undatum package create workbook.xlsx --table Sheet2 --output datapackage.json
undatum package create nested.jsonl --flatten-nested --output datapackage.json

# Create a package directory with data file copies
undatum package create data.csv --package-dir out/package

# Zip the materialized package directory
undatum package create data.csv --package-dir out/package --zip out/package.zip

# Add another resource to an existing package
undatum package add-resource out/package/datapackage.json new.csv

# Validate a package descriptor
undatum package validate out/package/datapackage.json

# Provide metadata and enable AI metadata generation
undatum package create data.csv --title "Sales data" --keywords sales,finance \
--autodoc --ai-provider openai --ai-model gpt-4o-mini

Subcommands:

  • create — generate a new descriptor (default workflow)
  • add-resource — append resources to an existing descriptor
  • validate — validate descriptor structure (full checks with pip install undatum[frictionless])

Metadata options:

  • --name, --title, --description, --keywords
  • --licenses (semicolon-separated entries, e.g. name=MIT;name=ODC-PDDL-1.0)
  • --sources (semicolon-separated entries, e.g. title=World Bank,path=https://...)
  • --contributors (semicolon-separated entries, e.g. title=Jane Doe,email=jane@example.com)
  • --version - Package version string

Features:

  • Frictionless profile: Emits profile: tabular-data-package with resource format/mediatype
  • Schema inference: Automatically infers field types, descriptions, and uniqueness constraints
  • Multiple resources: Package multiple files as separate resources
  • Remote URIs: Support for HTTP/HTTPS URLs as resource paths
  • Package directory: Bundle datapackage.json with data file copies
  • AI metadata: Use --autodoc to generate metadata with AI assistance (single-pass, no duplicate LLM calls)
  • Streaming-safe: Processes large datasets without loading everything into memory
  • Python SDK: Dataset.read("data.csv").package(output="datapackage.json")

Additional options:

  • --package-dir: Create a package directory with data file copies
  • --zip: Create a ZIP archive of the package directory (requires --package-dir)
  • --autodoc: Enable AI-powered metadata generation (reuses doc command logic)
  • --engine: Processing engine (auto or duckdb)
  • --delimiter, --encoding, --tagname, --start-line, --start-page: Passed through to analysis and sampling
  • --limit: Maximum objects to analyze for schema inference (default: 10000)
  • --sample-size: Number of sample records for metadata inference (default: 10)

Reference​

undatum package create​

undatum package create [OPTIONS] INPUT_FILES...
ArgumentDescription
INPUT_FILES...Input file(s) to package. (required)
OptionDescriptionDefault
-o, --output TEXTOutput datapackage.json path.
--package-dir TEXTPackage directory to materialize.
--name TEXTPackage name (slug).
--title TEXTPackage title.
--description TEXTPackage description.
--keywords TEXTComma-separated keywords.
--licenses TEXTLicenses (semicolon-separated entries, e.g. 'name=MIT;name=ODC-PDDL-1.0').
--sources TEXTSources (semicolon-separated entries, e.g. 'title=World Bank,path=https://...').
--contributors TEXTContributors (semicolon-separated entries, e.g. 'title=Jane Doe,email=jane@example.com').
--version TEXTPackage version string.
--sample-size INTEGERNumber of sample records to include in metadata inference.10
-n, --limit INTEGERMaximum number of objects to analyze for schema inference.10000
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--tagname TEXTXML tag name that contains individual records.
--start-line INTEGERLine number (0-based) to start reading from.0
--start-page INTEGERPage number (0-based) to start from for Excel files.0
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl', 'xml').
--autodoc / --no-autodocEnable AI-powered metadata generation.--no-autodoc
--lang TEXTLanguage for AI-generated metadata (default: 'English').English
--ai-provider TEXTAI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml).
--ai-model TEXTModel name to use (provider-specific).
--ai-base-url TEXTBase URL for AI API (optional).
--pii-mask-samples / --no-pii-mask-samplesMask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio.
--zip TEXTCreate a ZIP archive of the package directory (requires --package-dir).
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--verbose / --no-verboseEnable verbose logging output.--no-verbose

Deprecated spellings (removed in 2.0): --objects-limit → --limit.

undatum package add-resource​

undatum package add-resource [OPTIONS] PACKAGE_FILE INPUT_FILES...
ArgumentDescription
PACKAGE_FILEExisting datapackage.json to extend. (required)
INPUT_FILES...Input file(s) to add. (required)
OptionDescriptionDefault
--package-dir TEXTPackage directory containing data files (defaults to descriptor dir).
--sample-size INTEGERSample size for metadata inference.10
-n, --limit INTEGERMaximum objects to analyze.10000
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding.
--tagname TEXTXML record tag name.
--start-line INTEGERLine number (0-based) to start reading from.0
--start-page INTEGERExcel start page (0-based).0
-F, --format-in TEXTOverride input format.
--autodoc / --no-autodocEnable AI-powered metadata generation.--no-autodoc
--lang TEXTLanguage for AI metadata.English
--ai-provider TEXTAI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml).
--ai-model TEXTAI model name.
--ai-base-url TEXTAI API base URL.
--pii-mask-samples / --no-pii-mask-samplesMask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio.
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--verbose / --no-verboseEnable verbose logging output.--no-verbose

Deprecated spellings (removed in 2.0): --objects-limit → --limit.

undatum package validate​

undatum package validate [OPTIONS] PACKAGE_FILE
ArgumentDescription
PACKAGE_FILEPath to datapackage.json. (required)
OptionDescriptionDefault
--limit-rows INTEGERLimit rows validated per resource.
--check-data / --no-check-dataValidate resource data in addition to metadata.--check-data
--verbose / --no-verboseEnable verbose logging output.--no-verbose

See also shared options.