Skip to main content

apply

Applies a Python process(record) function, or a registered transform plugin, to each record. Either --script or --plugin is required.

Write with --output. A trailing path is not a positional argument.

# Script: the file must define process(record) -> dict
undatum apply --script transform.py data.jsonl --output output.jsonl

# Registered transform plugin
undatum apply data.jsonl --plugin example-transform --output out.jsonl

# Subset first, then transform
undatum apply --script transform.py --filter '`status` == "active"' data.jsonl --output out.jsonl

Install transform plugins via the undatum.plugins entry point; see Plugins.

Script contract​

The script is loaded with runpy.run_path. It must define process:

def process(record: dict) -> dict:
record["name"] = str(record.get("name", "")).strip()
return record

Missing process is a validation error. The function runs twice internally: a schema pass over at most 1000 records, then the write pass.

Options​

  • --script PATH — Python file with process
  • --plugin NAME — registered transform plugin instead of a script
  • --filter / --filter-expr — comparison expression before the transform (shared options)
  • --output — output path (stdout if omitted)
  • --format-in, --zipfile, --delimiter, --encoding, --start-page, --trust
  • --flatten-nested, --max-nested-depth, --keep-nested-parents / --no-keep-nested-parents, --on-error, --error-log, --table, --quotechar — shared options

Reference​

Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: streaming, two passes over the input · Engines: python

undatum apply [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
-f, --fields TEXTComma-separated list of field names (kept for compatibility).
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').utf8
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl').
--zipfile / --no-zipfileTreat input file as a ZIP archive.--no-zipfile
--script TEXTPath to Python script file containing transformation function.
--plugin TEXTName of a registered transform plugin to apply instead of a script.
--filter, --filter-expr TEXTFilter expression to apply before transformation.
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents

See also shared options.