apply
Applies a Python process(record) function, or a registered transform plugin, to each record. Either --script or --plugin is required.
Write with --output. A trailing path is not a positional argument.
# Script: the file must define process(record) -> dict
undatum apply --script transform.py data.jsonl --output output.jsonl
# Registered transform plugin
undatum apply data.jsonl --plugin example-transform --output out.jsonl
# Subset first, then transform
undatum apply --script transform.py --filter '`status` == "active"' data.jsonl --output out.jsonl
Install transform plugins via the undatum.plugins entry point; see Plugins.
Script contract
The script is loaded with runpy.run_path. It must define process:
def process(record: dict) -> dict:
record["name"] = str(record.get("name", "")).strip()
return record
Missing process is a validation error. The function runs twice internally: a schema pass over at most 1000 records, then the write pass.
Options
--script PATH— Python file withprocess--plugin NAME— registered transform plugin instead of a script--filter/--filter-expr— comparison expression before the transform (shared options)--output— output path (stdout if omitted)--format-in,--zipfile,--delimiter,--encoding,--start-page,--trust--flatten-nested,--max-nested-depth,--keep-nested-parents/--no-keep-nested-parents,--on-error,--error-log,--table,--quotechar— shared options
Reference
Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: streaming, two passes over the input · Engines: python
undatum apply [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file. (required) |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
-f, --fields TEXT | Comma-separated list of field names (kept for compatibility). | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | utf8 |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl'). | |
--zipfile / --no-zipfile | Treat input file as a ZIP archive. | --no-zipfile |
--script TEXT | Path to Python script file containing transformation function. | |
--plugin TEXT | Name of a registered transform plugin to apply instead of a script. | |
--filter, --filter-expr TEXT | Filter expression to apply before transformation. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
See also shared options.