select
Selects and reorders columns from files. Supports filtering, nested dot-notation fields, and engine selection. When the DuckDB engine is used, filter expressions are pushed to SQL when possible and results can be written directly via COPY for CSV, JSON, and Parquet output.
--filter is the documented flag (--filter-expr is an alias). Syntax: Basic usage.
undatum select --fields name,email,status data.jsonl
undatum select --fields name,email --filter '`status` == "active"' data.jsonl
undatum select --fields user.name,user.email --engine duckdb data.jsonl
undatum select --fields name,email --engine duckdb --output subset.csv data.jsonl
undatum select --fields name --table Sheet2 workbook.xlsx
undatum select --fields name,capital_city.lat --flatten-nested nested.jsonl
# SQL condition and a computed column
undatum select data.csv --where "amount > 100 AND city = 'Berlin'" --fields name,amount
undatum select data.csv --add "total = price * quantity" --add "big = amount > 200"
--where takes a SQL condition (DuckDB syntax) on typed values — text columns get the number, date or boolean type most of their values have — while the output keeps the original text; unknown columns are reported with suggestions. See SQL expressions. --add name = expression is repeatable; an existing name is replaced
in place.
Reference
Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: streaming · Engines: auto, duckdb, python
undatum select [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file. (required) |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
-f, --fields TEXT | Comma-separated list of field names to select and reorder. | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl', 'xlsx'). | |
-O, --format-out TEXT | Override output format (e.g., 'csv', 'jsonl'). | |
--zipfile / --no-zipfile | Treat input file as a ZIP archive. | --no-zipfile |
--filter, --filter-expr TEXT | Filter expression to apply (e.g., "status == 'active'"). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--duckdb-threads INTEGER | Number of threads for DuckDB engine. | |
--duckdb-memory TEXT | Memory limit for DuckDB (e.g., '4GB', '512MB'). | |
--duckdb-temp-dir TEXT | Temporary directory for DuckDB. | |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--where TEXT | Keep records where this SQL condition is true (DuckDB syntax), e.g. "amount > 100 AND city = 'Berlin'". Text values are typed automatically. | |
--add TEXT | Add a computed column: 'name = SQL expression', e.g. 'total = price * qty' (repeatable; an existing name is replaced in place). |
See also shared options.