Basic usage
The CLI entry points are undatum and the shorter data alias.
Pipelines: standard input and output
Use - as the input path to read standard input, and leave out --output to write to
standard output. Commands chain like other Unix tools:
cat data.csv | undatum sort - --by amount --numeric amount | undatum head - -n 10
undatum search events.jsonl --pattern ERROR --fields level | undatum dedup -
curl -s https://example.org/export.csv.gz | undatum count -
- Input format is detected from the first bytes (CSV, TSV, JSON, JSON Lines, Parquet)
together with
.gz,.zst,.bz2and.xzcompression. Pass--format-in(-F) when detection guesses wrong, for example-F tsv. - Output format on stdout follows the input's text format — CSV in, CSV out; JSON Lines
in, JSON Lines out. Binary or unknown inputs print JSON Lines. Override it with
--format-out(-O):undatum head data.csv -n 5 -O jsonl. - Binary formats (
-O parquet, ...) are written to stdout only when it is redirected or piped; on a terminal the command fails and suggests--output. - Logs and progress go to stderr, so they never mix with the data. Closing the pipe early
(
undatum head big.csv -n 1000000 | head -1) ends the command quietly with exit code 0.
Working with Compressed Files
undatum can process files inside compressed containers (ZIP, GZ, BZ2, XZ, ZSTD) with minimal memory usage.
# Process file inside ZIP archive
undatum headers --format-in jsonl data.zip
# Process XZ compressed file
undatum uniq --fields country --format-in jsonl data.jsonl.xz
Filtering Data
Filter rows with comparison expressions on commands that support --filter (select, frequency, uniq, plot, validate, split, and others). The same expression is pushed to DuckDB WHERE when possible, or evaluated in-process on the iterable path. For LIKE, IN, joins, and aggregations, use undatum sql.
# Filter by field value
undatum select --fields name,email --filter '`status` == "active"' data.jsonl
# Complex filters (AND/OR and &&/|| are both accepted)
undatum frequency --fields category --filter '`price` > 100 && `status` == "active"' data.jsonl
# DuckDB-accelerated select with SQL pushdown
undatum select --fields name,email --filter '`status` == "active"' --engine duckdb data.jsonl
# Unique values after a filter
undatum uniq --fields city --filter 'age >= 30' --engine duckdb data.jsonl
# Natural-language filter (translates to an expression; use --apply to run)
undatum ai filter "customers in California with orders over 1000" data.csv --apply
Filter syntax:
- Field names:
`fieldname`(backticks optional for simple names) - Strings:
"value"or'value' - Comparisons:
==,!=,>,<,>=,<= - Booleans:
AND/ORor&&/||(both work)
DuckDB pushdown: comparisons, AND/OR/&&/||, parentheses, and simple identifiers are translated to WHERE. The following are not supported on --filter (use undatum sql instead):
INlists and SQLLIKE- Nested dotted fields (
user.name) - Regex /
match
Migrating from older filters: AND/OR and &&/|| both work. Prefer double-quoted strings. There is no IN operator; write status == "active" || status == "pending". The experimental MistQL undatum query command is gone; use sql or select --filter.
For ad-hoc SQL over files, use sql or db query against a database URI.
Custom Encoding and Delimiters
CSV/TSV delimiters (comma, semicolon, tab, pipe) are auto-detected when --delimiter is omitted. --delimiter describes the input; written CSV uses commas and TSV uses tabs. convert, flatten, apply, and split default --encoding to utf8; pass --encoding when the file is not UTF-8.
Override when needed:
undatum headers --encoding cp1251 --delimiter ";" data.csv
undatum convert --encoding utf-8 --delimiter "," data.csv data.jsonl
Date Detection
Automatic date/datetime field detection:
undatum stats data.jsonl
# Date detection is on by default; disable with --no-checkdates
undatum stats --no-checkdates data.jsonl
This uses the qddate library to automatically identify and parse date fields.
See also: CLI reference, shared CLI options, performance.