Skip to main content

frequency

Calculates frequency distribution for specified fields. --fields is required.

Default stdout format is CSV. Use --format-out json (or an --output path ending in .json) for JSON.

undatum frequency --fields category data.jsonl
undatum frequency --fields status,region data.csv
undatum frequency --fields city --format-out json --output freq.json data.csv
undatum frequency --fields city workbook.xlsx --table Sheet2
undatum frequency --fields capital_city.lat nested.jsonl --flatten-nested

Reference​

Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: grows with the number of distinct values · Engines: auto, duckdb, python

undatum frequency [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
-f, --fields TEXTComma-separated list of field names to calculate frequency for.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride file type detection (e.g., 'csv', 'jsonl', 'xlsx').
--start-page INTEGERSheet index (0-based) for Excel files.0
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--threads INTEGERWorker processes for Python/iterable-engine frequency counting. Omit for sequential. For DuckDB, use --duckdb-threads.
--duckdb-threads INTEGERNumber of threads for DuckDB engine.
--duckdb-memory TEXTMemory limit for DuckDB (e.g., '4GB', '512MB').
--duckdb-temp-dir TEXTTemporary directory for DuckDB.
--filter, --filter-expr TEXTFilter expression to apply before counting frequencies.
-O, --format-out TEXTOutput format: 'csv' (default) or 'json' (also inferred from --output).
--jsonPrint the result as one JSON document (same as --format-out json).

Deprecated spellings (removed in 2.0): --filetype → --format-in.

See also shared options.