Skip to main content

sort

Sorts rows by one or more columns. Supports multiple sort keys, ascending/descending order, and numeric sorting.

Write with --output. A trailing path is not a positional argument (convert is the command that takes INPUT OUTPUT).

# Sort by single column ascending
undatum sort data.csv --by name --output output.csv

# Sort by multiple columns
undatum sort data.jsonl --by name,age --output output.jsonl

# Sort descending
undatum sort data.csv --by date --desc --output output.csv

# Numeric sort: --numeric takes a field list, not an output path
undatum sort data.csv --by price --numeric price --output output.csv
undatum sort workbook.xlsx --table Sheet2 --by city --output out.jsonl
undatum sort nested.jsonl --by capital_city.lat --flatten-nested --numeric capital_city.lat --output out.jsonl
  • --by — comma-separated sort fields
  • --desc — descending order
  • --numeric FIELD[,FIELD…] — treat those fields as numbers
  • --output — output path (stdout if omitted)
  • --low-memory — external merge sort (spill runs to disk)
  • --engine auto|duckdb|python
  • --format-in — override type detection (this command does not take --format-in)

Also accepts --table, --flatten-nested, --on-error, --error-log, and --quotechar (shared options).

Reference​

Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: in memory up to 100k rows, external merge sort on disk above (or with --low-memory); DuckDB spills to disk · Engines: auto, duckdb, python

undatum sort [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--by TEXTComma-separated list of field names to sort by.
--desc / --no-descSort in descending order.--no-desc
--numeric TEXTComma-separated list of field names to sort numerically.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride file type detection (e.g., 'csv', 'jsonl').
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--duckdb-threads INTEGERNumber of threads for DuckDB engine.
--duckdb-memory TEXTMemory limit for DuckDB (e.g., '4GB', '512MB').
--duckdb-temp-dir TEXTTemporary directory for DuckDB.
--low-memory / --no-low-memoryForce external merge sort (spill sorted runs to disk).--no-low-memory
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
-O, --format-out TEXTOutput format (e.g. csv, jsonl, parquet). Defaults to the --output extension, or to the input's text format on stdout.

Deprecated spellings (removed in 2.0): --filetype → --format-in.

See also shared options.