split
Splits datasets into multiple files based on chunk size or field values.
# Split by chunk size
undatum split --chunksize 10000 data.jsonl
# Split by field value
undatum split --fields category data.jsonl
undatum split workbook.xlsx --table Sheet2 --fields city --dirname out/
undatum split nested.jsonl --fields capital_city.lat --flatten-nested --dirname out/
undatum split events.jsonl --filter '`city` == "Berlin"' --dirname out/
Parts are written in the input's format (CSV in, CSV out), or in --format-out /
-O. Chunks are named <name>_1.<ext>, <name>_2.<ext>, ...; field splits name each file
after its value (north.csv; empty and missing values are __null__.csv).
Hive partitions
--hive writes one directory per value instead, in the layout DuckDB, Spark, Athena and
BigQuery read as partitioned tables:
undatum split data.csv --fields country --hive --dirname parts/
undatum split data.csv --fields country,city --hive --dirname parts/ -O parquet
parts/country=DE/city=Berlin/data_0.parquet
parts/country=FR/city=Paris/data_0.parquet
Values are percent-encoded (a/b → a%2Fb), missing and empty values go to
__HIVE_DEFAULT_PARTITION__, and the partition fields are left out of the files (readers
restore them from the path). At most --max-open-files files (default 128) are open at once;
with more distinct values, a partition that was closed continues in data_1.<ext>.
--filter uses the same comparison syntax as select. --dirname is the output directory;
--output can set a filename prefix; --gzipfile gzips the parts. See
shared CLI options.
Reference
Reads: any readable format · Writes: one file per chunk or per value · Memory: streaming · Engines: python
undatum split [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file. (required) |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Optional output file path prefix. If not specified, uses input filename. | |
-f, --fields TEXT | Comma-separated field names to split by (creates one file per unique value combination). | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | utf8 |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl'). | |
--zipfile / --no-zipfile | Treat input file as a ZIP archive. | --no-zipfile |
--gzipfile TEXT | Gzip compression option for output files. | |
--chunksize INTEGER | Number of records per chunk when splitting by size (default: 10000). | 10000 |
--filter, --filter-expr TEXT | Filter expression to apply before splitting. | |
--dirname TEXT | Directory path to write output files to. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--hive | With --fields, write field=value directories (Hive partitioning). | |
--max-open-files INTEGER | Partition files open at once; more distinct keys write additional part files. | 128 |
-O, --format-out TEXT | Format of the parts (default: the input's format). |
See also shared options.