Skip to main content

split

Splits datasets into multiple files based on chunk size or field values.

# Split by chunk size
undatum split --chunksize 10000 data.jsonl

# Split by field value
undatum split --fields category data.jsonl
undatum split workbook.xlsx --table Sheet2 --fields city --dirname out/
undatum split nested.jsonl --fields capital_city.lat --flatten-nested --dirname out/
undatum split events.jsonl --filter '`city` == "Berlin"' --dirname out/

Parts are written in the input's format (CSV in, CSV out), or in --format-out / -O. Chunks are named <name>_1.<ext>, <name>_2.<ext>, ...; field splits name each file after its value (north.csv; empty and missing values are __null__.csv).

Hive partitions​

--hive writes one directory per value instead, in the layout DuckDB, Spark, Athena and BigQuery read as partitioned tables:

undatum split data.csv --fields country --hive --dirname parts/
undatum split data.csv --fields country,city --hive --dirname parts/ -O parquet
parts/country=DE/city=Berlin/data_0.parquet
parts/country=FR/city=Paris/data_0.parquet

Values are percent-encoded (a/b → a%2Fb), missing and empty values go to __HIVE_DEFAULT_PARTITION__, and the partition fields are left out of the files (readers restore them from the path). At most --max-open-files files (default 128) are open at once; with more distinct values, a partition that was closed continues in data_1.<ext>.

--filter uses the same comparison syntax as select. --dirname is the output directory; --output can set a filename prefix; --gzipfile gzips the parts. See shared CLI options.

Reference​

Reads: any readable format · Writes: one file per chunk or per value · Memory: streaming · Engines: python

undatum split [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path prefix. If not specified, uses input filename.
-f, --fields TEXTComma-separated field names to split by (creates one file per unique value combination).
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').utf8
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl').
--zipfile / --no-zipfileTreat input file as a ZIP archive.--no-zipfile
--gzipfile TEXTGzip compression option for output files.
--chunksize INTEGERNumber of records per chunk when splitting by size (default: 10000).10000
--filter, --filter-expr TEXTFilter expression to apply before splitting.
--dirname TEXTDirectory path to write output files to.
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--hiveWith --fields, write field=value directories (Hive partitioning).
--max-open-files INTEGERPartition files open at once; more distinct keys write additional part files.128
-O, --format-out TEXTFormat of the parts (default: the input's format).

See also shared options.