convert
Converts data between any formats supported by iterabledata (140+, see undatum formats list). Reading and writing are handled by the iterabledata engine, including cloud URIs (s3://, gs://, az://). Use --recursive to bulk-convert a directory or glob pattern.
# XML to JSON Lines
undatum convert --tagname item data.xml data.jsonl
# CSV to Parquet
undatum convert data.csv data.parquet
# JSON Lines to CSV
undatum convert data.jsonl data.csv
# Convert from S3 to local
undatum convert s3://my-bucket/data.csv output.jsonl
# Bulk-convert a directory of CSVs to Parquet
undatum convert ./raw ./processed --recursive --to-ext parquet
undatum convert ./raw ./out --recursive --to-ext jsonl --filename-pattern "{stem}.converted.jsonl"
# Convert local to S3
undatum convert input.csv s3://my-bucket/output.parquet
# Convert S3 to S3
undatum convert s3://bucket/input.jsonl s3://bucket/output.parquet
Cloud storage: Input and output paths support s3://, gs:///gcs://, and az:///abfs:// URIs when the cloud extra is installed. See Cloud Storage Support.
--format-in/--format-out— override format detection--table/--sheet— named table or Excel sheet (keep--start-pagefor a 0-based index)--native-batch/--columns/--row-range— native columnar batch convert (auto with--low-memorywhen both formats support it);--batch-sizealso sizes native scanner chunks--profile fast|balanced|max— codec performance profile for compressed output--level N— explicit compression level for compressed output (overrides--profile; skips DuckDB COPY)--write-mode append|overwrite|error|ignore|create— lakehouse write mode (Delta / Iceberg / DuckLake / Lance)--row-group-size N— Parquet write row-group size (skips DuckDB COPY; pair with--batch-sizeif you need groups smaller than convert's write batches)--use-totals— use format-reported row totals for progress when available--trust— acknowledge pickle deserialization risk--on-error raise|skip|warn— parse-error policy for malformed rows (default: raise)--error-log PATH— append skipped/warned parse errors as JSONL--delimiter,--quotechar,--encoding,--tagname— passed through to the reader (delimiter auto-detected for CSV when omitted)--recursive/--to-ext/--filename-pattern— bulk-convert directories or globs ({name},{stem},{ext}in the output name)--flatten-data— flatten nested records to a flat schema (not--flatten; convert does not take--flatten-nested)--low-memory— streaming / native-batch path for large files--engine auto|duckdb|python--compression— output codec (for examplesnappy,gzip,brotli)--prefix-strip/--no-prefix-strip— XML namespace prefixes--start-line,--scan-limit— skip leading lines; cap format detection scan--strict-native/--no-native-batch— require or disable native bulk I/O--atomic— write to a temp file and rename on success (local paths only)--threads,--batch-size,--progress/--no-progress— throughput and feedback (--threadsis process-pool chunk parallelism for single-file Python-engine convert, and concurrent workers for--recursive)
Partitioned output
--partition-by writes OUTPUT as a directory with one subdirectory per value (Hive layout),
in any writable format:
undatum convert data.csv by_country --partition-by country -O parquet
undatum convert data.csv by_place --partition-by country,city -O csv --where "amount > 50"
duckdb -c "SELECT country, count(*) FROM read_parquet('by_country/**/*.parquet') GROUP BY 1"
CSV and Parquet output from DuckDB-readable inputs is written by DuckDB 1.5 or later
(COPY ... PARTITION_BY); other formats, inputs and DuckDB versions go through a streaming
writer with at most --max-open-files (default 128) open files. Both produce the same layout:
country=DE/data_0.parquet, percent-encoded values, __HIVE_DEFAULT_PARTITION__ for
missing and empty values, partition fields kept out of the files. The output directory must
be new or empty. split --hive writes the same layout.
convert takes two positional paths (INPUT OUTPUT). Reader and error-policy flags (--table, --on-error, --error-log, --quotechar, --trust): Shared CLI options.
Reference
Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: streaming · Engines: auto, duckdb, python
undatum convert [OPTIONS] INPUT_FILE OUTPUT
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file to convert. (required) |
OUTPUT | Path to output file. (required) |
| Option | Description | Default |
|---|---|---|
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--compression TEXT | Output compression codec (e.g. 'brotli', 'snappy', 'gzip'). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | utf8 |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
--flatten-data / --no-flatten-data | Flatten nested data structures into flat records. | --no-flatten-data |
--prefix-strip / --no-prefix-strip | Strip XML namespace prefixes from element names. | --prefix-strip |
--start-line INTEGER | Line number (0-based) to start reading from. | 0 |
--start-page INTEGER | Page number (0-based) to start from for Excel files. | 0 |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--tagname TEXT | XML tag name that contains individual records. | |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl', 'xml'). | |
-O, --format-out TEXT | Override output file format (e.g., 'csv', 'jsonl', 'parquet'). | |
--batch-size INTEGER | Number of records per conversion batch. | 50000 |
--scan-limit INTEGER | Records to sample for output schema detection. | 1000 |
--atomic / --no-atomic | Write to a temp file and rename on success (local output only). | --no-atomic |
--threads INTEGER | Worker processes for Python-engine chunk parallelism (single-file convert and bulk --recursive). Omit for sequential processing. DuckDB uses its own threading (see --engine / duckdb thread settings); not nested around DuckDB COPY. | |
--progress / --no-progress | Show progress bar. | --progress |
--low-memory / --no-low-memory | Prefer spill-to-disk / smaller batches for large-file conversion (DuckDB COPY when possible). | --no-low-memory |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--recursive / --no-recursive | Bulk-convert a directory or glob pattern; OUTPUT is treated as a directory. | --no-recursive |
--to-ext TEXT | Target extension for bulk conversion (e.g. 'parquet'). Defaults to --format-out. | |
--filename-pattern TEXT | Bulk output name pattern with {name}, {stem}, {ext} (used with --recursive; default replaces the extension via --to-ext). | |
--profile TEXT | Codec performance profile for compressed output: fast, balanced, or max. | |
--level INTEGER | Explicit compression level for compressed output (overrides --profile). Same codecargs compression_level as iterabledata / undatum repack. | |
--native-batch / --no-native-batch | Use native columnar batch conversion when both formats support it. Default: auto-enable with --low-memory. | |
--strict-native | Fail if native batch conversion is requested but unsupported. | |
--columns TEXT | Comma-separated columns to convert (native batch projection). | |
--row-range TEXT | Row range to convert as START:END (native batch). | |
--write-mode TEXT | Lakehouse write mode: append, overwrite, error, ignore, or create (Delta / Iceberg / DuckLake / Lance). | |
--row-group-size INTEGER | Parquet write row-group size (iterabledata row_group_size; defaults to the writer batch size when omitted). Skips DuckDB COPY. | |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--use-totals | Use format-reported row totals for convert progress when available. | |
--where TEXT | Keep records where this SQL condition is true (DuckDB syntax), e.g. "amount > 100 AND city = 'Berlin'". Text values are typed automatically. | |
--add TEXT | Add a computed column: 'name = SQL expression', e.g. 'total = price * qty' (repeatable; an existing name is replaced in place). | |
--partition-by TEXT | Write OUTPUT as a directory partitioned by these fields (Hive layout), e.g. year,month. | |
--max-open-files INTEGER | Partition files open at once; more distinct keys write additional part files. | 128 |
See also shared options.