Skip to main content

convert

Converts data between any formats supported by iterabledata (140+, see undatum formats list). Reading and writing are handled by the iterabledata engine, including cloud URIs (s3://, gs://, az://). Use --recursive to bulk-convert a directory or glob pattern.

# XML to JSON Lines
undatum convert --tagname item data.xml data.jsonl

# CSV to Parquet
undatum convert data.csv data.parquet

# JSON Lines to CSV
undatum convert data.jsonl data.csv

# Convert from S3 to local
undatum convert s3://my-bucket/data.csv output.jsonl

# Bulk-convert a directory of CSVs to Parquet
undatum convert ./raw ./processed --recursive --to-ext parquet
undatum convert ./raw ./out --recursive --to-ext jsonl --filename-pattern "{stem}.converted.jsonl"

# Convert local to S3
undatum convert input.csv s3://my-bucket/output.parquet

# Convert S3 to S3
undatum convert s3://bucket/input.jsonl s3://bucket/output.parquet

Cloud storage: Input and output paths support s3://, gs:///gcs://, and az:///abfs:// URIs when the cloud extra is installed. See Cloud Storage Support.

  • --format-in / --format-out — override format detection
  • --table / --sheet — named table or Excel sheet (keep --start-page for a 0-based index)
  • --native-batch / --columns / --row-range — native columnar batch convert (auto with --low-memory when both formats support it); --batch-size also sizes native scanner chunks
  • --profile fast|balanced|max — codec performance profile for compressed output
  • --level N — explicit compression level for compressed output (overrides --profile; skips DuckDB COPY)
  • --write-mode append|overwrite|error|ignore|create — lakehouse write mode (Delta / Iceberg / DuckLake / Lance)
  • --row-group-size N — Parquet write row-group size (skips DuckDB COPY; pair with --batch-size if you need groups smaller than convert's write batches)
  • --use-totals — use format-reported row totals for progress when available
  • --trust — acknowledge pickle deserialization risk
  • --on-error raise|skip|warn — parse-error policy for malformed rows (default: raise)
  • --error-log PATH — append skipped/warned parse errors as JSONL
  • --delimiter, --quotechar, --encoding, --tagname — passed through to the reader (delimiter auto-detected for CSV when omitted)
  • --recursive / --to-ext / --filename-pattern — bulk-convert directories or globs ({name}, {stem}, {ext} in the output name)
  • --flatten-data — flatten nested records to a flat schema (not --flatten; convert does not take --flatten-nested)
  • --low-memory — streaming / native-batch path for large files
  • --engine auto|duckdb|python
  • --compression — output codec (for example snappy, gzip, brotli)
  • --prefix-strip / --no-prefix-strip — XML namespace prefixes
  • --start-line, --scan-limit — skip leading lines; cap format detection scan
  • --strict-native / --no-native-batch — require or disable native bulk I/O
  • --atomic — write to a temp file and rename on success (local paths only)
  • --threads, --batch-size, --progress / --no-progress — throughput and feedback (--threads is process-pool chunk parallelism for single-file Python-engine convert, and concurrent workers for --recursive)

Partitioned output​

--partition-by writes OUTPUT as a directory with one subdirectory per value (Hive layout), in any writable format:

undatum convert data.csv by_country --partition-by country -O parquet
undatum convert data.csv by_place --partition-by country,city -O csv --where "amount > 50"
duckdb -c "SELECT country, count(*) FROM read_parquet('by_country/**/*.parquet') GROUP BY 1"

CSV and Parquet output from DuckDB-readable inputs is written by DuckDB 1.5 or later (COPY ... PARTITION_BY); other formats, inputs and DuckDB versions go through a streaming writer with at most --max-open-files (default 128) open files. Both produce the same layout: country=DE/data_0.parquet, percent-encoded values, __HIVE_DEFAULT_PARTITION__ for missing and empty values, partition fields kept out of the files. The output directory must be new or empty. split --hive writes the same layout.

convert takes two positional paths (INPUT OUTPUT). Reader and error-policy flags (--table, --on-error, --error-log, --quotechar, --trust): Shared CLI options.

Reference​

Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: streaming · Engines: auto, duckdb, python

undatum convert [OPTIONS] INPUT_FILE OUTPUT
ArgumentDescription
INPUT_FILEPath to input file to convert. (required)
OUTPUTPath to output file. (required)
OptionDescriptionDefault
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--compression TEXTOutput compression codec (e.g. 'brotli', 'snappy', 'gzip').
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').utf8
--verbose / --no-verboseEnable verbose logging output.--no-verbose
--flatten-data / --no-flatten-dataFlatten nested data structures into flat records.--no-flatten-data
--prefix-strip / --no-prefix-stripStrip XML namespace prefixes from element names.--prefix-strip
--start-line INTEGERLine number (0-based) to start reading from.0
--start-page INTEGERPage number (0-based) to start from for Excel files.0
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--tagname TEXTXML tag name that contains individual records.
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl', 'xml').
-O, --format-out TEXTOverride output file format (e.g., 'csv', 'jsonl', 'parquet').
--batch-size INTEGERNumber of records per conversion batch.50000
--scan-limit INTEGERRecords to sample for output schema detection.1000
--atomic / --no-atomicWrite to a temp file and rename on success (local output only).--no-atomic
--threads INTEGERWorker processes for Python-engine chunk parallelism (single-file convert and bulk --recursive). Omit for sequential processing. DuckDB uses its own threading (see --engine / duckdb thread settings); not nested around DuckDB COPY.
--progress / --no-progressShow progress bar.--progress
--low-memory / --no-low-memoryPrefer spill-to-disk / smaller batches for large-file conversion (DuckDB COPY when possible).--no-low-memory
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--recursive / --no-recursiveBulk-convert a directory or glob pattern; OUTPUT is treated as a directory.--no-recursive
--to-ext TEXTTarget extension for bulk conversion (e.g. 'parquet'). Defaults to --format-out.
--filename-pattern TEXTBulk output name pattern with {name}, {stem}, {ext} (used with --recursive; default replaces the extension via --to-ext).
--profile TEXTCodec performance profile for compressed output: fast, balanced, or max.
--level INTEGERExplicit compression level for compressed output (overrides --profile). Same codecargs compression_level as iterabledata / undatum repack.
--native-batch / --no-native-batchUse native columnar batch conversion when both formats support it. Default: auto-enable with --low-memory.
--strict-nativeFail if native batch conversion is requested but unsupported.
--columns TEXTComma-separated columns to convert (native batch projection).
--row-range TEXTRow range to convert as START:END (native batch).
--write-mode TEXTLakehouse write mode: append, overwrite, error, ignore, or create (Delta / Iceberg / DuckLake / Lance).
--row-group-size INTEGERParquet write row-group size (iterabledata row_group_size; defaults to the writer batch size when omitted). Skips DuckDB COPY.
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--use-totalsUse format-reported row totals for convert progress when available.
--where TEXTKeep records where this SQL condition is true (DuckDB syntax), e.g. "amount > 100 AND city = 'Berlin'". Text values are typed automatically.
--add TEXTAdd a computed column: 'name = SQL expression', e.g. 'total = price * qty' (repeatable; an existing name is replaced in place).
--partition-by TEXTWrite OUTPUT as a directory partitioned by these fields (Hive layout), e.g. year,month.
--max-open-files INTEGERPartition files open at once; more distinct keys write additional part files.128

See also shared options.