Skip to main content

replace

Performs string replacement in specified fields. Supports simple string replacement and regex-based replacement.

Write with --output. A trailing path is not a positional argument.

# Simple string replacement
undatum replace data.csv --field name --pattern "Mr\." --replacement "Mr" --output output.csv

# Regex replacement
undatum replace data.jsonl --field email --pattern "@old.com" --replacement "@new.com" --regex --output output.jsonl

# Global replacement (all occurrences; default is first match only)
undatum replace data.csv --field text --pattern "old" --replacement "new" --global-replace --output output.csv
  • --field, --pattern, --replacement
  • --regex — treat --pattern as regex
  • --global-replace — replace every match in the field (not --global)
  • --output — output path (stdout if omitted)

Also accepts --table, --flatten-nested, --on-error, --error-log, and --quotechar (shared options).

Reference​

Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: streaming · Engines: auto, duckdb, python

undatum replace [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--field TEXTField name to perform replacement in.
--pattern TEXTPattern to search for (string or regex).
--replacement TEXTReplacement string.
--regex / --no-regexTreat pattern as regex.--no-regex
--global-replace / --no-global-replaceReplace all occurrences (default: replace first only).--no-global-replace
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl').
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
-O, --format-out TEXTOutput format (e.g. csv, jsonl, parquet). Defaults to the --output extension, or to the input's text format on stdout.

See also shared options.