Skip to main content

join

Performs relational joins between two files. Supports inner, left, right, and full outer joins.

Write with --output. A trailing path is not a positional argument.

# Inner join by key field
undatum join data1.csv data2.csv --on email --type inner --output output.csv

# Left join (keep all rows from first file)
undatum join data1.jsonl data2.jsonl --on id --type left --output output.jsonl

# Right join (keep all rows from second file)
undatum join data1.csv data2.csv --on id --type right --output output.csv

# Full outer join (keep all rows from both files)
undatum join data1.jsonl data2.jsonl --on id --type full --output output.jsonl
undatum join workbook.xlsx other.xlsx --table Sheet2 --table2 Cities --on city --output out.jsonl
undatum join left.jsonl right.jsonl --on capital_city.lat --flatten-nested --output out.jsonl
  • --on — join key field(s)
  • --type inner|left|right|full
  • --output — output path (stdout if omitted)
  • --table / --sheet and --table2 / --sheet2 — named tables for each file
  • --filetype1 / --filetype2 — override type detection per file
  • --engine, --progress / --no-progress

Also accepts --flatten-nested, --on-error, --error-log, and --quotechar (shared options).

Reference​

Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: Python engine indexes the second file in memory; DuckDB works out of core · Engines: auto, duckdb, python

undatum join [OPTIONS] FILE1 FILE2
ArgumentDescription
FILE1Path to first input file. (required)
FILE2Path to second input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--on TEXTComma-separated list of key field names to join on.
--type TEXTJoin type: 'inner' (default), 'left', 'right', or 'full'.inner
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--verbose / --no-verboseEnable verbose logging output.--no-verbose
--filetype1 TEXTOverride file type detection for first file.
--filetype2 TEXTOverride file type detection for second file.
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--duckdb-threads INTEGERNumber of threads for DuckDB engine.
--duckdb-memory TEXTMemory limit for DuckDB (e.g., '4GB', '512MB').
--duckdb-temp-dir TEXTTemporary directory for DuckDB.
--progress / --no-progressShow progress bar.--no-progress
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for the first file.0
--table2, --sheet2 TEXTTable or sheet name for the second file (Excel, SQLite, lakehouse).
--start-page2 INTEGERSheet index (0-based) for the second file.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
-O, --format-out TEXTOutput format (e.g. csv, jsonl, parquet). Defaults to the --output extension, or to the input's text format on stdout.

See also shared options.