join
Performs relational joins between two files. Supports inner, left, right, and full outer joins.
Write with --output. A trailing path is not a positional argument.
# Inner join by key field
undatum join data1.csv data2.csv --on email --type inner --output output.csv
# Left join (keep all rows from first file)
undatum join data1.jsonl data2.jsonl --on id --type left --output output.jsonl
# Right join (keep all rows from second file)
undatum join data1.csv data2.csv --on id --type right --output output.csv
# Full outer join (keep all rows from both files)
undatum join data1.jsonl data2.jsonl --on id --type full --output output.jsonl
undatum join workbook.xlsx other.xlsx --table Sheet2 --table2 Cities --on city --output out.jsonl
undatum join left.jsonl right.jsonl --on capital_city.lat --flatten-nested --output out.jsonl
--on— join key field(s)--type inner|left|right|full--output— output path (stdout if omitted)--table/--sheetand--table2/--sheet2— named tables for each file--filetype1/--filetype2— override type detection per file--engine,--progress/--no-progress
Also accepts --flatten-nested, --on-error, --error-log, and --quotechar (shared options).
Reference
Reads: any readable format · Writes: any writable format; CSV/TSV/JSON/JSON Lines on stdout · Memory: Python engine indexes the second file in memory; DuckDB works out of core · Engines: auto, duckdb, python
undatum join [OPTIONS] FILE1 FILE2
| Argument | Description |
|---|---|
FILE1 | Path to first input file. (required) |
FILE2 | Path to second input file. (required) |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
--on TEXT | Comma-separated list of key field names to join on. | |
--type TEXT | Join type: 'inner' (default), 'left', 'right', or 'full'. | inner |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
--filetype1 TEXT | Override file type detection for first file. | |
--filetype2 TEXT | Override file type detection for second file. | |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--duckdb-threads INTEGER | Number of threads for DuckDB engine. | |
--duckdb-memory TEXT | Memory limit for DuckDB (e.g., '4GB', '512MB'). | |
--duckdb-temp-dir TEXT | Temporary directory for DuckDB. | |
--progress / --no-progress | Show progress bar. | --no-progress |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--start-page INTEGER | Sheet index (0-based) for the first file. | 0 |
--table2, --sheet2 TEXT | Table or sheet name for the second file (Excel, SQLite, lakehouse). | |
--start-page2 INTEGER | Sheet index (0-based) for the second file. | 0 |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
-O, --format-out TEXT | Output format (e.g. csv, jsonl, parquet). Defaults to the --output extension, or to the input's text format on stdout. |
See also shared options.