Skip to main content

diff

Compares two files and shows differences (added, removed, and changed rows).

# Compare files by key
undatum diff file1.csv file2.csv --key id
undatum diff workbook.xlsx other.xlsx --table Sheet2 --table2 Cities --key city
undatum diff nested1.jsonl nested2.jsonl --key name --flatten-nested

# Ignore order and show summary only (good for CI)
undatum diff file1.parquet file2.parquet --ignore-order --summary-only

# Output detailed diff to Markdown with numeric tolerance
undatum diff file1.csv file2.csv \
--key user_id \
--numeric-tolerance 0.001 \
--format-out markdown \
--output diff.md

# Fail CI when change thresholds are exceeded
undatum diff file1.csv file2.csv \
--key id \
--max-added-rows 10 \
--max-removed-rows 5 \
--max-changed-rows 0

Schema changes​

--schema compares the two schemas instead of the rows: added and removed fields, type and nullability changes, and likely renames (similar names or largely the same values).

undatum diff --schema data.csv data.jsonl
undatum diff --schema data.csv data.parquet --fail-on removed,type --json

--fail-on makes the command exit with 1 for the given kinds of change. The second argument can also be a declared schema (undatum schema --json output, JSON Schema, Frictionless). For many files against one baseline, see schema-drift.

Reference​

Reads: any readable format · Writes: report (see --format-out) · Memory: keys of both files in memory · Engines: python

undatum diff [OPTIONS] FILE1 FILE2
ArgumentDescription
FILE1Path to first input file. (required)
FILE2Path to second input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--key TEXTComma-separated list of key field names to compare on.
-O, --format-out TEXTDetailed output format: json, csv, markdown, html, or unified.
--summary-only / --no-summary-onlyShow summary only (suppress detailed output).--no-summary-only
--ignore-order / --no-ignore-orderTreat datasets as unordered sets when no key is provided.--no-ignore-order
--numeric-tolerance FLOATNumeric tolerance for float comparisons.
--ignore-case / --no-ignore-caseCase-insensitive comparison for strings.--no-ignore-case
--max-added-rows INTEGERFail if added rows exceed this threshold.
--max-removed-rows INTEGERFail if removed rows exceed this threshold.
--max-changed-rows INTEGERFail if changed rows exceed this threshold.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl').
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for the first file.0
--table2, --sheet2 TEXTTable or sheet name for the second file (Excel, SQLite, lakehouse).
--start-page2 INTEGERSheet index (0-based) for the second file.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--jsonPrint the result as one JSON document (same as --format-out json).
--schemaCompare the schemas (added, removed, retyped fields, renames) instead of rows.
--fail-on TEXTWith --schema: exit with 1 on these changes (added, removed, type, nullability, any).

Deprecated spellings (removed in 2.0): --format → --format-out, --output-format → --format-out.

See also shared options.