validate
Validates data against validation rules. Supports two modes: rich validation with rule files (recommended) and legacy single-rule mode (backward compatible).
Rich Validation with Rule Files
Use YAML/JSON rule files for comprehensive, reusable validation:
# Validate with rule file
undatum validate data.csv --rules validation-rules.yml
undatum validate workbook.xlsx --table Sheet2 --rules validation-rules.yml
undatum validate nested.jsonl --rules rules.yml --flatten-nested
undatum validate data.csv --rules rules.yml --filter '`status` == "active"'
# Filter by severity
undatum validate data.jsonl --rules rules.yml --severity error
# JSON output for CI/CD integration
undatum validate data.csv --rules rules.yml --format-out json
# Generate detailed violation report
undatum validate data.jsonl --rules rules.yml --violation-report violations.json
# Treat warnings as errors
undatum validate data.csv --rules rules.yml --fail-on-warnings
Rule File Format:
Rule files support field-level and cross-field validation with severity levels. Built-in
formats (country, currency, language, date, phone, iban, uuid, pattern, ...),
unique, and references to another file are listed on Validation rules
and by undatum validate --list-rules. -O jsonl prints one violation per line.
rules:
# Field-level rules
- field: email
name: Email Required
description: Email field must be present
required: true
type: string
format: email
severity: error
- field: age
name: Age Range
description: Age must be between 0 and 120
type: number
min: 0
max: 120
severity: warning
- field: status
name: Status Values
type: string
enum: [active, inactive, pending]
severity: error
# Cross-field validation
- type: cross-field
name: Date Range Validation
description: End date must be after start date
condition: "end_date >= start_date"
fields: [start_date, end_date]
severity: error
Rule Types:
- Required:
required: true- Field must be present and non-empty - Type:
type: string|number|integer|float|boolean- Value type validation - Format:
format: email|url|uuid- Format validation - Range:
min,maxfor numbers;min_length,max_lengthfor strings - Enum:
enum: [value1, value2, ...]- Whitelist validation - Pattern:
pattern: 'regex'- Regular expression validation - Custom:
custom: 'rule_name'- Use custom validation function from VALIDATION_RULEMAP - Cross-field:
type: cross-fieldwithconditionexpression
Severity Levels:
error: Hard errors that should block processingwarning: Soft warnings that don't block processinginfo: Informational violations
Violation Reporting:
The validation command provides comprehensive reporting:
- Summary statistics: Total violations by severity, by field, by rule
- Detailed violations: Record-level violation details with context
- JSON output: Machine-readable format for CI/CD integration
- Violation report file: Detailed JSON report with all violations
Example Rule Files:
Example rule files are available in examples/validation-rules/:
basic-validation.yml- Common field-level validation rulescross-field-validation.yml- Cross-field validation examplescomplex-validation.yml- Comprehensive validation scenario
Legacy Mode (Backward Compatible)
Simple single-rule validation for quick checks:
# Validate email addresses
undatum validate --rule common.email --fields email data.jsonl
# Validate Russian INN
undatum validate --rule ru.org.inn --fields VendorINN data.jsonl --mode stats
# Output invalid records
undatum validate --rule ru.org.inn --fields VendorINN data.jsonl --mode invalid
Available built-in validation rules:
common.email- Email address validationcommon.url- URL validationru.org.inn- Russian organization INN identifierru.org.ogrn- Russian organization OGRN identifierinteger- Integer validation
Validation Best Practices
- Use errors for critical issues: Fields that must be correct for data processing
- Use warnings for data quality: Issues that should be reviewed but don't block processing
- Organize rules by domain: Group related rules in separate files (e.g.,
user-validation.yml,order-validation.yml) - Version control rule files: Track rule changes and share across teams
- Use cross-field rules sparingly: They're more complex and slower to evaluate
- Test rules incrementally: Start with basic rules, add complexity as needed
Reference
Reads: any readable format · Writes: validation report or the (in)valid rows · Memory: streaming · Engines: python
undatum validate [OPTIONS] [INPUT_FILE]
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file (not needed with --list-rules). |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
-f, --fields TEXT | Comma-separated list of field names to validate (legacy mode). | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | utf8 |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl'). | |
--zipfile / --no-zipfile | Treat input file as a ZIP archive. | --no-zipfile |
--rule TEXT | Validation rule name (legacy mode, e.g., 'common.email', 'common.url'). | |
--filter, --filter-expr TEXT | Filter expression to apply before validation. | |
--mode TEXT | Legacy --rule output: 'invalid' (default), 'valid', 'all', or 'stats' (JSON counts). | invalid |
--rules TEXT | Path to YAML/JSON rule file for rich validation. | |
--severity TEXT | Filter violations by severity: 'error', 'warning', 'info', or 'all' (default). | all |
-O, --format-out TEXT | Output format: 'text' (default) or 'json'. | text |
--violation-report TEXT | Path to write detailed violation report (JSON format). | |
--fail-on-warnings / --no-fail-on-warnings | Treat warnings as errors (exit with non-zero code). | --no-fail-on-warnings |
--max-violations INTEGER | Maximum number of violations to display (default: 10 for text, 100 for JSON). | |
--progress / --no-progress | Show progress bar. | --no-progress |
--threads INTEGER | Worker processes for rule-file validation chunk parallelism. Omit for sequential processing. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--json | Print the result as one JSON document (same as --format-out json). | |
--list-rules | List the built-in rules (format: ...) and exit. |
Deprecated spellings (removed in 2.0): --output-format → --format-out.
See also shared options.