stats / profile
Generates statistics and profiling metrics about your dataset. DuckDB is selected automatically for supported formats (CSV, JSONL, JSON, Parquet) and is typically much faster.
profile is a deprecated alias of stats (it prints a warning and will be removed in 2.0).
# Basic statistics
undatum stats data.jsonl
# Date detection is on by default; disable with --no-checkdates
undatum stats data.csv --no-checkdates
# Force DuckDB
undatum stats data.parquet --engine duckdb
# Machine-readable JSON (also used when --output ends in .json)
undatum stats data.csv --format-out json --output stats.json
# HTML or Markdown profiling report (also inferred from --output extension)
undatum stats data.csv --format-out html --output profile.html
undatum stats data.csv --output profile.md
# Nested JSONL: unfold dict fields onto dotted paths
undatum stats nested.jsonl --flatten-nested --format-out json
undatum stats nested.jsonl --flatten-nested --max-nested-depth 2
undatum stats nested.jsonl --flatten-nested --no-keep-nested-parents
# Named Excel sheet
undatum stats workbook.xlsx --table Sheet2
What is filled depends on the engine. DuckDB populates missing rates, type categories, and distribution stats. --engine python still reports field names, uniqueness, and lengths; mean / median / stddev / type_category / missing_rate are often empty.
DuckDB statistics include:
- Field types and array flags
- Missing value rates (count and percentage)
- Cardinality (distinct counts and percentages)
- Type inference:
categorical,numerical, ortext - Distribution for numerical fields: mean, median, min, max, stddev (not percentiles)
- Unique value counts and percentages
- Min/max/average lengths
- Date field detection (
--checkdates/--no-checkdates; default on)
Other options: --dictshare, --threads, --progress / --no-progress, --zipfile, --engine auto|duckdb|iterable, --format-out json|html|markdown. Also accepts --table, --flatten-nested, --on-error, --error-log, --quotechar, and --trust (shared options).
Profiling metrics (DuckDB)
Missing value analysis: count and percentage of missing/null values per field. Example: 5 (2.5%) means 5 missing values out of 200 records.
Cardinality: distinct count and percentage of distinct values. High cardinality (IDs, timestamps) vs low cardinality (status, category).
Type inference:
- categorical — low cardinality, typically string-like
- numerical — numeric types
- text — everything else (there is no
Mixedcategory)
Distribution (numerical fields): mean (μ), median (m), min, max, standard deviation. Example: μ=42.5, m=40.0.
Use cases
# Profile dataset to identify quality issues
undatum stats data.csv
# Look for high missing rates, unexpected cardinality, and min/max outliers
# Understand structure before processing
undatum stats data.jsonl --format-out json --output profile.json
Reference
Reads: any readable format · Writes: report (see --format-out) · Memory: streaming aggregates; distinct values grow with cardinality · Engines: auto, duckdb, python
undatum stats [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file. (required) |
| Option | Description | Default |
|---|---|---|
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
--dictshare INTEGER | Dictionary share threshold (0-100) for type detection. | |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl'). | |
-O, --format-out TEXT | Output format: 'json', 'html', or 'markdown' (also inferred from --output extension). | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
--zipfile / --no-zipfile | Treat input file as a ZIP archive. | --no-zipfile |
--checkdates / --no-checkdates | Enable automatic date field detection. | --checkdates |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | |
--progress / --no-progress | Show progress bar (default: True). | --progress |
--no-progress / --no-no-progress | Disable progress bar (for non-interactive use). | --no-no-progress |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--threads INTEGER | Worker processes for Python/iterable-engine chunk parallelism. Omit for sequential. For DuckDB, prefer --duckdb-threads / engine settings. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--json | Print the result as one JSON document (same as --format-out json). |
See also shared options.