Skip to main content

stats / profile

Generates statistics and profiling metrics about your dataset. DuckDB is selected automatically for supported formats (CSV, JSONL, JSON, Parquet) and is typically much faster.

profile is a deprecated alias of stats (it prints a warning and will be removed in 2.0).

# Basic statistics
undatum stats data.jsonl


# Date detection is on by default; disable with --no-checkdates
undatum stats data.csv --no-checkdates

# Force DuckDB
undatum stats data.parquet --engine duckdb

# Machine-readable JSON (also used when --output ends in .json)
undatum stats data.csv --format-out json --output stats.json

# HTML or Markdown profiling report (also inferred from --output extension)
undatum stats data.csv --format-out html --output profile.html
undatum stats data.csv --output profile.md

# Nested JSONL: unfold dict fields onto dotted paths
undatum stats nested.jsonl --flatten-nested --format-out json
undatum stats nested.jsonl --flatten-nested --max-nested-depth 2
undatum stats nested.jsonl --flatten-nested --no-keep-nested-parents

# Named Excel sheet
undatum stats workbook.xlsx --table Sheet2

What is filled depends on the engine. DuckDB populates missing rates, type categories, and distribution stats. --engine python still reports field names, uniqueness, and lengths; mean / median / stddev / type_category / missing_rate are often empty.

DuckDB statistics include:

  • Field types and array flags
  • Missing value rates (count and percentage)
  • Cardinality (distinct counts and percentages)
  • Type inference: categorical, numerical, or text
  • Distribution for numerical fields: mean, median, min, max, stddev (not percentiles)
  • Unique value counts and percentages
  • Min/max/average lengths
  • Date field detection (--checkdates / --no-checkdates; default on)

Other options: --dictshare, --threads, --progress / --no-progress, --zipfile, --engine auto|duckdb|iterable, --format-out json|html|markdown. Also accepts --table, --flatten-nested, --on-error, --error-log, --quotechar, and --trust (shared options).

Profiling metrics (DuckDB)​

Missing value analysis: count and percentage of missing/null values per field. Example: 5 (2.5%) means 5 missing values out of 200 records.

Cardinality: distinct count and percentage of distinct values. High cardinality (IDs, timestamps) vs low cardinality (status, category).

Type inference:

  • categorical — low cardinality, typically string-like
  • numerical — numeric types
  • text — everything else (there is no Mixed category)

Distribution (numerical fields): mean (μ), median (m), min, max, standard deviation. Example: μ=42.5, m=40.0.

Use cases​

# Profile dataset to identify quality issues
undatum stats data.csv

# Look for high missing rates, unexpected cardinality, and min/max outliers

# Understand structure before processing
undatum stats data.jsonl --format-out json --output profile.json

Reference​

Reads: any readable format · Writes: report (see --format-out) · Memory: streaming aggregates; distinct values grow with cardinality · Engines: auto, duckdb, python

undatum stats [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--dictshare INTEGERDictionary share threshold (0-100) for type detection.
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl').
-O, --format-out TEXTOutput format: 'json', 'html', or 'markdown' (also inferred from --output extension).
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--verbose / --no-verboseEnable verbose logging output.--no-verbose
--zipfile / --no-zipfileTreat input file as a ZIP archive.--no-zipfile
--checkdates / --no-checkdatesEnable automatic date field detection.--checkdates
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--progress / --no-progressShow progress bar (default: True).--progress
--no-progress / --no-no-progressDisable progress bar (for non-interactive use).--no-no-progress
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--threads INTEGERWorker processes for Python/iterable-engine chunk parallelism. Omit for sequential. For DuckDB, prefer --duckdb-threads / engine settings.
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--jsonPrint the result as one JSON document (same as --format-out json).

See also shared options.