analyze
Analyzes data files and provides human-readable insights about structure, encoding, fields, and data types. With --autodoc, automatically generates field descriptions and dataset summaries using AI.
--autodoc uses the same iterabledata providers as ai doc; prefer ai doc for block-based documentation.
# Basic analysis
undatum analyze data.jsonl
# With AI-powered documentation
undatum analyze data.jsonl --autodoc
# Using a supported autodoc provider
undatum analyze data.jsonl --autodoc --ai-provider openai --ai-model gpt-4o-mini
# Output to file (format inferred from --output, or set --format-out)
undatum analyze data.jsonl --output report.yaml --autodoc
# Named Excel sheet
undatum analyze workbook.xlsx --table Sheet2
# Nested JSONL: unfold dict fields onto dotted paths
undatum analyze nested.jsonl --flatten-nested
Output includes:
- File type, encoding, compression
- Number of records and fields
- Field types and structure
- Per-field uniqueness statistics (unique count, total count, uniqueness %)
- Table detection for nested data (JSON/XML)
- AI-generated field descriptions (with
--autodoc) - AI-generated dataset summary (with
--autodoc)
Read and scan options:
--delimiter— CSV/TSV separator (auto-detected when omitted: comma, semicolon, tab, or pipe)--quotechar— CSV quote character--encoding— file encoding--engine—auto(default),duckdb, orpython--table/--sheet— named Excel sheet or multi-table source--limit— max records to scan (default 10000)--no-scan/--no-stats— skip structure scan or uniqueness stats--use-pandas— use pandas for the analysis path--format-out/-O—text,json,yaml, ormarkdown(also inferred from--output);--jsonis short for-O json--lang— language for--autodoctext (defaultEnglish)
--autodoc providers (iterabledata's iterable.ai, the same as ai):
openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity,
openai-compatible
export OPENAI_API_KEY=sk-...
undatum analyze data.csv --autodoc --ai-provider openai --ai-model gpt-4o-mini
export ANTHROPIC_API_KEY=sk-ant-...
undatum analyze data.csv --autodoc --ai-provider anthropic
export GEMINI_API_KEY=...
undatum analyze data.csv --autodoc --ai-provider gemini
export OPENROUTER_API_KEY=sk-or-...
undatum analyze data.csv --autodoc --ai-provider openrouter --ai-model openai/gpt-4o-mini
# Local (no API key)
undatum analyze data.csv --autodoc --ai-provider ollama --ai-model llama3.2
undatum analyze data.csv --autodoc --ai-provider lmstudio --ai-model local-model
Without --ai-provider or UNDATUM_AI_PROVIDER, the first API key found in the environment
picks the provider (OPENAI_API_KEY, OPENROUTER_API_KEY, PERPLEXITY_API_KEY,
ANTHROPIC_API_KEY, GEMINI_API_KEY). --ai-model and --ai-base-url override the provider
defaults (OLLAMA_BASE_URL, LMSTUDIO_BASE_URL). A local provider that is not listening is
detected within 2 seconds; a failing provider skips the AI content with one warning.
Before a data sample goes to a remote provider, values in columns that look like personal data
are replaced with ***; --no-pii-mask-samples turns this off (see
PII masking).
Configuration (lowest to highest): environment (UNDATUM_AI_PROVIDER, provider API keys), undatum.yaml / ~/.undatum/config.yaml, then CLI flags. Inspect with undatum config show. See AI documentation and config.
Reference
Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample of --limit records · Engines: auto, duckdb, python
undatum analyze [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file to analyze. (required) |
| Option | Description | Default |
|---|---|---|
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--use-pandas / --no-use-pandas | Use pandas for data processing (may use more memory). | --no-use-pandas |
-O, --format-out TEXT | Output format: 'text' (default), 'json', 'yaml', or 'markdown'. Also inferred from --output (.json/.yaml/.yml/.md). | |
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
--autodoc / --no-autodoc | Enable AI-powered automatic field and dataset documentation. | --no-autodoc |
--lang TEXT | Language for AI-generated documentation (default: 'English'). | English |
--ai-provider TEXT | AI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml). | |
--ai-model TEXT | Model name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI). | |
--ai-base-url TEXT | Base URL for AI API (optional, uses provider-specific defaults if not specified). | |
--pii-mask-samples / --no-pii-mask-samples | Mask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio. | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | |
-n, --limit INTEGER | Maximum number of records to scan for schema inference. | 10000 |
--ignore-errors / --no-ignore-errors | Ignore parse errors in CSV/JSON files (default: True). | --ignore-errors |
--no-scan / --no-no-scan | Return file metadata only; skip structure scan. | --no-no-scan |
--no-stats / --no-no-stats | Skip uniqueness statistics in field analysis. | --no-no-stats |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
--json | Print the result as one JSON document (same as --format-out json). |
Deprecated spellings (removed in 2.0): --objects-limit → --limit, --outtype → --format-out.
See also shared options.