Skip to main content

analyze

Analyzes data files and provides human-readable insights about structure, encoding, fields, and data types. With --autodoc, automatically generates field descriptions and dataset summaries using AI.

--autodoc uses the same iterabledata providers as ai doc; prefer ai doc for block-based documentation.

# Basic analysis
undatum analyze data.jsonl

# With AI-powered documentation
undatum analyze data.jsonl --autodoc

# Using a supported autodoc provider
undatum analyze data.jsonl --autodoc --ai-provider openai --ai-model gpt-4o-mini

# Output to file (format inferred from --output, or set --format-out)
undatum analyze data.jsonl --output report.yaml --autodoc

# Named Excel sheet
undatum analyze workbook.xlsx --table Sheet2

# Nested JSONL: unfold dict fields onto dotted paths
undatum analyze nested.jsonl --flatten-nested

Output includes:

  • File type, encoding, compression
  • Number of records and fields
  • Field types and structure
  • Per-field uniqueness statistics (unique count, total count, uniqueness %)
  • Table detection for nested data (JSON/XML)
  • AI-generated field descriptions (with --autodoc)
  • AI-generated dataset summary (with --autodoc)

Read and scan options:

  • --delimiter — CSV/TSV separator (auto-detected when omitted: comma, semicolon, tab, or pipe)
  • --quotechar — CSV quote character
  • --encoding — file encoding
  • --engine — auto (default), duckdb, or python
  • --table / --sheet — named Excel sheet or multi-table source
  • --limit — max records to scan (default 10000)
  • --no-scan / --no-stats — skip structure scan or uniqueness stats
  • --use-pandas — use pandas for the analysis path
  • --format-out / -O — text, json, yaml, or markdown (also inferred from --output); --json is short for -O json
  • --lang — language for --autodoc text (default English)

--autodoc providers (iterabledata's iterable.ai, the same as ai):

openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity, openai-compatible

export OPENAI_API_KEY=sk-...
undatum analyze data.csv --autodoc --ai-provider openai --ai-model gpt-4o-mini

export ANTHROPIC_API_KEY=sk-ant-...
undatum analyze data.csv --autodoc --ai-provider anthropic

export GEMINI_API_KEY=...
undatum analyze data.csv --autodoc --ai-provider gemini

export OPENROUTER_API_KEY=sk-or-...
undatum analyze data.csv --autodoc --ai-provider openrouter --ai-model openai/gpt-4o-mini

# Local (no API key)
undatum analyze data.csv --autodoc --ai-provider ollama --ai-model llama3.2
undatum analyze data.csv --autodoc --ai-provider lmstudio --ai-model local-model

Without --ai-provider or UNDATUM_AI_PROVIDER, the first API key found in the environment picks the provider (OPENAI_API_KEY, OPENROUTER_API_KEY, PERPLEXITY_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY). --ai-model and --ai-base-url override the provider defaults (OLLAMA_BASE_URL, LMSTUDIO_BASE_URL). A local provider that is not listening is detected within 2 seconds; a failing provider skips the AI content with one warning.

Before a data sample goes to a remote provider, values in columns that look like personal data are replaced with ***; --no-pii-mask-samples turns this off (see PII masking).

Configuration (lowest to highest): environment (UNDATUM_AI_PROVIDER, provider API keys), undatum.yaml / ~/.undatum/config.yaml, then CLI flags. Inspect with undatum config show. See AI documentation and config.

Reference​

Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample of --limit records · Engines: auto, duckdb, python

undatum analyze [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file to analyze. (required)
OptionDescriptionDefault
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--use-pandas / --no-use-pandasUse pandas for data processing (may use more memory).--no-use-pandas
-O, --format-out TEXTOutput format: 'text' (default), 'json', 'yaml', or 'markdown'. Also inferred from --output (.json/.yaml/.yml/.md).
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--autodoc / --no-autodocEnable AI-powered automatic field and dataset documentation.--no-autodoc
--lang TEXTLanguage for AI-generated documentation (default: 'English').English
--ai-provider TEXTAI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml).
--ai-model TEXTModel name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI).
--ai-base-url TEXTBase URL for AI API (optional, uses provider-specific defaults if not specified).
--pii-mask-samples / --no-pii-mask-samplesMask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
-n, --limit INTEGERMaximum number of records to scan for schema inference.10000
--ignore-errors / --no-ignore-errorsIgnore parse errors in CSV/JSON files (default: True).--ignore-errors
--no-scan / --no-no-scanReturn file metadata only; skip structure scan.--no-no-scan
--no-stats / --no-no-statsSkip uniqueness statistics in field analysis.--no-no-stats
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--jsonPrint the result as one JSON document (same as --format-out json).

Deprecated spellings (removed in 2.0): --objects-limit → --limit, --outtype → --format-out.

See also shared options.