Skip to main content

ai

AI-assisted workflows backed by iterabledata's iterable.ai stack. Subcommands: doc, filter, plan, and suggest.

Configure providers via undatum.yaml, ~/.undatum/config.yaml, environment variables, or CLI flags. The ai * subcommands use iterabledata's provider set (OpenAI, Anthropic, Gemini, Azure OpenAI, OpenRouter, Ollama, LM Studio, Perplexity). Legacy analyze --autodoc / schema --autodoc / doc --autodoc only accept openai, openrouter, ollama, lmstudio, and perplexity — see analyze.

For block-based documentation with schema enrichment, prefer ai doc over legacy analyze --autodoc / schema --autodoc.

ai doc​

Block-based dataset documentation. Default blocks: general, schema, quality, examples, statistics, agent_skill, codebook.

undatum ai doc data.csv
undatum ai doc data.csv --format-out json --blocks general,schema,quality
undatum ai doc workbook.xlsx --tables Sheet2 --cache --pii-mask-samples
undatum ai doc data.csv --context '{"title": "City register"}'
undatum ai doc data.csv --progress --job-id run-42
undatum ai doc data.csv --sample-size 20 --no-detect-constraints --no-statistics
undatum ai doc data.csv --temperature 0.2 --max-tokens 2048

Notable flags: --format, --blocks, --tables, --cache, --pii-mask-samples, --context, --progress, --job-id, --sample-size, --detect-constraints / --no-detect-constraints, --statistics / --no-statistics, --temperature, --max-tokens, --include-field-descriptions, --validate-output.

ai filter​

Translate a natural-language or expression filter. Use --apply to stream matching rows.

undatum ai filter "active users in New York" data.csv --apply
undatum ai filter "city is Dushanbe" workbook.xlsx --table Sheet2 --apply
undatum ai filter "age > 30" data.csv --sample-size 500
undatum ai filter "name == 'Alice'" quoted.csv --quotechar "'"
undatum ai filter "lat > 40" nested.jsonl --flatten-nested --apply

--flatten-nested, --max-nested-depth, and --keep-nested-parents apply when unfolding nested fields for schema context and --apply. See shared options.

ai plan​

Produce a declarative conversion plan between two paths or format ids. No conversion is performed. Both arguments are positional (not --to).

undatum ai plan data.csv data.parquet
undatum ai plan data.json data.geojson --use-llm

--use-llm adds LLM reasoning on top of catalog metadata.

ai suggest​

Suggest a transform spec from a natural-language goal. --apply writes transformed rows (confirm unless --yes).

undatum ai suggest data.csv "normalize phone numbers"
undatum ai suggest data.csv "normalize phone numbers" --sample-size 20
undatum ai suggest data.csv "rename id to user_id" --apply --yes --output out.jsonl

Reference​

undatum ai doc​

undatum ai doc [OPTIONS] FILENAME
ArgumentDescription
FILENAMEInput data file or source to document. (required)
OptionDescriptionDefault
-o, --output TEXTOutput file path. Prints to stdout if not given.
-O, --format-out TEXTOutput format: markdown, json, yaml, html, text.markdown
--blocks TEXTComma-separated documentation blocks (e.g. 'general,schema,quality').
--language TEXTLanguage for generated content.English
--pii-detect / --no-pii-detectDetect PII fields (requires metacrafter).--no-pii-detect
--pii-mask-samplesMask detected PII in sample rows sent to the LLM.
--semantic-types / --no-semantic-typesDetect semantic types (requires metacrafter).--no-semantic-types
--tables TEXTComma-separated table/sheet names for multi-table sources.
--cacheCache generated documentation by content hash.
--include-field-descriptionsGenerate per-field descriptions (non-block generate path).
--validate-outputValidate JSON documentation against Pydantic models.
--context TEXTJSON object of extra prompt context, e.g. '{"title": "Sales"}'.
--progressPrint documentation generation stages to stderr.
--sample-size INTEGEROverride sample row count for documentation (engine default if omitted).
--detect-constraints / --no-detect-constraintsDetect min/max/length/enum constraints during schema inference (block path).--detect-constraints
--statistics / --no-statisticsCompute statistics used by statistics/quality/codebook blocks.--statistics
--temperature FLOATLLM sampling temperature (engine default if omitted).
--max-tokens INTEGERMaximum tokens per LLM documentation block (engine default if omitted).
--job-id TEXTStable job identifier for documentation progress and JSON results.
--provider TEXTLLM provider override.
--model TEXTModel name override.
--api-key TEXTAPI key override.
--base-url TEXTBase URL override (local providers).
--verbose / --no-verboseEnable verbose logging output.--no-verbose

Deprecated spellings (removed in 2.0): --format → --format-out.

undatum ai filter​

undatum ai filter [OPTIONS] EXPRESSION [FILENAME]
ArgumentDescription
EXPRESSIONNatural-language or DSL filter (e.g. 'rows where age > 30'). (required)
FILENAMEOptional data file (used for schema context and --apply).
OptionDescriptionDefault
--apply / --no-applyApply the translated filter and output matching rows as JSONL.--no-apply
-o, --output TEXTOutput file for --apply (default stdout).
--provider TEXTLLM provider override.
--model TEXTModel name override.
--api-key TEXTAPI key override.
--base-url TEXTBase URL override (local providers).
--verbose / --no-verboseEnable verbose logging output.--no-verbose
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--sample-size INTEGERRows to sample when inferring schema context for a file (engine default 10000).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths for schema context and --apply.
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents

undatum ai plan​

undatum ai plan [OPTIONS] SOURCE TARGET
ArgumentDescription
SOURCESource file/path. (required)
TARGETTarget file/path. (required)
OptionDescriptionDefault
--use-llm / --no-use-llmUse LLM reasoning in addition to catalog metadata.--no-use-llm
--provider TEXTLLM provider override.
--model TEXTModel name override.
--api-key TEXTAPI key override.
--base-url TEXTBase URL override (local providers).
--verbose / --no-verboseEnable verbose logging output.--no-verbose

undatum ai suggest​

undatum ai suggest [OPTIONS] FILENAME GOAL
ArgumentDescription
FILENAMEInput data file. (required)
GOALNatural-language description of the desired result. (required)
OptionDescriptionDefault
--provider TEXTLLM provider override.
--model TEXTModel name override.
--api-key TEXTAPI key override.
--base-url TEXTBase URL override (local providers).
--applyApply the suggested spec and write transformed rows.
-o, --output TEXTOutput file for --apply (JSONL; default stdout).
-y, --yesDo not prompt before applying the transform.
--sample-size INTEGEROverride sample row count sent to the suggestion engine (default 5).
--verbose / --no-verboseEnable verbose logging output.--no-verbose
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).

See also shared options.