Skip to main content

doc

Generates dataset documentation with schema, statistics, and samples in Markdown (default), JSON, YAML, or text. Supports AI-powered descriptions with --autodoc. Also available as the document alias.

# Markdown documentation (default)
undatum doc data.jsonl

# JSON documentation with samples
undatum doc data.jsonl --format-out json --sample-size 5 --output report.json
undatum doc nested.jsonl --flatten-nested --format-out json

# With AI-powered descriptions
undatum doc data.csv --autodoc --ai-provider openai --ai-model gpt-4o-mini

Output includes:

  • Dataset metadata and summary counts
  • Schema fields with types and descriptions
  • Field-level uniqueness statistics (when available)
  • Sample records (configurable via --sample-size)

Extended metadata and PII options:

  • --semantic-types: annotate fields with semantic types (requires metacrafter CLI)
  • --pii-detect: detect PII fields and include a PII summary (requires metacrafter CLI)
  • --pii-mask-samples: redact detected PII values in samples (use with --pii-detect); with --autodoc, samples sent to remote AI providers are masked by default and --no-pii-mask-samples turns that off
# Semantic typing and PII summary
undatum doc data.csv --semantic-types --pii-detect --format-out json

# Mask PII values in samples
undatum doc data.csv --pii-detect --pii-mask-samples --format-out json

Optional dependencies:

  • metacrafter (for semantic types and PII detection)
  • langdetect (for language detection in metadata)

Reference​

Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample plus streaming statistics · Engines: auto, duckdb, python

undatum doc [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file to document. (required)
OptionDescriptionDefault
-O, --format-out TEXTOutput format: 'markdown' (default), 'json', 'yaml', or 'text'.markdown
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--sample-size INTEGERNumber of sample records to include (default: 10).10
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
--tagname TEXTXML tag name that contains individual records.
--start-line INTEGERLine number (0-based) to start reading from.0
--start-page INTEGERPage number (0-based) to start from for Excel files.0
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl').
--autodoc / --no-autodocEnable AI-powered automatic field and dataset documentation.--no-autodoc
--lang TEXTLanguage for AI-generated documentation (default: 'English').English
--ai-provider TEXTAI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml).
--ai-model TEXTModel name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI).
--ai-base-url TEXTBase URL for AI API (optional, uses provider-specific defaults if not specified).
--semantic-types / --no-semantic-typesEnable semantic type annotations using Metacrafter.--no-semantic-types
--pii-detect / --no-pii-detectEnable PII detection using Metacrafter.--no-pii-detect
--pii-mask-samples / --no-pii-mask-samplesRedact detected PII in sample records (with --pii-detect) and mask likely PII in rows sent to the AI provider (--autodoc masks for remote providers by default).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents

Deprecated spellings (removed in 2.0): --format → --format-out.

See also shared options.