doc
Generates dataset documentation with schema, statistics, and samples in Markdown (default), JSON, YAML, or text. Supports AI-powered descriptions with --autodoc. Also available as the document alias.
# Markdown documentation (default)
undatum doc data.jsonl
# JSON documentation with samples
undatum doc data.jsonl --format-out json --sample-size 5 --output report.json
undatum doc nested.jsonl --flatten-nested --format-out json
# With AI-powered descriptions
undatum doc data.csv --autodoc --ai-provider openai --ai-model gpt-4o-mini
Output includes:
- Dataset metadata and summary counts
- Schema fields with types and descriptions
- Field-level uniqueness statistics (when available)
- Sample records (configurable via
--sample-size)
Extended metadata and PII options:
--semantic-types: annotate fields with semantic types (requiresmetacrafterCLI)--pii-detect: detect PII fields and include a PII summary (requiresmetacrafterCLI)--pii-mask-samples: redact detected PII values in samples (use with--pii-detect); with--autodoc, samples sent to remote AI providers are masked by default and--no-pii-mask-samplesturns that off
# Semantic typing and PII summary
undatum doc data.csv --semantic-types --pii-detect --format-out json
# Mask PII values in samples
undatum doc data.csv --pii-detect --pii-mask-samples --format-out json
Optional dependencies:
metacrafter(for semantic types and PII detection)langdetect(for language detection in metadata)
Reference
Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample plus streaming statistics · Engines: auto, duckdb, python
undatum doc [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file to document. (required) |
| Option | Description | Default |
|---|---|---|
-O, --format-out TEXT | Output format: 'markdown' (default), 'json', 'yaml', or 'text'. | markdown |
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
--sample-size INTEGER | Number of sample records to include (default: 10). | 10 |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
-d, --delimiter TEXT | CSV delimiter character (auto-detected when omitted). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--encoding TEXT | File encoding (e.g., 'utf8', 'latin1'). | |
--tagname TEXT | XML tag name that contains individual records. | |
--start-line INTEGER | Line number (0-based) to start reading from. | 0 |
--start-page INTEGER | Page number (0-based) to start from for Excel files. | 0 |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
-F, --format-in TEXT | Override input file format detection (e.g., 'csv', 'jsonl'). | |
--autodoc / --no-autodoc | Enable AI-powered automatic field and dataset documentation. | --no-autodoc |
--lang TEXT | Language for AI-generated documentation (default: 'English'). | English |
--ai-provider TEXT | AI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml). | |
--ai-model TEXT | Model name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI). | |
--ai-base-url TEXT | Base URL for AI API (optional, uses provider-specific defaults if not specified). | |
--semantic-types / --no-semantic-types | Enable semantic type annotations using Metacrafter. | --no-semantic-types |
--pii-detect / --no-pii-detect | Enable PII detection using Metacrafter. | --no-pii-detect |
--pii-mask-samples / --no-pii-mask-samples | Redact detected PII in sample records (with --pii-detect) and mask likely PII in rows sent to the AI provider (--autodoc masks for remote providers by default). | |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --keep-nested-parents |
Deprecated spellings (removed in 2.0): --format → --format-out.
See also shared options.