schema-bulk
Extracts schemas from multiple files at once using a glob pattern or directory path. Either extracts distinct unique schemas (--mode distinct, default) or one schema per file (--mode perfile).
The CLI command is schema-bulk (hyphen), not schema_bulk.
--autodoc uses the same iterabledata providers as analyze and ai: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity and openai-compatible. Field descriptions send field names only.
# Distinct schemas across all CSV files in a directory
undatum schema-bulk "data/*.csv" --output schemas/
# One schema per file, JSON Schema format
undatum schema-bulk data/ --mode perfile --format jsonschema --output schemas/
# With AI-powered field documentation
undatum schema-bulk "data/*.jsonl" --autodoc --output schemas/
Reference
Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample of records per file · Engines: auto, duckdb, python
undatum schema-bulk [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Glob pattern or directory path for input files (e.g., 'data/*.csv' or 'data/'). (required) |
| Option | Description | Default |
|---|---|---|
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-O, --format-out TEXT | Output format: 'text' (default), 'json', or 'yaml'. | text |
--format TEXT | Schema format: 'yaml' (default), 'json', 'cerberus', 'jsonschema', 'avro', or 'parquet'. Overrides outtype when specified. | |
-o, --output TEXT | Output directory path for schema files. | |
--mode TEXT | Extraction mode: 'distinct' (extract unique schemas, default) or 'perfile' (one schema per file). | distinct |
--autodoc / --no-autodoc | Enable AI-powered automatic field documentation. | --no-autodoc |
--lang TEXT | Language for AI-generated documentation (default: 'English'). | English |
--ai-provider TEXT | AI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml). | |
--ai-model TEXT | Model name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI). | |
--ai-base-url TEXT | Base URL for AI API (optional, uses provider-specific defaults if not specified). | |
--pii-mask-samples / --no-pii-mask-samples | Mask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio. | |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). |
Deprecated spellings (removed in 2.0): --outtype → --format-out.
See also shared options.