Skip to main content

schema-bulk

Extracts schemas from multiple files at once using a glob pattern or directory path. Either extracts distinct unique schemas (--mode distinct, default) or one schema per file (--mode perfile).

The CLI command is schema-bulk (hyphen), not schema_bulk.

--autodoc uses the same iterabledata providers as analyze and ai: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity and openai-compatible. Field descriptions send field names only.

# Distinct schemas across all CSV files in a directory
undatum schema-bulk "data/*.csv" --output schemas/

# One schema per file, JSON Schema format
undatum schema-bulk data/ --mode perfile --format jsonschema --output schemas/

# With AI-powered field documentation
undatum schema-bulk "data/*.jsonl" --autodoc --output schemas/

Reference​

Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample of records per file · Engines: auto, duckdb, python

undatum schema-bulk [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEGlob pattern or directory path for input files (e.g., 'data/*.csv' or 'data/'). (required)
OptionDescriptionDefault
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-O, --format-out TEXTOutput format: 'text' (default), 'json', or 'yaml'.text
--format TEXTSchema format: 'yaml' (default), 'json', 'cerberus', 'jsonschema', 'avro', or 'parquet'. Overrides outtype when specified.
-o, --output TEXTOutput directory path for schema files.
--mode TEXTExtraction mode: 'distinct' (extract unique schemas, default) or 'perfile' (one schema per file).distinct
--autodoc / --no-autodocEnable AI-powered automatic field documentation.--no-autodoc
--lang TEXTLanguage for AI-generated documentation (default: 'English').English
--ai-provider TEXTAI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml).
--ai-model TEXTModel name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI).
--ai-base-url TEXTBase URL for AI API (optional, uses provider-specific defaults if not specified).
--pii-mask-samples / --no-pii-mask-samplesMask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio.
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).

Deprecated spellings (removed in 2.0): --outtype → --format-out.

See also shared options.