Skip to main content

schema

Generates data schemas from files. Supports multiple output formats including YAML, JSON, Cerberus, JSON Schema, Avro, and Parquet.

# Generate schema in default YAML format
undatum schema data.jsonl

# Generate schema in JSON Schema format
undatum schema data.jsonl --format jsonschema

# Generate schema in Avro format
undatum schema data.jsonl --format avro

# Generate schema in Parquet format
undatum schema data.jsonl --format parquet

# Generate Cerberus schema (for backward compatibility with deprecated `scheme` command)
undatum schema data.jsonl --format cerberus

# Save to file
undatum schema data.jsonl --output schema.yaml

# Nested JSONL: unfold dict fields onto dotted paths
undatum schema nested.jsonl --flatten-nested --format jsonschema
undatum schema nested.jsonl --flatten-nested --max-nested-depth 2
undatum schema nested.jsonl --flatten-nested --keep-nested-parents

# Named Excel sheet
undatum schema workbook.xlsx --table Sheet2

# Validate rows against the inferred schema (not rule-pack validation)
undatum schema data.jsonl --validate --format-out json
undatum schema data.jsonl --validate --strict
undatum schema data.jsonl --validate --sample-size 500

# Generate schema with AI-powered field documentation
undatum schema data.jsonl --autodoc --output schema.yaml

# Machine-readable schema (undatum.schema/1), usable as a drift or quality baseline
undatum schema data.jsonl --json

# Compare schemas: same as `diff --schema` and `schema-drift`
undatum schema diff data.csv data.jsonl
undatum schema drift data.csv data.jsonl --fail-on removed,type

undatum schema diff OLD NEW and undatum schema drift FILES... are spellings of diff --schema and schema-drift; they take those commands' options.

Supported schema formats:

  • yaml (default) - YAML format with full schema details
  • json - JSON format with full schema details
  • cerberus - Cerberus validation schema format (for backward compatibility with deprecated scheme command)
  • jsonschema - JSON Schema (W3C/IETF standard) - Use for API validation, OpenAPI specs, and tool integration
  • avro - Apache Avro schema format - Use for Kafka message schemas and Hadoop data pipelines
  • parquet - Parquet schema format - Use for data lake schemas and Parquet file metadata

Use cases:

  • JSON Schema: API documentation, data validation in web applications, OpenAPI specifications
  • Avro: Kafka message schemas, Hadoop ecosystem integration, schema registry compatibility
  • Parquet: Data lake schemas, Parquet file metadata, analytics pipeline definitions
  • Cerberus: Python data validation (schema --format cerberus; the legacy scheme command is deprecated)

Examples:

# Generate JSON Schema for API documentation
undatum schema api_data.jsonl --format jsonschema --output api_schema.json

# Generate Avro schema for Kafka
undatum schema events.jsonl --format avro --output events.avsc

# Generate Parquet schema for data lake
undatum schema data.csv --format parquet --output schema.json

# Generate Cerberus schema (deprecated, use schema command instead)
undatum schema data.jsonl --format cerberus --output validation_schema.json

Note: The scheme command is deprecated. Use undatum schema --format cerberus instead. The scheme command will show a deprecation warning but continues to work for backward compatibility.

Reference​

Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample of records · Engines: auto, duckdb, python

undatum schema [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-O, --format-out TEXTOutput format: 'text' (default), 'json', or 'yaml'.text
--format TEXTSchema format: 'yaml' (default), 'json', 'cerberus', 'jsonschema', 'avro', or 'parquet'. Overrides outtype when specified.
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
--autodoc / --no-autodocEnable AI-powered automatic field documentation.--no-autodoc
--lang TEXTLanguage for AI-generated documentation (default: 'English').English
--ai-provider TEXTAI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml).
--ai-model TEXTModel name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI).
--ai-base-url TEXTBase URL for AI API (optional, uses provider-specific defaults if not specified).
--pii-mask-samples / --no-pii-mask-samplesMask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio.
-e, --engine [auto|duckdb|python]Processing engine: auto (default), duckdb, or python.
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--start-page INTEGERSheet index (0-based) for Excel files.0
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--no-keep-nested-parents
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--validateValidate rows against an inferred schema (iterabledata schema.validate). Use undatum validate for rule packs.
--strictWith --validate, flag extra fields not present in the inferred schema.
--sample-size INTEGERWith --validate, rows to sample when inferring the schema (engine default 10000).
--jsonPrint the result as one JSON document (same as --format-out json).

Deprecated spellings (removed in 2.0): --outtype → --format-out.

See also shared options.