schema
Generates data schemas from files. Supports multiple output formats including YAML, JSON, Cerberus, JSON Schema, Avro, and Parquet.
# Generate schema in default YAML format
undatum schema data.jsonl
# Generate schema in JSON Schema format
undatum schema data.jsonl --format jsonschema
# Generate schema in Avro format
undatum schema data.jsonl --format avro
# Generate schema in Parquet format
undatum schema data.jsonl --format parquet
# Generate Cerberus schema (for backward compatibility with deprecated `scheme` command)
undatum schema data.jsonl --format cerberus
# Save to file
undatum schema data.jsonl --output schema.yaml
# Nested JSONL: unfold dict fields onto dotted paths
undatum schema nested.jsonl --flatten-nested --format jsonschema
undatum schema nested.jsonl --flatten-nested --max-nested-depth 2
undatum schema nested.jsonl --flatten-nested --keep-nested-parents
# Named Excel sheet
undatum schema workbook.xlsx --table Sheet2
# Validate rows against the inferred schema (not rule-pack validation)
undatum schema data.jsonl --validate --format-out json
undatum schema data.jsonl --validate --strict
undatum schema data.jsonl --validate --sample-size 500
# Generate schema with AI-powered field documentation
undatum schema data.jsonl --autodoc --output schema.yaml
# Machine-readable schema (undatum.schema/1), usable as a drift or quality baseline
undatum schema data.jsonl --json
# Compare schemas: same as `diff --schema` and `schema-drift`
undatum schema diff data.csv data.jsonl
undatum schema drift data.csv data.jsonl --fail-on removed,type
undatum schema diff OLD NEW and undatum schema drift FILES... are spellings of
diff --schema and schema-drift;
they take those commands' options.
Supported schema formats:
yaml(default) - YAML format with full schema detailsjson- JSON format with full schema detailscerberus- Cerberus validation schema format (for backward compatibility with deprecatedschemecommand)jsonschema- JSON Schema (W3C/IETF standard) - Use for API validation, OpenAPI specs, and tool integrationavro- Apache Avro schema format - Use for Kafka message schemas and Hadoop data pipelinesparquet- Parquet schema format - Use for data lake schemas and Parquet file metadata
Use cases:
- JSON Schema: API documentation, data validation in web applications, OpenAPI specifications
- Avro: Kafka message schemas, Hadoop ecosystem integration, schema registry compatibility
- Parquet: Data lake schemas, Parquet file metadata, analytics pipeline definitions
- Cerberus: Python data validation (
schema --format cerberus; the legacyschemecommand is deprecated)
Examples:
# Generate JSON Schema for API documentation
undatum schema api_data.jsonl --format jsonschema --output api_schema.json
# Generate Avro schema for Kafka
undatum schema events.jsonl --format avro --output events.avsc
# Generate Parquet schema for data lake
undatum schema data.csv --format parquet --output schema.json
# Generate Cerberus schema (deprecated, use schema command instead)
undatum schema data.jsonl --format cerberus --output validation_schema.json
Note: The scheme command is deprecated. Use undatum schema --format cerberus instead. The scheme command will show a deprecation warning but continues to work for backward compatibility.
Reference
Reads: any readable format · Writes: report (see --format-out) · Memory: bounded: a sample of records · Engines: auto, duckdb, python
undatum schema [OPTIONS] INPUT_FILE
| Argument | Description |
|---|---|
INPUT_FILE | Path to input file. (required) |
| Option | Description | Default |
|---|---|---|
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
-O, --format-out TEXT | Output format: 'text' (default), 'json', or 'yaml'. | text |
--format TEXT | Schema format: 'yaml' (default), 'json', 'cerberus', 'jsonschema', 'avro', or 'parquet'. Overrides outtype when specified. | |
-o, --output TEXT | Optional output file path. If not specified, prints to stdout. | |
--autodoc / --no-autodoc | Enable AI-powered automatic field documentation. | --no-autodoc |
--lang TEXT | Language for AI-generated documentation (default: 'English'). | English |
--ai-provider TEXT | AI provider: openai, anthropic, gemini, azure, openrouter, ollama, lmstudio, perplexity or openai-compatible (default: UNDATUM_AI_PROVIDER or the ai: section of undatum.yaml). | |
--ai-model TEXT | Model name to use (provider-specific, e.g., 'gpt-4o-mini' for OpenAI). | |
--ai-base-url TEXT | Base URL for AI API (optional, uses provider-specific defaults if not specified). | |
--pii-mask-samples / --no-pii-mask-samples | Mask likely PII in sample rows sent to the AI provider with --autodoc. Default: masked for remote providers, unmasked for ollama and lmstudio. | |
-e, --engine [auto|duckdb|python] | Processing engine: auto (default), duckdb, or python. | |
--table, --sheet TEXT | Table or sheet name for multi-table sources (Excel, SQLite, lakehouse). | |
--quotechar TEXT | CSV quote character (iterabledata default '"' when omitted). | |
--start-page INTEGER | Sheet index (0-based) for Excel files. | 0 |
--flatten-nested | Unfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat). | |
--max-nested-depth INTEGER | With --flatten-nested, maximum nest depth to unfold (engine default 5). | |
--keep-nested-parents / --no-keep-nested-parents | With --flatten-nested, keep parent dict/array fields alongside dotted children. | --no-keep-nested-parents |
--trust | Acknowledge pickle deserialization risk when reading pickle sources. | |
--on-error TEXT | Parse-error policy: raise (default), skip, or warn. | |
--error-log TEXT | Append parse errors as JSONL (use with --on-error skip or warn). | |
--validate | Validate rows against an inferred schema (iterabledata schema.validate). Use undatum validate for rule packs. | |
--strict | With --validate, flag extra fields not present in the inferred schema. | |
--sample-size INTEGER | With --validate, rows to sample when inferring the schema (engine default 10000). | |
--json | Print the result as one JSON document (same as --format-out json). |
Deprecated spellings (removed in 2.0): --outtype → --format-out.
See also shared options.