Skip to main content

headers

Extracts field names from data files (CSV, JSON Lines, BSON, XML, Excel, and other iterabledata sources).

Default scan --limit is 10000. --format-out json (or --output ending in .json) writes JSON instead of text.

undatum headers data.jsonl
undatum headers data.csv --limit 50000
undatum headers data.csv --format-out json --output fields.json
undatum headers workbook.xlsx --table Sheet2

Reference​

Reads: any readable format · Writes: field names (text or JSON) · Memory: streaming · Engines: python

undatum headers [OPTIONS] INPUT_FILE
ArgumentDescription
INPUT_FILEPath to input file. (required)
OptionDescriptionDefault
-o, --output TEXTOptional output file path. If not specified, prints to stdout.
-f, --fields TEXTField filter (kept for API compatibility, not currently used).
-d, --delimiter TEXTCSV delimiter character (auto-detected when omitted).
--quotechar TEXTCSV quote character (iterabledata default '"' when omitted).
--encoding TEXTFile encoding (e.g., 'utf8', 'latin1').
-n, --limit INTEGERMaximum number of records to scan for field detection.10000
--verbose / --no-verboseEnable verbose logging output.--no-verbose
-F, --format-in TEXTOverride input file format detection (e.g., 'csv', 'jsonl', 'xml').
-O, --format-out TEXTOverride output format (e.g., 'csv', 'json').
--zipfile / --no-zipfileTreat input file as a ZIP archive.--no-zipfile
--filter-expr TEXTFilter expression (kept for API compatibility, not currently used).
--table, --sheet TEXTTable or sheet name for multi-table sources (Excel, SQLite, lakehouse).
--start-page INTEGERSheet index (0-based) for Excel files.0
--trustAcknowledge pickle deserialization risk when reading pickle sources.
--on-error TEXTParse-error policy: raise (default), skip, or warn.
--error-log TEXTAppend parse errors as JSONL (use with --on-error skip or warn).
--flatten-nestedUnfold nested dict / array-of-dict fields into dotted paths (e.g. city.lat).
--max-nested-depth INTEGERWith --flatten-nested, maximum nest depth to unfold (engine default 5).
--keep-nested-parents / --no-keep-nested-parentsWith --flatten-nested, keep parent dict/array fields alongside dotted children.--keep-nested-parents
--jsonPrint the result as one JSON document (same as --format-out json).

See also shared options.