Skip to main content

Python SDK reference

Generated from the docstrings of undatum.sdk. For a guided introduction see Python SDK.

from undatum import Dataset

Dataset​

Lazy, chainable dataset.

from undatum import Dataset
ds = Dataset.read("data.jsonl")
ds = ds.fill("age", value=0).dedup(keys=["user_id"])
stats = ds.stats()
ds.write("output.parquet")

Dataset.read(path: str, **options) -> Dataset (classmethod)​

Read data from a file or cloud URI.

ParameterDescription
pathFile path or cloud URI (s3://, gs://, az://, ...)
**optionsReader options (encoding, delimiter, format_in, table, flatten_nested, on_error, ...)

Returns: Dataset instance

ds = Dataset.read("data.csv")
ds = Dataset.read("s3://bucket/data.jsonl", encoding="utf8")
ds = Dataset.read("workbook.xlsx", table="Sheet2")
ds = Dataset.read("nested.jsonl", flatten_nested=True)

Dataset.from_records(records: Iterable[dict]) -> Dataset (classmethod)​

Wrap in-memory records (a list can be iterated any number of times).

Dataset.from_records([{"a": 1}, {"a": 2}]).count()
2

Dataset.source (property)​

Path or URI the records are read from (None for in-memory records).

Dataset.steps (property)​

The operations of the plan, in order.

Dataset.explain() -> str​

Describe the plan, one line per step.

Dataset.collect() -> QueryResult​

Run the plan and return every record as a list.

Dataset.write(path: str, **options) -> None​

Write dataset to a file or cloud URI.

ParameterDescription
pathOutput file path or cloud URI (s3://, gs://, az://, ...)
**optionsOutput options (format_out; for a plain conversion without steps also compression, level, profile, ...)
ds.write("output.jsonl")
ds.write("s3://bucket/output.parquet", format_out="parquet")
ds.write("az://container/output.csv")

Dataset.close() -> None​

Release resources (kept for compatibility; plans create no temporary files).

Dataset.fill(fields: str | list[str] | None = None, value: Any = None, strategy: str | None = None, **options) -> Dataset​

Fill empty or null values.

ParameterDescription
fieldsField name(s) to fill (all fields when None)
valueConstant value (and fallback for forward/backward)
strategyconstant (default), forward or backward
**optionsIgnored (kept for compatibility)

Returns: New Dataset with the step added

ds = ds.fill("age", value=0)
ds = ds.fill(["name", "email"], value="N/A")
ds = ds.fill("status", strategy="forward")

Dataset.dedup(keys: list[str] | None = None, keep: str = 'first', **options) -> Dataset​

Remove duplicate records.

ParameterDescription
keysKey fields (all fields when None)
keepfirst or last
**optionslow_memory, temp_dir

Returns: New Dataset with the step added

ds = ds.dedup(keys=["user_id"])

Dataset.sort(by: str | list[str], desc: bool = False, numeric: bool | list[str] = False, **options) -> Dataset​

Sort records (stable; external merge sort above 100,000 records).

ParameterDescription
byField name(s) to sort by
descDescending order
numericTrue to compare every by field as a number, or a list of fields
**optionstemp_dir for merge runs

Returns: New Dataset with the step added

ds = ds.sort("age", desc=True, numeric=True)

Dataset.filter(pattern: str | None = None, fields: list[str] | None = None, query: str | None = None, **options) -> Dataset​

Keep records matching a regular expression or a comparison expression.

ParameterDescription
patternRegular expression searched in fields (all fields when None)
fieldsFields searched by pattern
queryComparison expression such as age > 30 AND city == "Berlin"
**optionsignore_case for pattern

Returns: New Dataset with the step added

ds = ds.filter(pattern="error", fields=["message"])
ds = ds.filter(query="age > 30")

Dataset.select(fields: str | list[str], filter_expr: str | None = None, **options) -> Dataset​

Keep some fields (dotted paths keep nested values) and optionally filter.

ParameterDescription
fieldsField name(s) to keep
filter_exprOptional comparison expression
**optionsoutput writes the result there right away (compatibility)

Returns: New Dataset with the step added

ds = ds.select(["name", "email"], filter_expr="age > 30")

Dataset.join(other: Dataset | str, keys: str | list[str], join_type: str = 'inner', **options) -> Dataset​

Hash join with another dataset or file (the other side is held in memory).

ParameterDescription
otherDataset or file path
keysJoin key field(s)
join_typeinner, left, right or full
**optionsReader options for other when it is a path (table2, ...)

Returns: New Dataset with the step added

ds = ds1.join(ds2, keys=["user_id"], join_type="left")

Dataset.sample(n: int | None = None, percent: float | None = None, **options) -> Dataset​

Random sample of n records or percent of them.

ParameterDescription
nNumber of records
percentShare of the records (0-100)
**optionsseed for a reproducible sample

Returns: New Dataset with the step added

ds = ds.sample(n=100, seed=1)

Dataset.mask(fields: str | list[str], method: str = 'redact', salt: str | None = None, **options) -> Dataset​

Mask sensitive fields.

ParameterDescription
fieldsField name(s) to mask
methodredact, hash or randomize
saltSalt for hash
**optionsIgnored (kept for compatibility)

Returns: New Dataset with the step added

ds = ds.mask(["email", "phone"], method="hash", salt="s3cret")

Dataset.rename(mapping: dict[str, str] | None = None, pattern: str | None = None, replacement: str = '', **options) -> Dataset​

Rename fields by exact mapping and/or a regular expression.

ParameterDescription
mappingDict of old_name to new_name
patternRegular expression matched against field names
replacementReplacement for pattern
**optionsIgnored (kept for compatibility)

Returns: New Dataset with the step added

Dataset.replace(field: str, pattern: str, replacement: str = '', *, regex: bool = False, global_replace: bool = True) -> Dataset​

Replace text in one field.

ParameterDescription
fieldField whose values change
patternText, or a regular expression with regex=True
replacementReplacement text
regexTreat pattern as a regular expression
global_replaceReplace every occurrence (default) or only the first

Returns: New Dataset with the step added

Dataset.explode(field: str, separator: str = ',', **options) -> Dataset​

One record per part of a delimited field.

ParameterDescription
fieldField to split
separatorSeparator (default: comma)
**optionsIgnored (kept for compatibility)

Returns: New Dataset with the step added

Dataset.enum(field: str = 'row_id', enum_type: str = 'number', start: int = 1, value: Any = None, **options) -> Dataset​

Add row numbers, UUIDs or a constant field.

ParameterDescription
fieldField name for generated values (default: row_id)
enum_typenumber, uuid or constant
startFirst number
valueConstant for enum_type="constant"
**optionsIgnored (kept for compatibility)

Returns: New Dataset with the step added

Dataset.reverse(**options) -> Dataset​

Reverse record order (spills to disk for large inputs).

ParameterDescription
**optionsIgnored (kept for compatibility)

Returns: New Dataset with the step added

Dataset.exclude(other: Dataset | str, on: str | list[str] | None = None) -> Dataset​

Drop records whose key appears in other.

ParameterDescription
otherDataset or file path with the keys to exclude
onKey fields (all fields when None)

Returns: New Dataset with the step added

Dataset.concat(*others) -> Dataset​

Append the records of other datasets or files.

Returns: New Dataset with the step added

Dataset.limit(n: int) -> Dataset​

Keep the first n records (lazy; :meth:head returns a list).

Returns: New Dataset with the step added

Dataset.slice(start: int = 0, end: int | None = None) -> Dataset​

Records start (inclusive) to end (exclusive), 0-based.

Returns: New Dataset with the step added

Dataset.fixlengths(strategy: str = 'pad', value: Any = '') -> Dataset​

Give every record the same, alphabetically ordered fields.

ParameterDescription
strategypad or truncate
valueValue for missing fields

Returns: New Dataset with the step added

Dataset.transpose() -> Dataset​

One record per field: field plus row_0, row_1, ...

Returns: New Dataset with the step added

Dataset.count(**options) -> int​

Number of records.

ParameterDescription
**optionsReader options applied for this call (e.g. flatten_nested)

Returns: Number of records

Dataset.read("data.csv").count()
1000

Dataset.head(n: int = 10, **options) -> list[dict]​

First n records.

ParameterDescription
nNumber of records
**optionsReader options applied for this call

Returns: List of records

Dataset.tail(n: int = 10, **options) -> list[dict]​

Last n records.

ParameterDescription
nNumber of records
**optionsReader options applied for this call

Returns: List of records

Dataset.uniq(fields: str | list[str]) -> list[dict]​

Distinct combinations of fields, in first-seen order.

Returns: List of dicts with the field values

Dataset.frequency(fields: str | list[str]) -> list[dict]​

How often each combination of fields occurs, most frequent first.

Returns: List of dicts with the field values and count

Dataset.validate(rules: str | list[dict[str, Any]]) -> dict[str, Any]​

Check every record against validation rules.

ParameterDescription
rulesRules file (YAML/JSON, see undatum validate) or a list of rule dicts

Returns: {"records": n, "valid": n_valid, "violations": [...]}

Dataset.schema(limit: int = 1000) -> dict[str, Any]​

Infer a schema (field types, nested structure) from the first limit records.

Returns: Schema dict as produced by undatum schema

Dataset.quality(rules: str | None = None, schema: str | None = None, thresholds: str | None = None) -> dict[str, Any]​

Data quality report: field profiles, schema differences, rule violations, verdict.

ParameterDescription
rulesValidation rule file (YAML/JSON).
schemaExpected schema (undatum schema --json output, JSON Schema, Frictionless schema, or a data file).
thresholdsThresholds file; report["verdict"]["passed"] is False when one fails.

Returns: The undatum.quality/1 report (without the schema id key).

Dataset.read("data.csv").quality()["summary"]["rows"]
1000

Dataset.stats(**options) -> dict[str, Any]​

Statistics profile (types, uniqueness, lengths, distributions).

ParameterDescription
**optionscheckdates, engine, flatten_nested, ...

Returns: Profile with count, num_fields, fieldtypes, fields, dictkeys and per-field details under debug

Dataset.read("data.csv").stats()["count"]
1000

Dataset.package(output: str | None = None, package_dir: str | None = None, **options) -> dict[str, Any]​

Generate a Frictionless Data Package descriptor for this dataset.

ParameterDescription
outputOutput datapackage.json path
package_dirOptional directory to materialize the package
**optionsPackaging options (autodoc, metadata, ...)

Returns: Dictionary with package, output_file and optional archive_path

Dataset.convert_many(source: str, dest: str, *, to_ext: str | None = None, filename_pattern: str | None = None, **options) -> Any (classmethod)​

Bulk-convert a directory or glob of files.

Same as undatum convert --recursive; to_ext is required unless format_out is set.

ParameterDescription
sourceDirectory path or glob pattern
destOutput directory
to_extTarget extension (e.g. "jsonl")
filename_patternOutput name pattern with {name}, {stem} and {ext}
**optionsConversion options (encoding, quotechar, profile, level, ...)

Returns: iterabledata BulkConversionResult

Dataset.convert_many("./raw", "./out", to_ext="jsonl")

Dataset.to_pandas(chunksize: int | None = None) -> Any​

Convert to a pandas DataFrame (or an iterator of them with chunksize).

Dataset.to_polars(chunksize: int | None = None) -> Any​

Convert to a Polars DataFrame (needs undatum[polars]).

Dataset.to_dask(chunksize: int = 1000000) -> Any​

Convert to a Dask DataFrame (needs undatum[dask]).

Dataset.as_dataclasses(dataclass_type: type, skip_empty: bool = True) -> Iterator[Any]​

Iterate records as instances of a dataclass.

ParameterDescription
dataclass_typeThe dataclass to build
skip_emptySkip empty records

Yields: Instances of dataclass_type

Dataset.as_pydantic(model_type: type, skip_empty: bool = True, validate: bool = True) -> Iterator[Any]​

Iterate records as pydantic models (needs pydantic).

ParameterDescription
model_typeThe model class
skip_emptySkip empty records
validateValidate fields

Yields: Instances of model_type

StatsResult​

Statistics profile with both mapping and attribute access.

Returned by Dataset.stats(). Existing callers that treat the result as a dict keep working; attribute access is provided for the common fields.

StatsResult.count (property)​

Number of records profiled.

StatsResult.num_fields (property)​

Number of fields found.

StatsResult.fields (property)​

Per-field statistics.

StatsResult.fieldtypes (property)​

Detected type distribution per field.

StatsResult.dictkeys (property)​

Fields with few distinct values (dictionary candidates).

StatsResult.debug (property)​

Engine diagnostics (timings, engine used).

QueryResult​

List of records returned by Dataset query-style methods (head/tail).

QueryResult.to_dicts() -> list[dict[str, Any]]​

Return the records as plain dicts (non-mapping rows become {"value": row}).