Python SDK reference
Generated from the docstrings of undatum.sdk. For a guided introduction see
Python SDK.
from undatum import Dataset
Dataset
Lazy, chainable dataset.
from undatum import Dataset
ds = Dataset.read("data.jsonl")
ds = ds.fill("age", value=0).dedup(keys=["user_id"])
stats = ds.stats()
ds.write("output.parquet")
Dataset.read(path: str, **options) -> Dataset (classmethod)
Read data from a file or cloud URI.
| Parameter | Description |
|---|---|
path | File path or cloud URI (s3://, gs://, az://, ...) |
**options | Reader options (encoding, delimiter, format_in, table, flatten_nested, on_error, ...) |
Returns: Dataset instance
ds = Dataset.read("data.csv")
ds = Dataset.read("s3://bucket/data.jsonl", encoding="utf8")
ds = Dataset.read("workbook.xlsx", table="Sheet2")
ds = Dataset.read("nested.jsonl", flatten_nested=True)
Dataset.from_records(records: Iterable[dict]) -> Dataset (classmethod)
Wrap in-memory records (a list can be iterated any number of times).
Dataset.from_records([{"a": 1}, {"a": 2}]).count()
2
Dataset.source (property)
Path or URI the records are read from (None for in-memory records).
Dataset.steps (property)
The operations of the plan, in order.
Dataset.explain() -> str
Describe the plan, one line per step.
Dataset.collect() -> QueryResult
Run the plan and return every record as a list.
Dataset.write(path: str, **options) -> None
Write dataset to a file or cloud URI.
| Parameter | Description |
|---|---|
path | Output file path or cloud URI (s3://, gs://, az://, ...) |
**options | Output options (format_out; for a plain conversion without steps also compression, level, profile, ...) |
ds.write("output.jsonl")
ds.write("s3://bucket/output.parquet", format_out="parquet")
ds.write("az://container/output.csv")
Dataset.close() -> None
Release resources (kept for compatibility; plans create no temporary files).
Dataset.fill(fields: str | list[str] | None = None, value: Any = None, strategy: str | None = None, **options) -> Dataset
Fill empty or null values.
| Parameter | Description |
|---|---|
fields | Field name(s) to fill (all fields when None) |
value | Constant value (and fallback for forward/backward) |
strategy | constant (default), forward or backward |
**options | Ignored (kept for compatibility) |
Returns: New Dataset with the step added
ds = ds.fill("age", value=0)
ds = ds.fill(["name", "email"], value="N/A")
ds = ds.fill("status", strategy="forward")
Dataset.dedup(keys: list[str] | None = None, keep: str = 'first', **options) -> Dataset
Remove duplicate records.
| Parameter | Description |
|---|---|
keys | Key fields (all fields when None) |
keep | first or last |
**options | low_memory, temp_dir |
Returns: New Dataset with the step added
ds = ds.dedup(keys=["user_id"])
Dataset.sort(by: str | list[str], desc: bool = False, numeric: bool | list[str] = False, **options) -> Dataset
Sort records (stable; external merge sort above 100,000 records).
| Parameter | Description |
|---|---|
by | Field name(s) to sort by |
desc | Descending order |
numeric | True to compare every by field as a number, or a list of fields |
**options | temp_dir for merge runs |
Returns: New Dataset with the step added
ds = ds.sort("age", desc=True, numeric=True)
Dataset.filter(pattern: str | None = None, fields: list[str] | None = None, query: str | None = None, **options) -> Dataset
Keep records matching a regular expression or a comparison expression.
| Parameter | Description |
|---|---|
pattern | Regular expression searched in fields (all fields when None) |
fields | Fields searched by pattern |
query | Comparison expression such as age > 30 AND city == "Berlin" |
**options | ignore_case for pattern |
Returns: New Dataset with the step added
ds = ds.filter(pattern="error", fields=["message"])
ds = ds.filter(query="age > 30")
Dataset.select(fields: str | list[str], filter_expr: str | None = None, **options) -> Dataset
Keep some fields (dotted paths keep nested values) and optionally filter.
| Parameter | Description |
|---|---|
fields | Field name(s) to keep |
filter_expr | Optional comparison expression |
**options | output writes the result there right away (compatibility) |
Returns: New Dataset with the step added
ds = ds.select(["name", "email"], filter_expr="age > 30")
Dataset.join(other: Dataset | str, keys: str | list[str], join_type: str = 'inner', **options) -> Dataset
Hash join with another dataset or file (the other side is held in memory).
| Parameter | Description |
|---|---|
other | Dataset or file path |
keys | Join key field(s) |
join_type | inner, left, right or full |
**options | Reader options for other when it is a path (table2, ...) |
Returns: New Dataset with the step added
ds = ds1.join(ds2, keys=["user_id"], join_type="left")
Dataset.sample(n: int | None = None, percent: float | None = None, **options) -> Dataset
Random sample of n records or percent of them.
| Parameter | Description |
|---|---|
n | Number of records |
percent | Share of the records (0-100) |
**options | seed for a reproducible sample |
Returns: New Dataset with the step added
ds = ds.sample(n=100, seed=1)
Dataset.mask(fields: str | list[str], method: str = 'redact', salt: str | None = None, **options) -> Dataset
Mask sensitive fields.
| Parameter | Description |
|---|---|
fields | Field name(s) to mask |
method | redact, hash or randomize |
salt | Salt for hash |
**options | Ignored (kept for compatibility) |
Returns: New Dataset with the step added
ds = ds.mask(["email", "phone"], method="hash", salt="s3cret")
Dataset.rename(mapping: dict[str, str] | None = None, pattern: str | None = None, replacement: str = '', **options) -> Dataset
Rename fields by exact mapping and/or a regular expression.
| Parameter | Description |
|---|---|
mapping | Dict of old_name to new_name |
pattern | Regular expression matched against field names |
replacement | Replacement for pattern |
**options | Ignored (kept for compatibility) |
Returns: New Dataset with the step added
Dataset.replace(field: str, pattern: str, replacement: str = '', *, regex: bool = False, global_replace: bool = True) -> Dataset
Replace text in one field.
| Parameter | Description |
|---|---|
field | Field whose values change |
pattern | Text, or a regular expression with regex=True |
replacement | Replacement text |
regex | Treat pattern as a regular expression |
global_replace | Replace every occurrence (default) or only the first |
Returns: New Dataset with the step added
Dataset.explode(field: str, separator: str = ',', **options) -> Dataset
One record per part of a delimited field.
| Parameter | Description |
|---|---|
field | Field to split |
separator | Separator (default: comma) |
**options | Ignored (kept for compatibility) |
Returns: New Dataset with the step added
Dataset.enum(field: str = 'row_id', enum_type: str = 'number', start: int = 1, value: Any = None, **options) -> Dataset
Add row numbers, UUIDs or a constant field.
| Parameter | Description |
|---|---|
field | Field name for generated values (default: row_id) |
enum_type | number, uuid or constant |
start | First number |
value | Constant for enum_type="constant" |
**options | Ignored (kept for compatibility) |
Returns: New Dataset with the step added
Dataset.reverse(**options) -> Dataset
Reverse record order (spills to disk for large inputs).
| Parameter | Description |
|---|---|
**options | Ignored (kept for compatibility) |
Returns: New Dataset with the step added