Installation
Iterable Data is a Python library for reading and writing data files row by row in a consistent, iterator-based interface. It provides a unified API for working with various data formats (CSV, JSON, Parquet, XML, etc.) similar to csv.DictReader but supporting many more formats.
Requirements
- Python 3.10 or higher
Install from PyPI
The PyPI package is iterabledata. The import package is iterable:
pip install iterabledata
from iterable import open_iterable
Install from Source
To install the latest development version from source:
git clone https://github.com/datenoio/iterabledata.git
cd iterabledata
pip install -e ".[dev]"
Optional Dependencies
Some formats and engines need extras. Install only what you use:
pip install iterabledata[parquet] # Parquet, Arrow, GeoParquet
pip install iterabledata[excel] # XLS / XLSX
pip install iterabledata[xml] # XML, ZIPXML, several geospatial XML formats
pip install iterabledata[duckdb] # DuckDB engine and .duckdb files
pip install iterabledata[compression] # zstd, brotli, lz4, snappy, lzo, 7z
pip install iterabledata[geospatial] # GeoJSON, Shapefile, FlatGeobuf, Fiona/GDAL
pip install iterabledata[cloud] # s3://, gs://, az:// via fsspec
pip install iterabledata[db] # SQL/NoSQL engines and ingest
pip install iterabledata[ai] # LLM documentation generation
pip install iterabledata[mcp] # iterable-mcp stdio server
Extras by area
| Extra | What it enables |
|---|---|
parquet, orc, avro, vortex, npy, ion, bson, cbor | Columnar / binary analytics formats |
excel, xlsb, ods | Spreadsheets |
xml, html, rdf | Markup and RDF |
geospatial, lidar, mvt, topojson, dxf | Spatial formats |
stats, rdata, mat, hdf5, zarr, netcdf, cdf, fst | Scientific / stats |
alignment, bio | SAM/BAM/CRAM, genomic VCF, BED/GFF/GTF extras |
geophysical | SEG-Y, GRIB2, miniSEED |
lakehouse, ducklake, paimon, paimon-row, paimon-mosaic, paimon-table | Lakehouse tables and Paimon files |
otlp, protobuf | OpenTelemetry and generic protobuf |
db, db-sql, db-nosql, db-ingest | Database engines and ingest |
dataframes, pydantic | pandas/Polars/Dask bridges and typed models |
cloud | S3, GCS, Azure |
warc, dbf, graph, access, ics, ldif, hocon, feed, pcap, html | Archives, graphs, Access, vCard/iCal, logs |
ai, anthropic, google-genai, langchain, mcp, agents | LLM and agent surfaces |
compression | Optional codecs (see Codecs) |
all | Everything except dev |
dev | Tests, ruff, mypy, pre-commit |
The full pin list is in pyproject.toml [project.optional-dependencies]. Format pages name the extra they need.
Verify Installation
You can verify the installation by importing the library:
from iterable import open_iterable
# If this runs without errors, installation was successful
print("Iterable Data installed successfully!")
Next Steps
- Quick Start Guide - Get up and running quickly
- When to use IterableData - vs pandas and the standard library
- Cookbook - prompt-shaped recipes
- Basic Usage - Learn common patterns
- Supported Formats - See all available formats