Skip to main content

Installation

Iterable Data is a Python library for reading and writing data files row by row in a consistent, iterator-based interface. It provides a unified API for working with various data formats (CSV, JSON, Parquet, XML, etc.) similar to csv.DictReader but supporting many more formats.

Requirements​

  • Python 3.10 or higher

Install from PyPI​

The PyPI package is iterabledata. The import package is iterable:

pip install iterabledata
from iterable import open_iterable

Install from Source​

To install the latest development version from source:

git clone https://github.com/datenoio/iterabledata.git
cd iterabledata
pip install -e ".[dev]"

Optional Dependencies​

Some formats and engines need extras. Install only what you use:

pip install iterabledata[parquet]      # Parquet, Arrow, GeoParquet
pip install iterabledata[excel] # XLS / XLSX
pip install iterabledata[xml] # XML, ZIPXML, several geospatial XML formats
pip install iterabledata[duckdb] # DuckDB engine and .duckdb files
pip install iterabledata[compression] # zstd, brotli, lz4, snappy, lzo, 7z
pip install iterabledata[geospatial] # GeoJSON, Shapefile, FlatGeobuf, Fiona/GDAL
pip install iterabledata[cloud] # s3://, gs://, az:// via fsspec
pip install iterabledata[db] # SQL/NoSQL engines and ingest
pip install iterabledata[ai] # LLM documentation generation
pip install iterabledata[mcp] # iterable-mcp stdio server

Extras by area​

ExtraWhat it enables
parquet, orc, avro, vortex, npy, ion, bson, cborColumnar / binary analytics formats
excel, xlsb, odsSpreadsheets
xml, html, rdfMarkup and RDF
geospatial, lidar, mvt, topojson, dxfSpatial formats
stats, rdata, mat, hdf5, zarr, netcdf, cdf, fstScientific / stats
alignment, bioSAM/BAM/CRAM, genomic VCF, BED/GFF/GTF extras
geophysicalSEG-Y, GRIB2, miniSEED
lakehouse, ducklake, paimon, paimon-row, paimon-mosaic, paimon-tableLakehouse tables and Paimon files
otlp, protobufOpenTelemetry and generic protobuf
db, db-sql, db-nosql, db-ingestDatabase engines and ingest
dataframes, pydanticpandas/Polars/Dask bridges and typed models
cloudS3, GCS, Azure
warc, dbf, graph, access, ics, ldif, hocon, feed, pcap, htmlArchives, graphs, Access, vCard/iCal, logs
ai, anthropic, google-genai, langchain, mcp, agentsLLM and agent surfaces
compressionOptional codecs (see Codecs)
allEverything except dev
devTests, ruff, mypy, pre-commit

The full pin list is in pyproject.toml [project.optional-dependencies]. Format pages name the extra they need.

Verify Installation​

You can verify the installation by importing the library:

from iterable import open_iterable

# If this runs without errors, installation was successful
print("Iterable Data installed successfully!")

Next Steps​