Apache Arrow/Feather Format
Description
Apache Arrow is a columnar in-memory data format designed for efficient data transfer and analytics. Feather is a file format built on Arrow that provides fast, language-agnostic columnar storage. It's optimized for speed and is ideal for data exchange between Python, R, and other languages.
File Extensions
.arrow- Apache Arrow files.feather- Feather format files (alias for Arrow)
Implementation Details
Reading
The Arrow/Feather implementation:
- Uses PyArrow library for reading
- Loads entire table into memory
- Supports batch reading for efficient iteration
- Preserves data types and schema
Writing
Writing support:
- Buffers records before writing
- Uses PyArrow Feather writer
- Flushes data on close
- Maintains schema from data
Key Features
- Columnar format: Efficient for analytical workloads
- Fast I/O: Optimized for speed
- Type preservation: Maintains data types
- Cross-language: Compatible with R, Java, and other languages
- Totals support: Can count total rows
- Batch processing: Efficient batch reading
Usage
from iterable import open_iterable
# Basic reading
with open_iterable('data.arrow') as source:
for row in source:
print(row)
# Writing
dest = open_iterable('output.feather', mode='w', iterableargs={
'batch_size': 10000
})
dest.write({'id': 1, 'name': 'John'})
dest.close() # Important: closes and flushes data
Parameters
batch_size(int): Batch size for reading (default:1024)
Limitations
- PyArrow dependency: Requires
pyarrowpackage - Memory usage: Entire file is loaded into memory when reading
- Flat data only: Designed for tabular/flat data structures
- Write buffering: Must call
close()to flush buffered data - File path requirement: Feather writing may require file path rather than stream
Compression Support
Arrow/Feather files can be compressed with all supported codecs:
- GZip (
.arrow.gz) - BZip2 (
.arrow.bz2) - LZMA (
.arrow.xz) - LZ4 (
.arrow.lz4) - ZIP (
.arrow.zip) - Brotli (
.arrow.br) - ZStandard (
.arrow.zst)
Performance Considerations
- Fast I/O: Feather format is optimized for speed
- Memory: Entire file loaded into memory for reading
- Batch size: Larger batches improve iteration performance
Use Cases
- Data exchange: Fast transfer between Python and R
- Temporary storage: Intermediate format in data pipelines
- Analytics: Fast columnar data access
Error Handling
- Missing dependency: optional libraries raise
ImportErrorwith an install hint (pip install 'iterabledata[<extra>]'when an extra exists). - Write mode: read-only formats raise
WriteNotSupportedErrororValueErrorwhen opened withmode="w". - Bad or unsupported input: may raise
ValueError,OSError, or library-specific errors. - See Troubleshooting for decoding, detection, and engine issues.