ORC Format
Description
ORC (Optimized Row Columnar) is a columnar storage format designed for Hadoop workloads. It provides efficient compression and fast columnar access, making it ideal for analytical queries. ORC files store data in a column-oriented format with built-in compression and indexing.
File Extensions
.orc- ORC files
Implementation Details
Reading
The ORC implementation:
- Uses
pyorclibrary for reading - Supports schema inference from file
- Converts ORC records to Python dictionaries
- Efficient columnar reading
Writing
Writing support:
- Requires schema or keys specification
- Supports compression (default: level 5)
- Writes ORC files with schema
- Efficient columnar writing
Key Features
- Columnar storage: Efficient for analytical queries
- Compression: Built-in compression support
- Schema support: Requires schema definition for writing
- Totals support: Can count total rows
- Type preservation: Maintains data types
Usage
from iterable import open_iterable
# Basic reading
with open_iterable('data.orc') as source:
for row in source:
print(row)
# Writing with schema
with open_iterable('output.orc', mode='w', iterableargs={
'keys': ['id', 'name', 'age'],
'compression': 5 # Compression level (0-9)
}) as dest:
dest.write({'id': '1', 'name': 'John', 'age': '30'})
Parameters
keys(list[str]): Column names (required for writing if schema not provided)schema(list[str]): ORC schema definition (optional, overrides keys)compression(int): Compression level 0-9 (default:5)
Limitations
- pyorc dependency: Requires
pyorcpackage - Schema required: Must specify schema or keys for writing
- Flat data only: Designed for tabular/flat data structures
- Binary format: Not human-readable
- Hadoop ecosystem: Primarily designed for Hadoop environments
Compression Support
ORC has built-in compression (separate from file-level compression):
- Compression levels 0-9
- Default compression level: 5
ORC files can also be compressed with file-level codecs:
- GZip (
.orc.gz) - BZip2 (
.orc.bz2) - LZMA (
.orc.xz) - LZ4 (
.orc.lz4) - ZIP (
.orc.zip) - Brotli (
.orc.br) - ZStandard (
.orc.zst)
Use Cases
- Hadoop ecosystems: Data storage in Hadoop
- Data warehousing: Analytical data storage
- ETL pipelines: Intermediate format for data transformation
- Big data: Large-scale data processing
Error Handling
- Missing dependency: optional libraries raise
ImportErrorwith an install hint (pip install 'iterabledata[<extra>]'when an extra exists). - Write mode: read-only formats raise
WriteNotSupportedErrororValueErrorwhen opened withmode="w". - Bad or unsupported input: may raise
ValueError,OSError, or library-specific errors. - See Troubleshooting for decoding, detection, and engine issues.
Related Formats
- Parquet - Similar columnar format
- Arrow - Another columnar format
- Delta Lake - Transactional layer over Parquet