extract
Extracts tables or text from PDF/DOC/DOCX/XLS/XLSX files and outputs CSV, JSON, NDJSON, Parquet, or a Frictionless Data Package. PDF extraction supports table, text, or OCR modes.
pip install "undatum[extract]"
# PDF tables to CSV
undatum extract report.pdf --format-out csv --output report.csv
# Extract tables from multiple files
undatum extract data/*.pdf --format-out parquet --output-dir out/
# PDF text extraction for specific pages
undatum extract report.pdf --method text --pages 1-3 --format-out ndjson --output report.ndjson
Optional extra: pip install "undatum[extract]" installs:
pdfplumber— PDF tables/textpdf2image+pytesseract— OCR (--method ocr; system Tesseract still required)textract— legacy.doc
Reference
Reads: PDF, DOCX, DOC, HTML and images (OCR) · Writes: CSV, JSON, JSON Lines, Parquet or a data package · Memory: one document at a time · Engines: python
undatum extract [OPTIONS] INPUT_FILES...
| Argument | Description |
|---|---|
INPUT_FILES... | Input file(s) to extract from. (required) |
| Option | Description | Default |
|---|---|---|
-O, --format-out TEXT | Output format: csv, json, ndjson, parquet, datapackage. | csv |
-o, --output TEXT | Output file path (single table only). | |
--output-dir TEXT | Output directory for multiple tables. | |
--method TEXT | Extraction method: tables, text, ocr. | |
--pages TEXT | PDF pages to extract (e.g., 1-3,7,10-12). | |
--ocr / --no-ocr | Enable OCR for scanned PDFs. | --no-ocr |
--flatten / --no-flatten | Flatten multiple tables into one output table. | --no-flatten |
--verbose / --no-verbose | Enable verbose logging output. | --no-verbose |
Deprecated spellings (removed in 2.0): --output-format → --format-out.
See also shared options.