Read Existing Data
Use Curator’s JsonlReader, ParquetReader, and LanceReader to read existing datasets into a pipeline, then optionally add processing stages.
JSONL Reader
Parquet Reader
Lance Reader
:sync: jsonl
Example: Read JSONL and Filter
Reader Configuration
Common Parameters
JsonlReader and ParquetReader support these configuration options:
JSONL Engine Selection
JsonlReader parses JSONL directly into a PyArrow table by default. Choose the engine through
read_kwargs when your dataset requires different inference behavior.
The direct engine does not silently fall back to pandas. If PyArrow cannot infer a consistent schema—for example, when one JSON field contains both numbers and strings—select the pandas engine explicitly:
The direct engine accepts these additional options:
The direct reader processes 8 MiB chunks and retries with larger chunks, up to 256 MiB, for rows containing large payloads such as base64-encoded images or PDFs.
Parquet-Specific Features
ParquetReader provides these optimizations:
- PyArrow Engine: Uses
pyarrowengine by default for better performance - Storage Options: Supports cloud storage via
storage_optionsinread_kwargs - Schema Handling: Automatic schema inference and validation
- Columnar Efficiency: Optimized for reading specific columns
Performance Tips
- Use
fieldsparameter to read required columns for better performance - Set
files_per_partitionbased on your cluster size and memory constraints - Use
blocksizefor fine-grained control over partition sizes
Output Integration
These readers produce DocumentBatch tasks that integrate seamlessly with:
- Processing Stages: Apply filters, transformations, and quality checks
- Writer Stages: Export to JSONL, Parquet, or other formats
- Analysis Tools: Convert to Pandas/PyArrow for inspection and debugging