Curate TextLoad Data

Read Existing Data

View as Markdown

Use Curator’s JsonlReader, ParquetReader, and LanceReader to read existing datasets into a pipeline, then optionally add processing stages.

:sync: jsonl

Example: Read JSONL and Filter

from nemo_curator.core.client import RayClient
from nemo_curator.pipeline import Pipeline
from nemo_curator.stages.text.io.reader import JsonlReader
from nemo_curator.stages.text.filters import ScoreFilter
from nemo_curator.stages.text.filters.heuristic import WordCountFilter
# Initialize Ray client
ray_client = RayClient()
ray_client.start()
# Create pipeline for processing existing JSONL files
pipeline = Pipeline(name="jsonl_data_processing")
# Read JSONL files
reader = JsonlReader(
file_paths="/path/to/data",
files_per_partition=4,
fields=["text", "url"] # Only read specific columns
)
pipeline.add_stage(reader)
# Add filtering stage
word_filter = ScoreFilter(
filter_obj=WordCountFilter(min_words=50, max_words=1000),
text_field="text"
)
pipeline.add_stage(word_filter)
# Add more stages to pipeline...
# Execute pipeline
results = pipeline.run()
# Stop Ray client
ray_client.stop()

Reader Configuration

Common Parameters

JsonlReader and ParquetReader support these configuration options:

ParameterTypeDescriptionDefault
file_pathsstr | list[str]File paths or glob patterns to readRequired
files_per_partitionint | NoneNumber of files per partition. Overrides blocksize if both are provided.None
blocksizeint | str | NoneTarget partition size (e.g., “128MB”). Ignored if files_per_partition is provided.None
fieldslist[str] | NoneColumn names to read (column selection)None (all columns)
read_kwargsdict[str, Any] | NoneExtra arguments for the underlying readerNone

JSONL Engine Selection

JsonlReader parses JSONL directly into a PyArrow table by default. Choose the engine through read_kwargs when your dataset requires different inference behavior.

read_kwargs["engine"]Output backing typeUse when
Omitted or "pyarrow_direct"pyarrow.TableRecommended for throughput and Arrow-native processing
"pandas"pandas.DataFrameYou need pandas-specific options or inference, such as convert_dates or mixed-type object columns
Another pandas engine, such as "ujson" or "pyarrow"pandas.DataFrameYou explicitly need an engine supported by pandas.read_json

The direct engine does not silently fall back to pandas. If PyArrow cannot infer a consistent schema—for example, when one JSON field contains both numbers and strings—select the pandas engine explicitly:

reader = JsonlReader(
file_paths="/path/to/data",
read_kwargs={
"engine": "pandas",
"convert_dates": ["created_at"],
},
)

The direct engine accepts these additional options:

OptionDescriptionDefault
pyarrow_block_sizeInitial number of bytes PyArrow processes per parser block8 MiB
pyarrow_max_block_sizeMaximum parser block used when retrying an unusually large JSON object256 MiB

The direct reader processes 8 MiB chunks and retries with larger chunks, up to 256 MiB, for rows containing large payloads such as base64-encoded images or PDFs.

Parquet-Specific Features

ParquetReader provides these optimizations:

  • PyArrow Engine: Uses pyarrow engine by default for better performance
  • Storage Options: Supports cloud storage via storage_options in read_kwargs
  • Schema Handling: Automatic schema inference and validation
  • Columnar Efficiency: Optimized for reading specific columns

Performance Tips

  • Use fields parameter to read required columns for better performance
  • Set files_per_partition based on your cluster size and memory constraints
  • Use blocksize for fine-grained control over partition sizes

Output Integration

These readers produce DocumentBatch tasks that integrate seamlessly with:

  • Processing Stages: Apply filters, transformations, and quality checks
  • Writer Stages: Export to JSONL, Parquet, or other formats
  • Analysis Tools: Convert to Pandas/PyArrow for inspection and debugging