Text Quickstart
Quickstart for NeMo Curator Text Curation
Use this quickstart to install NeMo Curator, run a text quality-filtering pipeline, and inspect its task count. For full setup options, refer to the Installation Guide.
Prerequisites
Prepare the following environment:
- Python 3.11, 3.12, or 3.13.
- Ubuntu 20.04 or 22.04,
packaging22.0 or later, Git, anduvfor package management. - Use an NVIDIA GPU with CUDA 12 only for GPU-accelerated operations; most text modules run without a GPU.
- For GPU modules, use a Volta-class or newer GPU with compute capability 7.0 or later.
Quickstart Steps
Complete these steps to install Curator, prepare input data, and run a filtering pipeline.
-
Install NeMo Curator from PyPI when the needed release is available; use source for development or an unpublished release. A local container is also available for deployments that require one.
PyPI Installation
Source Installation
Build an Image
For a specific release, verify that its package version is available on PyPI and matches the release notes.
For
text_cuda12, use this tested dependency override. Standardpip installand plainuv pip installwithout the override are not supported. -
Prepare a directory and one JSONL input file for your pipeline.
Add JSONL files to
~/nemo_curator/data/sample/. Each line must be a JSON object with at leasttextandidfields. Refer to Read Existing Data for input options. -
Create and run a text filtering pipeline.
-
Confirm that the pipeline prints
Pipeline completed successfully!and a task count. The number of processed tasks depends on your input files and filters.
If you download datasets or models from Hugging Face, set HF_TOKEN to reduce rate limiting. Create a token at Hugging Face Settings.
Next Steps
Explore these resources to build on the text pipeline: