Get Started

Text Quickstart

View as Markdown

Quickstart for NeMo Curator Text Curation

Use this quickstart to install NeMo Curator, run a text quality-filtering pipeline, and inspect its task count. For full setup options, refer to the Installation Guide.

Prerequisites

Prepare the following environment:

  • Python 3.11, 3.12, or 3.13.
  • Ubuntu 20.04 or 22.04, packaging 22.0 or later, Git, and uv for package management.
  • Use an NVIDIA GPU with CUDA 12 only for GPU-accelerated operations; most text modules run without a GPU.
  • For GPU modules, use a Volta-class or newer GPU with compute capability 7.0 or later.

Quickstart Steps

Complete these steps to install Curator, prepare input data, and run a filtering pipeline.

  1. Install NeMo Curator from PyPI when the needed release is available; use source for development or an unpublished release. A local container is also available for deployments that require one.

    For a specific release, verify that its package version is available on PyPI and matches the release notes.

    curl -LsSf https://astral.sh/uv/0.12.1/install.sh | sh
    source $HOME/.local/bin/env
    uv venv
    source .venv/bin/activate
    curl -O https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/requirements/text_cuda12-overrides.txt
    uv pip install \
    --override text_cuda12-overrides.txt \
    --torch-backend cu129 \
    --extra-index-url https://wheels.vllm.ai/0.22.0/cu129 \
    "nemo-curator[text_cuda12]"

    For text_cuda12, use this tested dependency override. Standard pip install and plain uv pip install without the override are not supported.

  2. Prepare a directory and one JSONL input file for your pipeline.

    mkdir -p ~/nemo_curator/data/sample
    mkdir -p ~/nemo_curator/data/curated
    cat > ~/nemo_curator/data/sample/example.jsonl <<'EOF'
    {"id":"example-1","text":"NeMo Curator provides a repeatable pipeline for preparing training data. This sample contains enough words to pass the minimum word count filter. Teams can load, filter, deduplicate, and transform text datasets with reusable processing stages. The same pipeline can run locally or across a distributed cluster. Teams can reuse its stages across dataset preparation workflows."}
    EOF

    Add JSONL files to ~/nemo_curator/data/sample/. Each line must be a JSON object with at least text and id fields. Refer to Read Existing Data for input options.

  3. Create and run a text filtering pipeline.

    from nemo_curator.pipeline import Pipeline
    from nemo_curator.stages.text.io.reader import JsonlReader
    from nemo_curator.stages.text.io.writer import JsonlWriter
    from nemo_curator.stages.text.filters import ScoreFilter
    from nemo_curator.stages.text.filters.heuristic import WordCountFilter, NonAlphaNumericFilter
    pipeline = Pipeline(
    name="text_curation_pipeline",
    description="Basic text quality filtering pipeline"
    )
    pipeline.add_stage(
    JsonlReader(
    file_paths="~/nemo_curator/data/sample/",
    files_per_partition=4,
    fields=["text", "id"]
    )
    )
    pipeline.add_stage(
    ScoreFilter(
    filter_obj=WordCountFilter(min_words=50, max_words=100000),
    text_field="text",
    score_field="word_count"
    )
    )
    pipeline.add_stage(
    ScoreFilter(
    filter_obj=NonAlphaNumericFilter(max_non_alpha_numeric_to_text_ratio=0.25),
    text_field="text",
    score_field="non_alpha_score"
    )
    )
    pipeline.add_stage(JsonlWriter("~/nemo_curator/data/curated"))
    results = pipeline.run()
    print(f"Pipeline completed successfully! Processed {len(results) if results else 0} tasks.")
  4. Confirm that the pipeline prints Pipeline completed successfully! and a task count. The number of processed tasks depends on your input files and filters.

If you download datasets or models from Hugging Face, set HF_TOKEN to reduce rate limiting. Create a token at Hugging Face Settings.

export HF_TOKEN="your_token_here"

Next Steps

Explore these resources to build on the text pipeline: