Curate TextLoad Data

Nemotron-Parse PDF Pipeline

View as Markdown

Convert PDF datasets into interleaved Parquet output using NVIDIA’s Nemotron-Parse vision-language model. Unlike traditional text-only PDF parsers, Nemotron-Parse extracts text, images, and reading order in one pass — producing rows directly compatible with the interleaved dataset format.

How it Works

NemotronParsePDFReader is a composite stage that expands into four underlying sub-stages:

  1. PDFPartitioningStage — reads a JSONL manifest of PDF entries and packs them into FileGroupTask objects.
  2. PDFPreprocessStage — extracts PDF bytes from the configured source, renders pages to images with scale-to-fit safeguarding against OOM on large pages.
  3. NemotronParseHTTPClientStage (recommended) or NemotronParseInferenceStage — calls an OpenAI-compatible InferenceServer, or runs vLLM/Hugging Face Transformers in process.
  4. NemotronParsePostprocessStage — parses model output, aligns images and captions, crops images, and emits the final interleaved rows.

The output is interleaved Parquet ready to be filtered with Interleaved Filters and written to MINT-1T-style WebDataset shards.

Before You Start

Choose your PDF source and confirm the prerequisites:

  • GPU: Required. Use one inference-server model replica per GPU.
  • Inference server: Recommended for production. It lets the serving layer batch page requests across pipeline tasks while rendering and postprocessing scale independently. Internal PDF-pipeline comparisons found this to be the better default; benchmark your corpus before further tuning.
  • In-process inference: Retained for local validation and debugging. Use vLLM when possible, or set backend="hf" when vLLM is unavailable.
  • pypdfium2: Required Python dependency for PDF rendering. Installed automatically with the interleaved_cpu or interleaved_cuda12 extras (e.g., uv sync --extra interleaved_cuda12).
  • OpenCV: Required for PDF rendering plus the postprocessing color conversion and resizing paths. Install the optional extra with uv pip install "nemo-curator[cv2]" (or add --extra cv2 to a source checkout).
  • Manifest: A JSONL file listing the PDFs to process. Each line should specify the PDF location relative to the source directory you choose.

Choosing a PDF Source

Pass exactly one of pdf_dir, zip_base_dir, or jsonl_base_dir so the preprocess stage knows where to find the PDF bytes:

ParameterSource LayoutWhen to Use
pdf_dirA directory of .pdf filesLocal or mounted directories of standalone PDFs
zip_base_dirA CC-MAIN-2021-31-PDF-UNTRUNCATED zip hierarchyCommon Crawl PDF dumps
jsonl_base_dirJSONL-encoded PDF datasets where each line carries the PDF bytesGitHub-hosted PDF datasets, custom JSONL collections

Inference Topology

TopologyWhen to Use
Dynamo InferenceServer + NemotronParseHTTPClientStage (recommended)Production pipelines. Dynamo owns GPU replicas and batches requests across pipeline tasks.
In-process vllmSmall local runs or debugging without a separate server lifecycle.
In-process hfCompatibility fallback when vLLM is unavailable.

For the recommended topology, use a fixed HTTP stage pool of 4 * num_gpus workers and start with inference_batch_size=32 concurrent page requests per worker. We validated both 32 and 64 on 8 H100 GPUs and retained 32 because 64 improved aggregate page throughput only slightly while increasing individual request latency. Because the best value depends on the GPU and corpus, benchmark 64 on the target workload and keep it only when it improves throughput without request failures or out-of-memory errors. The in-process stage retries transient port collisions when starting vLLM; non-retryable startup failures fail immediately.


Usage

The tutorial entry point starts Dynamo, waits for its OpenAI-compatible endpoint, runs the PDF pipeline, and stops the server. From a source checkout, install both dependency groups:

uv sync --extra interleaved_cuda12 --extra inference_server

Dynamo starts local etcd and nats-server processes. Images built from the current repository Dockerfile include both binaries. For a source environment or another image, run the Dynamo dependency installation script before you run the entry point.

Create a manifest with one PDF filename per line:

{"file_name": "document.pdf"}

Start Dynamo and run the pipeline with one command:

uv run python tutorials/interleaved/nemotron_parse_pdf/main.py \
--manifest ./pdfs.jsonl \
--pdf-dir /data/pdfs \
--output-dir ./parsed_pdfs \
--model-path nvidia/NVIDIA-Nemotron-Parse-v1.2 \
--inference-batch-size 32

The entry point detects the Ray-visible GPUs and configures one DynamoVLLMModelConfig replica per GPU. It passes the resulting server.endpoint to create_nemotron_parse_pdf_pipeline, fixes the HTTP stage pool at 4 * num_gpus workers, and uses Ray Data for the pipeline. Use CUDA_VISIBLE_DEVICES or your Ray cluster resources to select the GPUs. The entry point defaults to 32 concurrent requests per HTTP worker; pass --inference-batch-size 64 to compare the higher concurrency on your corpus.

The server configuration uses Dynamo’s TCP request plane, enables multimodal vLLM, and sets the frontend and model-worker options needed by Nemotron-Parse. The shared create_nemotron_parse_inference_server helper owns those model-specific settings so tutorials and benchmarks use the same configuration. It does not use dyn_chat_processor. See the runnable main.py source for the complete lifecycle and Inference Server for the underlying configuration objects. For executor options, refer to Execution Backends.

main.py always starts Dynamo with vLLM and does not expose a backend option. For local validation without an inference server, run inprocess.py with --backend vllm or --backend hf.

The in-process and Ray Serve paths automatically select Triton attention on Ampere GPUs, following the Nemotron-Parse model’s A100/A10 guidance. They also select Triton on Blackwell with vLLM versions before 0.23, avoiding the affected FlashInfer/TRTLLM path. Ray Serve bases this choice on the driver-visible GPU and assumes its architecture matches the serving replicas. Explicit settings always take precedence.

For programmatic use, the helper returns an unstarted InferenceServer; use it as a context manager to start it, wait for health, and stop it:

from nemo_curator.backends.utils import get_available_cpu_gpu_resources
from nemo_curator.stages.interleaved.pdf.nemotron_parse import (
NemotronParsePDFReader,
create_nemotron_parse_inference_server,
)
_, available_gpus = get_available_cpu_gpu_resources(init_and_shutdown=True)
num_gpus = int(available_gpus)
with create_nemotron_parse_inference_server(
model_path="nvidia/NVIDIA-Nemotron-Parse-v1.2",
num_replicas=num_gpus,
) as server:
reader = NemotronParsePDFReader(
manifest_path="./pdfs.jsonl",
pdf_dir="/data/pdfs",
inference_server_endpoint=server.endpoint,
inference_server_client_num_workers=4 * num_gpus,
inference_batch_size=32,
)

Example: CC-MAIN PDF Dump

Parse a Common Crawl PDF dump from its zip hierarchy:

NemotronParsePDFReader(
manifest_path="./cc_pdfs.jsonl",
zip_base_dir="/data/CC-MAIN-2021-31-PDF-UNTRUNCATED",
file_names_field="cc_pdf_file_names",
pdfs_per_task=20,
inference_server_endpoint=server.endpoint,
inference_server_model_name=model_name,
inference_server_client_num_workers=4 * num_gpus,
inference_batch_size=32,
)

Example: JSONL-Encoded PDFs

Parse a JSONL-encoded dataset (e.g., GitHub-hosted PDFs where each line contains the bytes):

NemotronParsePDFReader(
manifest_path="./github_pdfs.jsonl",
jsonl_base_dir="/data/github_pdfs",
inference_server_endpoint=server.endpoint,
inference_server_model_name=model_name,
inference_server_client_num_workers=4 * num_gpus,
inference_batch_size=32,
)

Parameters

ParameterTypeDefaultDescription
manifest_pathstr | NoneNoneJSONL manifest listing PDF entries.
pdf_dirstr | NoneNoneDirectory containing .pdf files.
zip_base_dirstr | NoneNoneRoot directory of CC-MAIN PDF zip hierarchy.
jsonl_base_dirstr | NoneNoneRoot directory of JSONL-encoded PDF datasets.
model_pathstr"nvidia/NVIDIA-Nemotron-Parse-v1.2"Local path or HF repo ID for the Nemotron-Parse weights.
backendstr"vllm"In-process inference backend (vllm or hf); ignored by the HTTP stage.
pdfs_per_taskint10Number of PDFs grouped into each FileGroupTask.
max_pdfsint | NoneNoneHard cap on total PDFs processed (debug aid).
dpiint300Render DPI for PDF pages.
max_pagesint50Maximum pages rendered per PDF; longer PDFs are truncated.
inference_batch_sizeint4Pages per HF pass or concurrent page requests per HTTP worker. The HTTP stage warns below the validated starting point of 32.
max_num_seqsint64Maximum concurrent vLLM sequences.
max_tokensint8192Maximum output tokens generated per page.
text_in_picboolFalseWhen True, treat embedded text within rendered images as part of the text content.
enforce_eagerboolFalseDisable vLLM compilation for compatibility with restricted environments.
min_crop_pxint10Minimum dimension (pixels) for cropped image regions.
dataset_namestr"pdf_dataset"Logical dataset label written to output rows.
file_name_fieldstr"file_name"Manifest field naming a single PDF file.
file_names_fieldstr"cc_pdf_file_names"Manifest field naming a list of PDF files (CC-MAIN layout).
url_fieldstr"url"Manifest field for the source URL passthrough.
inference_server_endpointstr | NoneNoneOpenAI-compatible endpoint. Set this to use the recommended HTTP stage instead of in-process inference.
inference_server_model_namestr | NoneNoneServed model name; defaults to model_path.
inference_server_client_num_workersint4Fixed HTTP stage workers. Set to 4 * num_gpus for production.
inference_server_request_timeout_sfloat300.0Per-page HTTP request timeout.
inference_server_max_retriesint3Retries for transient HTTP failures.

Failure Behavior

PDF read and render failures happen in PDFPreprocessStage, before inference, so they behave identically for in-process and inference-server runs: the bad PDF is skipped, and a task with no renderable PDFs returns no output. After vLLM or HTTP retries—or the Hugging Face single-page fallback—are exhausted, inference failures raise the task instead of emitting a partially parsed document. The executor’s failure policy determines whether the overall run stops or records that task as failed.

Tune In-Process vLLM

Use this section only for the in-process fallback. For the recommended inference-server topology, configure vLLM through DynamoVLLMModelConfig as shown in the Inference Server guide.

NemotronParseInferenceStage.engine_kwargs is None by default. It passes additional settings to NeMo Curator’s shared vLLM initializer: vLLM engine settings are forwarded to vllm.LLM, while helper settings such as max_port_retries control initialization itself. Use this field when you compose the pipeline from individual stages and need controls that are not exposed by NemotronParsePDFReader:

from nemo_curator.pipeline import Pipeline
from nemo_curator.stages.interleaved.io import InterleavedParquetWriterStage
from nemo_curator.stages.interleaved.pdf.nemotron_parse import (
NemotronParseInferenceStage,
NemotronParsePostprocessStage,
PDFPartitioningStage,
PDFPreprocessStage,
)
pipeline = Pipeline(name="tuned_nemotron_parse")
pipeline.add_stage(PDFPartitioningStage(manifest_path="./pdfs.jsonl", pdfs_per_task=10))
pipeline.add_stage(PDFPreprocessStage(pdf_dir="/data/pdfs", max_pages=50))
pipeline.add_stage(
NemotronParseInferenceStage(
backend="vllm",
max_num_seqs=64,
engine_kwargs={
"gpu_memory_utilization": 0.90,
"max_num_batched_tokens": 16384,
},
)
)
pipeline.add_stage(NemotronParsePostprocessStage(min_crop_px=10))
pipeline.add_stage(
InterleavedParquetWriterStage(
path="./parsed_pdfs",
materialize_on_write=False,
)
)

The installed vLLM version performs final validation of forwarded engine settings. Check the matching vLLM documentation before using additional keys.

Keys in engine_kwargs take precedence over the stage’s max_num_seqs and enforce_eager values. Prefer the dedicated stage fields for those two settings and reserve engine_kwargs for other vLLM controls so the effective configuration remains clear.

Common tuning controls include:

KeyEffectTuning guidance
gpu_memory_utilizationFraction of GPU memory reserved by the vLLM engine.Lower it when the engine competes with other GPU workloads; increase cautiously when KV-cache capacity is limiting throughput.
max_num_batched_tokensMaximum tokens scheduled in one iteration.Increase for throughput when memory permits; reduce after scheduler or memory pressure.
dtypeModel weight data type. The helper default is "bfloat16".Change only when the model and GPU support the selected type.
limit_mm_per_promptPer-prompt multimodal limits. The helper default is {"image": 1}.Keep one image per prompt for the current page-level pipeline.
max_port_retriesvLLM engine startup attempts. The helper default is 3.Increase only for nodes with frequent transient MASTER_PORT collisions.

Ray Data Fanout

PDFPartitioningStage is a Ray Data fanout stage. It reads the manifest on one worker, emits one FileGroupTask per pdfs_per_task group, and Ray Data repartitions the result to one emitted task per block. Downstream preprocess and inference stages can then consume those blocks in parallel instead of receiving the whole manifest as one block.

This behavior is automatic with RayDataExecutor; no stage-spec override is required. pdfs_per_task still controls the work in each emitted task:

  • Lower values create more tasks and expose more downstream parallelism, with more scheduling overhead.
  • Higher values reduce scheduling overhead but can leave GPUs idle when the number of tasks is smaller than the available workers.
  • max_pdfs is applied before tasks are created, so it remains useful for small validation runs.

Xenna uses its own task dispatch and does not consume the Ray Data fanout marker.

Output Format

Each output row represents a single item (text, image, or metadata) from a parsed PDF page. Rows sharing a sample_id belong to the same document. Example output JSON:

{
"sample_id": "doc_42",
"position": 0,
"modality": "text",
"text_content": "# Introduction\n\nThis paper investigates...",
"binary_content": null,
"source_files": ["pdf_42.pdf"],
"url": "https://example.com/pdf_42.pdf"
}
{
"sample_id": "doc_42",
"position": 1,
"modality": "image",
"text_content": null,
"binary_content": "<bytes>",
"source_files": ["pdf_42.pdf"]
}
{
"sample_id": "doc_42",
"position": 2,
"modality": "text",
"text_content": "Figure 1 shows the architecture...",
"binary_content": null,
"source_files": ["pdf_42.pdf"]
}

Output Schema

ColumnTypeDescription
sample_idstringPDF identifier; rows sharing a sample_id belong to the same document.
positionintZero-based item position within the sample, used to reconstruct ordering.
modalitystringOne of text, image, or metadata.
text_contentstring | nullText payload for text and metadata rows.
binary_contentbytes | nullImage payload for image rows.
source_fileslist[string]Source PDF files that produced this row (for lineage tracking).

The output is directly compatible with Interleaved IO readers and writers — the schema matches INTERLEAVED_SCHEMA exactly.

Inspect Inference Metrics

Both Nemotron-Parse inference stages record additive custom metrics on each output task. Aggregate the final pipeline results with TaskPerfUtils:

import time
from nemo_curator.tasks.utils import TaskPerfUtils
started = time.perf_counter()
results = pipeline.run(executor)
wall_time_s = time.perf_counter() - started
metrics = TaskPerfUtils.aggregate_task_metrics(results, prefix="task")
metric_prefix = "task_nemotron_parse_inference_custom"
valid_pages = metrics.get(f"{metric_prefix}.num_valid_pages_sum", 0.0)
output_tokens = metrics.get(f"{metric_prefix}.total_output_tokens_sum", 0.0)
pages_per_second = valid_pages / wall_time_s if wall_time_s else 0.0
output_tokens_per_second = output_tokens / wall_time_s if wall_time_s else 0.0
print(f"{pages_per_second:.2f} pages/s")
print(f"{output_tokens_per_second:.2f} output tokens/s")

aggregate_task_metrics() appends _sum, _mean, and _std to each flattened metric. Use _sum for additive counts and total durations. Measure end-to-end throughput against pipeline wall time rather than the sum of per-task inference times, because tasks can run concurrently.

Metric Reference

Custom metricBackendsDescription
image_load_timevLLM, HFSeconds spent decoding page-image bytes for the task.
num_input_pagesvLLM, HF, HTTPPage rows presented to the inference stage.
num_valid_pagesvLLM, HF, HTTPPages sent to the model.
num_skipped_pagesvLLM, HF, HTTPPages skipped because image data was unavailable or invalid.
vllm_inference_timevLLMSeconds spent in vLLM generation, including retried inference attempts.
inference_server_request_timeHTTPSeconds spent waiting for page requests in the HTTP stage.
total_prompt_tokensvLLM, HTTPPrompt tokens reported across valid pages.
total_output_tokensvLLM, HTTPGenerated token count across valid pages.
total_output_charsvLLM, HTTPCharacters in generated text across valid pages.
num_output_length_truncatedvLLM, HTTPCompletions whose finish reason was length.
num_empty_outputsvLLM, HTTPRequests with no completion or blank completion text.
vllm_retriesvLLMInference-engine resets after generation failures. This does not count startup port-collision retries.

The HF backend records image-loading and page-count metrics, but it does not expose token, character, truncation, or retry metrics. The HTTP stage reports server-request time rather than in-process image-loading or vLLM-engine time.

Use the quality signals alongside throughput. A high num_output_length_truncated value means outputs are reaching the configured max_tokens limit (8,192 by default) and warrants inspection of those pages; num_empty_outputs and num_skipped_pages identify model-output and image failures that raw pages-per-second figures can hide. HTTP request failures are not counted as partial successes: after retries are exhausted, the task raises.

Retry Behavior

There are three separate retry paths:

  1. Engine startup: create_vllm_llm() chooses a new MASTER_PORT and retries up to max_port_retries=3 times for direct address-in-use errors or vLLM v1’s wrapped Engine core initialization failed error. Retries wait two to five seconds with jitter. Known non-retryable failures, including out-of-memory, device-side assertion, and invalid configuration errors, are raised immediately.
  2. Inference: vLLM generation is attempted up to three times. After a failed attempt, the stage resets the engine before retrying. Successful retries contribute to the vllm_retries task metric; the final exception is raised after the third failed attempt.
  3. HTTP requests: NemotronParseHTTPClientStage uses the existing AsyncOpenAIClient exponential-backoff retry behavior. Configure the retry count with inference_server_max_retries; a request that still fails raises the task so the pipeline cannot silently emit a partial document.

If startup retries are exhausted, first check the worker log for the original failure. For repeated port collisions, reduce the number of vLLM replicas starting simultaneously or set a larger stage-level max_port_retries through engine_kwargs. Do not mask CUDA out-of-memory or invalid-model errors by increasing retries; tune memory-related engine settings or correct the model configuration instead.

Render Timeout

The preprocess stage replaces signal.SIGALRM with a multiprocessing fork-based timeout (_RENDER_TIMEOUT_S = 60 by default). This is required because Xenna runs stage workers inside Ray actor processes on non-main threads, where SIGALRM raises ValueError: signal only works in main thread. The forked child inherits the PDF bytes via copy-on-write and is killed if it exceeds the timeout, reliably escaping any hung C-extension code inside pypdfium2.

You don’t need to configure this — it works automatically. If you find legitimate PDFs that take longer than 60 seconds to render, the constant lives at nemo_curator/stages/interleaved/pdf/nemotron_parse/preprocess.py.

Benchmarking

A standalone benchmark script ships at benchmarking/scripts/nemotron_parse_pdf_benchmark.py. It uses TaskPerfUtils to report stage-normalized per-GPU pages, input tokens, and output tokens per second for both in-process and inference-server runs. The benchmark harness still records full-script exec_time_s, including inference-server startup; time_taken_s covers pipeline.run(), and inference_server_startup_s reports server startup separately. Use a representative manifest and the public tutorial arguments to compare configurations before scaling to your full corpus; the repository’s nightly benchmark orchestration is not required.

Best Practices

  • Use Dynamo serving for production: keep one model replica per GPU and four HTTP workers per GPU. Start at 32 concurrent requests per worker, then benchmark any change on the target GPU and corpus. Retain in-process vLLM and HF for small validation runs or debugging.
  • Cap max_pages for outliers: very long PDFs (1000+ pages) can dominate runtime. The default 50 pages handles most academic papers and articles; raise to 200+ for book-length sources.
  • Tune pdfs_per_task for parallelism: smaller values (5–10) parallelize better across many GPUs; larger values (20–50) reduce per-task overhead on smaller clusters.
  • Set enforce_eager=True in restricted environments: vLLM’s torch.compile path can fail on certain hosts. Disabling compilation trades throughput for compatibility.
  • Pair with interleaved filters: PDF parsing produces noisy output. Chain with the Interleaved Filters (blur, CLIP score) to drop low-quality samples before training.
  • Inference Server — Dynamo lifecycle, model configuration, and troubleshooting.
  • Interleaved IO — readers and writers that consume the Parquet output of this pipeline.
  • Interleaved Filters — sample-level filters to apply after parsing.
  • Common Crawl — companion source for web-scale PDF input via CC-MAIN dumps.