NeMo Data Designer Integration
NeMo Data Designer (NDD) is a declarative data generation framework that integrates with NeMo Curator to scale synthetic data pipelines. Instead of writing imperative large language model (LLM) call logic, you define the columns, structured-field sampling, and LLM. NDD handles execution, batching, and token metric collection automatically.
How It Works
NeMo Curator wraps NDD through the DataDesignerStage, which accepts a DataDesignerConfigBuilder or a YAML config file. The stage:
- Takes input records from a
DocumentBatch. - Passes them to NDD as a seed dataset.
- Calls
DataDesigner.preview()to generate new columns for samplers, expressions, and LLM text. - Returns the enriched dataset as a new
DocumentBatchwith token usage metrics.
Prerequisites
Install the NDD dependency:
The data-designer package ships only in the synthetic data generation (SDG) extras: sdg_cpu and sdg_cuda12 for the GPU stack. It is not part of the text extras.
For local model serving on GPU, install sdg_cuda12, which pulls in cuda12, sdg_cpu, and inference_server together:
DataDesignerStage
The DataDesignerStage is the core integration point between NeMo Curator and NDD.
Parameters
The DataDesignerStage constructor accepts the following parameters:
Metrics
DataDesignerStage automatically collects and reports:
ndd_running_time: Wall-clock time for the NDDpreview()callnum_input_records/num_output_records: Record counts before and after generationinput_tokens_median_per_record/output_tokens_median_per_record: Median token counts across all LLM columns
Distributed Execution and Empty Batches
Ray can serialize and copy DataDesignerStage to workers even though the live Data Designer client contains process-local objects that are not serializable. The stage serializes its configuration builder and model providers, omits the live client, and recreates the client inside the destination worker. This process is done automatically via internal __getstate__, __setstate__, and _init_data_designer methods. Construct the stage normally, and do not create or replace its internal Data Designer client.
An upstream stage can legitimately produce an empty DocumentBatch. For example, a filter might remove every record in one input partition while other partitions still contain data. In that case, DataDesignerStage:
- Skips
DataDesigner.preview(), so it makes no generation or judge requests for that batch. - Returns an empty
DocumentBatchwhile preserving its dataset name, metadata, and stage-performance history. - Reports
0for running time, input and output record counts, and median input and output tokens.
This behavior makes it safe to place filtering stages before Data Designer or between multiple Data Designer stages without special-casing fully filtered partitions.
Building a Configuration
NDD configurations use a builder pattern. You add columns of three types:
For complete guidance on building an NDD configuration, refer to the NDD config builder reference.
Sampler Columns
Generate structured data using built-in samplers (Faker names, UUIDs, dates):
Expression Columns
Derive values from other columns using Jinja templates:
LLM Text Columns
Generate text using an LLM with prompts that reference other columns:
End-to-End Example
This example generates synthetic medical notes from seed symptom data using a local InferenceServer:
Using a Remote Provider
To use NVIDIA NIM or another hosted endpoint instead of a local server, configure the ModelProvider with the remote URL and API key:
NDD-Backed Nemotron-CC Stages
The Nemotron-CC synthetic data stages have NDD-backed equivalents that replace the AsyncOpenAIClient with NDD execution. These stages accept the same input_field, output_field, and prompt parameters, but route generation through DataDesignerStage internally.
The following stages provide NDD-backed implementations:
These stages inherit from NDDBaseSyntheticStage, which auto-builds an NDD config from the prompt fields. You configure the LLM through model_configs and model_providers instead of an AsyncOpenAIClient:
YAML Configuration
Instead of building configs in Python, you can define the entire NDD configuration in a YAML file and pass it to DataDesignerStage:
This is useful for reproducible pipelines where the generation config is versioned alongside data artifacts.
Next Steps
Refer to the following pages for related configuration and workflows:
- Inference Server: Co-locate model serving with your pipeline.
- Nemotron-CC Pipelines: Explore advanced text transformation tasks.
- Synthetic Data Generation: Review all synthetic data generation capabilities.