LLM Judge Workflow
Many data-quality decisions are difficult to express as a regular expression, numeric threshold, or conventional classifier. A large language model (LLM) judge can assess whether a response is correct, relevant, well-written, and supported by the supplied context.
LLMJudgeWorkflow turns that kind of semantic evaluation into a Curator data pipeline. You define the evidence and rubric in Jinja and YAML. Curator serves the judge model locally, and NeMo Data Designer produces a structured judgment for each record. Curator then writes or filters the results at dataset scale.
An LLM judge is an evaluation instrument, not ground truth. Validate prompts, rubrics, thresholds, and model behavior against a manually reviewed sample before using its results to filter a large dataset.
For a complete worked example, follow the LLM Judge Runner tutorial.
Why Use an LLM Judge?
Use an LLM judge when quality depends on meaning, context, or comparison and a deterministic rule would discard too much information. Typical evaluation patterns include:
The workflow does not assume a particular task or row schema. A Jinja prompt can use any fields in a JSON Lines (JSONL) or Parquet row, and the YAML configuration defines the output rubric.
Why Run It Through Curator?
A direct Python loop can call a model, but it leaves ingestion, serving, concurrency, recovery, and result filtering to the application. LLMJudgeWorkflow composes those concerns from existing Curator components:
- Dataset I/O and partitioning: Curator readers turn JSONL or Parquet files into tasks that Ray Data can schedule, and writers emit enriched output partitions.
- Local model serving: Curator starts every configured model through Dynamo/vLLM behind one OpenAI-compatible endpoint and stops the server after success or failure.
- Structured judgments: NeMo Data Designer renders the prompt from each row, calls the selected model, and adds rubric scores and reasoning under a predictable judge column.
- Pipeline composition: Multiple judge groups can run as ordered stages with separate worker counts and runtime environments. Filters run directly after the stage that produced their score.
- Operational controls: The workflow exposes Curator checkpointing, file partitioning, an optional FastText language gate, and separate controls for client concurrency and model parallelism.
This design keeps the evaluation definition—evidence, rubric, output names, and filtering policy—in reviewable text files instead of embedding it in pipeline code.
How It Works
The YAML file defines model serving, judge columns, score rubrics, execution stages, and optional filters. The Jinja files define what evidence from each input row the model sees. LLMJudgeWorkflow translates both into a running Curator pipeline:
The configuration separates concepts that affect different parts of the run:
Execution Walkthrough
One call to workflow.run() creates two connected systems: a local inference service that owns the GPUs, and a Curator pipeline that reads records and sends judge requests to that service. The following steps trace the actual control flow in workflow.py.
1. Load the YAML and Validate References
The work begins during LLMJudgeWorkflow construction, before a model is loaded or a GPU is allocated:
_load_yaml() requires the document root to be a mapping. _validate_filter_references() builds an index of every configured judge and score. It rejects filters that name an unknown output, use an unsupported operator, or appear before their judge runs. These checks fail during construction instead of after the inference server starts.
Resolving config_path also establishes the base directory for prompt_path and system_prompt_path. For example, prompt_path: prompts/quality.jinja is loaded relative to the YAML file, not relative to the shell’s working directory.
2. Turn Serving YAML Into Typed Dynamo Objects
Suppose the configuration contains this serving definition:
_start_inference_server() maps it to Curator’s serving API as follows:
The fields have distinct jobs:
server.start() launches all configured models behind one OpenAI-compatible endpoint. When it returns, server.endpoint is the URL that the judge stages call. If dynamo_server.subprocess_env.PYTHONPATH is relative, the workflow first resolves it against the YAML directory. This resolution lets configuration-specific model patches live beside the YAML.
num_replicas, tensor_parallel_size, num_workers, and max_parallel_requests act at four different layers: model replicas, GPUs per replica, Data Designer stage workers, and concurrent client requests.
3. Convert Each Judge Stage Into Data Designer Configuration
With the server running, _build_judge_stages() processes execution.stages in YAML order. For each stage, it calls build_config_builder() with the shared endpoint and only the judges assigned to that stage:
The helper builds three kinds of Data Designer object. The following abridged excerpt shows how one configured model and judge are translated:
ModelConfig tells Data Designer which API model name to request and how to make requests. LLMJudgeColumnConfig defines one output column: it loads the Jinja prompt, binds the judge to a model alias, and constrains the response to the configured scores. The full helper also attaches an optional system prompt and trace settings. ModelProvider is the connection between those logical model aliases and the running InferenceServer endpoint.
When a row reaches this stage, Data Designer uses it as a seed record and renders the Jinja variables from its fields. Data Designer then makes one structured model request for each judge. If a judge defines multiple scores, that request returns the scores together.
4. Put Filters Immediately After Their Producing Judge
Filters can be declared at the YAML top level even when the workflow contains multiple judge stages. _place_filters() finds which stage produces each referenced judge and assigns the filter to that boundary:
For example, a filter on response_quality.overall_quality is placed after the stage that creates response_quality, not automatically at the end of the pipeline. _build_filter_stages() turns that declaration into a Curator Filter whose comparison function reads judge_result[score_name]["score"]. A missing result or incompatible comparison type fails the condition and removes the row.
This placement matters when a later stage contains another expensive judge: records rejected by the first rubric never generate requests in the later stage.
5. Assemble the Curator Pipeline in Execution Order
build_pipeline() chooses readers and writers from input_format and output_format. It then converts each judge-stage description into a DataDesignerStage and appends that stage’s filters immediately afterward:
The resulting data path is:
The optional FastText gate runs before any judge. It is useful when the input contains mixed languages and the rubric is valid for only one language. Removing other records at this point avoids unnecessary LLM requests.
If a language or score filter removes every row from one partition, a downstream DataDesignerStage passes that empty batch through without calling DataDesigner.preview() or sending judge requests. It records zero input, output, timing, and token metrics for that batch, while non-empty partitions continue through the pipeline normally.
6. Execute the Pipeline and Always Stop Inference
Finally, run() connects the pieces, executes with RayDataExecutor, and records the output tasks and elapsed time. The following abridged example shows the lifecycle and is not a standalone program:
The finally block is the ownership boundary: LLMJudgeWorkflow starts the local server, so it also stops it after either success or failure. Ray has a different lifecycle. The workflow expects the caller to start Ray before run() and stop it afterward, as shown in Run the Workflow.
Helper Responsibilities and API Boundary
The workflow delegates each part of the execution flow to the following helpers:
Only LLMJudgeWorkflow is exported from nemo_curator.eval.llm_judge. The other functions explain the implementation but are not stable application APIs. Even the non-underscored assembly helpers are not re-exported.
Input and Output Contract
Suppose an input row contains an instruction and a generated response:
If the YAML defines a judge named response_quality with a score named overall_quality, the output row retains those fields and adds the structured result:
Judge names and score names are therefore part of the output schema. Choose them before a large run and keep them stable for filters and downstream analysis.
When Not to Use an LLM Judge
Prefer a deterministic filter or classifier when the decision can be measured directly—for example, character count, language ID, exact duplication, or a known metadata value. Deterministic stages are cheaper, reproducible, and easier to debug.
Do not use judge scores as an unexplained quality oracle. Model behavior can change with candidate order, prompt wording, context truncation, model version, or completion budget. Use human-reviewed calibration data, include explicit unresolved or tie outcomes when the evidence can be ambiguous, and retain reasoning while developing the rubric.
Prerequisites
Before you run the workflow, confirm the following requirements:
- Linux and one or more NVIDIA GPUs for the local Dynamo/vLLM server.
- Input records in JSONL or Parquet format.
- A Hugging Face-format model that the installed vLLM version supports, supplied as a repository ID or local weights path.
- Jinja prompts whose variables match the fields in the input records.
Install both CUDA extras. text_cuda12 provides the text-curation dependencies, including the optional FastText language gate. sdg_cuda12 provides NeMo Data Designer and the local inference server.
For a package installation, use Curator’s text dependency override file:
If you are working from a Curator source checkout instead, run:
Refer to Install NeMo Curator for supported installation methods and CUDA requirements.
Run the Workflow
First, create the judge YAML and Jinja prompt files. Then create the workflow, start Ray, and call run():
LLMJudgeWorkflow manages the model server lifecycle, but it does not start or stop Ray. Use a RayClient around each call as shown.
The repository also provides a generic CLI runner. It accepts the same workflow inputs, so after creating the YAML and Jinja files described in the next section, you can run your evaluation without writing Python:
Configure Judges
Each judge configuration consists of a YAML file and one or more Jinja prompt files. Paths such as prompt_path and system_prompt_path resolve relative to the YAML file, so keep copied examples together in one directory.
Create Configuration Files With the Authoring Skill
The Curator repository includes an llm-judge-config authoring skill that helps a coding agent create or revise the YAML and Jinja files. The skill is authoring guidance: LLMJudgeWorkflow does not load or execute the skill file at runtime.
From a Curator source checkout, ask a coding agent that can read the repository to use the skill file. Include the concrete evaluation contract in your request:
- The input row fields and which fields can be
null. - The decision the judge should make and the evidence it can use.
- The required judge names, score names, and allowed option values.
- Whether the evaluation is pointwise, pairwise, or both, including tie and insufficient-evidence behavior.
- For pairwise evaluation, whether to test both candidate orders for position bias.
- How repeated or multi-model judgments should be aggregated and when disagreement requires review.
- The model or local weights path and the available GPU resources.
For example, send the following request after replacing every placeholder with details from your evaluation:
Review the generated files before running them. Confirm the following items:
- Jinja field names match fields in the input records.
- Jinja expressions guard nullable fields.
- Judge and score names match downstream consumers.
- Categorical YAML keys retain their intended types.
- Prompt truncation leaves enough model context for the completion.
- GPU settings match the available allocation.
Run the configuration on a manually reviewed sample before scaling it up.
YAML Structure
The following skeleton serves one model and adds one structured judge column:
Model Settings
The model configuration accepts the following settings:
Approximate GPU use is num_replicas * tensor_parallel_size for each model. num_replicas creates independent model servers for horizontal throughput. tensor_parallel_size assigns multiple GPUs to each replica.
dynamo_server is optional and is passed to DynamoServerConfig. Use it for server-wide settings such as routing, existing etcd or NATS endpoints, and subprocess environment variables.
inference_server is also optional and is passed to InferenceServer. Increase health_check_timeout_s when large model weights need longer than the default startup window. Refer to Inference Server for all supported fields and disaggregated-serving examples.
Judge and Score Settings
Each item under judges makes one LLM call per input row. The call returns all scores for that judge in one structured response. The judge settings are as follows:
Score option keys can be numbers or short labels. Quote strings that YAML can interpret as special values so the parser does not convert their types. Examples include "yes", "no", "true", "false", "on", "off", and "null". Downstream filters compare option values without coercion.
Write Prompts for Input Records
Jinja variables read fields from the current JSONL or Parquet row. Guard nullable fields and delimit untrusted content so the model treats it as evidence rather than instructions:
Choose truncation limits from representative input lengths and the judge model’s context window. Leave room for the system prompt, structured-output instructions, and max_tokens. Increasing max_model_len does not truncate oversized inputs automatically.
For blind pairwise evaluation, use neutral labels such as candidate_a and candidate_b. If position bias matters, evaluate both candidate orders on a calibration sample and map the second result back to the original candidates before comparing them.
Organize Execution Stages
Every item in execution.stages becomes a separate DataDesignerStage and runs in listed order:
Group judges in one stage when they share a dependency graph or similar generation costs. Split them when you need filters between judge groups, independent worker counts, different runtime environments, or explicit Curator pipeline boundaries.
A later prompt can reference an earlier result:
Within one execution stage, NeMo Data Designer detects this dependency. Across execution stages, place the producing judge in an earlier stage. The workflow does not reorder cross-stage dependencies automatically.
Use execution.stages[].num_workers to set the Ray/Data Designer workers for that stage. It does not set model replicas or directly limit requests: each worker can submit up to the model’s max_parallel_requests.
Filter by Judge Results
Add top-level filters to retain rows that meet all configured conditions:
Supported operators are eq, ne, gt, gte, lt, lte, in, and not_in. The workflow places each top-level filter immediately after the stage that produces its judge column. You can instead put a filter under an execution stage when it must run at that later boundary.
During workflow construction, the workflow rejects filters that name an unknown judge, unknown score, or unsupported operator. For a stage-local filter, it also rejects a judge produced by a later stage. A row is removed when its judge result does not contain the configured score or when the comparison types are incompatible.
Interpret the Results
The output path contains one or more JSONL or Parquet part files. Load the entire directory instead of assuming that one input file produces one output file. Keep a stable identifier such as document_id in every row when you must join results to another dataset.
Interpret each part of a judge result separately:
- The score is the machine-readable decision constrained by the rubric’s option values. Use it for analysis or a configured filter only after calibration.
- The reasoning is a short explanation returned with the score. During prompt development, use it to find misunderstood criteria, missing evidence, and responses that latched onto irrelevant text. It is not an independent verification that the score is correct.
- A missing output row can indicate that NeMo Data Designer rejected a malformed structured response, that a completion was truncated, or that a configured filter removed the row. Compare input and output counts and inspect traces on a small retry before scaling.
When multiple models apply the same rubric, extract each judge’s nested score into a separate analysis column. Agreement can identify easy cases. Disagreement identifies records, rubric wording, or model behavior that needs review. Do not average the scores until you understand the cause of the disagreement.
run() returns a WorkflowRunResult containing the output tasks and total_time metadata. Detailed token and record-count metrics are logged by the underlying Data Designer stages.
Workflow Parameters
The LLMJudgeWorkflow constructor accepts the following parameters:
Calibrate, Scale, and Troubleshoot
Use the following practices when you move beyond an initial test run:
- Build a calibration set. Start with records that a human has reviewed. Include clear passes, clear failures, boundary cases, missing fields, and cases where
unresolvedis the right answer. - Test prompt sensitivity. Swap pairwise candidates, remove source-identifying names, change formatting without changing meaning, and insert instruction-like text into evidence. Change one factor at a time. A score that changes for an irrelevant reason exposes a prompt or rubric weakness.
- Inspect before filtering. Run without score filters first, review score distributions and reasoning, and choose thresholds from observed behavior rather than from the numeric scale alone.
- Check the full data path. Confirm rendered prompts, structured outputs, context-length errors, malformed responses, and input/output row counts before increasing concurrency.
- Tune the correct layer.
max_parallel_requestscontrols client concurrency,num_replicascontrols horizontal serving capacity,tensor_parallel_sizecontrols GPUs per replica, andnum_workerscontrols stage workers. - Use multiple input files. One JSONL or Parquet file is one source task. Multiple files allow pipeline stages and Slurm job-array elements to make progress independently.
- Use traces selectively.
all_messagesduplicates prompt content in output and can substantially increase output size. - Preserve checkpoint identity. Reuse a checkpoint directory only when retrying the same logical input and configuration.
- Check completion budgets. A response cut off by a small
max_tokensvalue can fail structured-output validation and result in a missing row.
Troubleshoot Common Failures
Use the following table to diagnose common workflow failures:
For distributed execution patterns, refer to Multi-Node Ray on Slurm and SLURM Job Arrays. Refer to the LLM judge source directory for the implementation. The LLM judge tutorial source contains the maintained examples.