LLM Judge Runner
LLM Judge Runner
In this tutorial, you use LLMJudgeWorkflow to run a large language model (LLM) as a judge. The judge compares two text extractions of the same Common Crawl document. Curator prepares each record with raw HTML, jusText output, and Trafilatura output. A locally served LLM selects the most useful representation and scores content fidelity, boilerplate removal, and semantic disagreement.
This tutorial teaches you how to:
- Prepare a small, reproducible comparison dataset with Curator.
- Adapt the supplied Qwen judge configuration to your model and GPUs.
- Run structured LLM judgments over JSONL records.
- Inspect the winning extraction, rubric scores, reasoning, and traces.
- Decide whether the rubric is ready for a larger evaluation.
For the reusable workflow architecture and complete configuration contract, refer to LLM Judge Workflow.
The supplied prompts, score descriptions, model settings, and thresholds demonstrate the integration. They are not a calibrated extraction-quality benchmark. Validate them against human-reviewed records before using the results to select or filter training data.
Before You Start
Before you start the tutorial, confirm the following requirements:
- A Linux checkout of the Curator repository.
- Access to Common Crawl over HTTPS, or
s5cmdaccess when using--use-aws-to-download. - One or more NVIDIA GPUs with enough aggregate memory for the selected judge model.
- A Hugging Face-format instruction model available as local weights or a repository ID supported by the installed vLLM version.
This tutorial requires both CUDA extras: text_cuda12 for Common Crawl processing and text extraction, and sdg_cuda12 for NeMo Data Designer and local inference serving. text_cuda12 also supplies the optional FastText language gate, although the main tutorial command does not enable that gate. From the Curator source checkout, create or update the environment with:
Run this command and the remaining tutorial commands from the repository root. Refer to Install NeMo Curator for CUDA requirements and other supported installation methods.
The language gate is optional. To enable it, download FastText’s lid.176.bin language-identification model:
Then add --language en --fasttext-langid-model-path models/fasttext/lid.176.bin to the judge-runner command in step 3. Omit those flags to skip the gate. The main tutorial path does not require this download.
Files Used in This Tutorial
The example separates data preparation, workflow execution, serving configuration, and prompt text so that each part can be changed independently:
The file names in this table are relative to tutorials/eval/llm_judge. Run all shell commands in the rest of the tutorial from the Curator repository root. The Jinja paths inside a judge YAML are instead resolved relative to that YAML file.
Concepts Overview: What the Example Evaluates
The preparation script downloads Common Crawl WARC data and applies both extractors to every decodable record:
The single-model YAML runs two judges per row:
qwen3_8_27b_text_extraction_judgmentchoosescandidate_a,candidate_b,raw, ornone, then scores content fidelity and boilerplate removal from 1 through 5.qwen3_8_27b_text_extraction_disagreementscores how materially the two extracted candidates differ from 1 through 5.
The prompts keep extractor identities hidden from the model: candidate_a is jusText and candidate_b is Trafilatura. This mapping is fixed in the supplied Jinja files and YAML comments.
1. Prepare a Small Comparison Dataset
Download one WARC file and process at most 100 records:
The script starts a local RayClient, runs a RayDataExecutor pipeline, and writes JSONL parts beneath data/cc_extractions. It retains a row when one extractor or language detection fails, so the extracted fields and language can legitimately be null.
The language field created here is preparation metadata: lang_detect() attempts to label each downloaded document, but the preparation pipeline does not filter by that label. This is separate from LLMJudgeWorkflow’s optional FastText gate, which runs before the LLM judges only when you pass --language and --fasttext-langid-model-path to the runner.
List the generated parts and inspect a record:
Confirm that the record contains raw_text, justext_text, and trafilatura_text. The supplied prompts are null-safe and truncate raw HTML to 12,000 characters and each candidate to 8,000 characters.
The defaults intentionally keep this first run small, but one Common Crawl WARC can still take time to download and process. Keep the downloaded WARC directory if you expect to repeat the preparation step.
2. Copy and Configure the Judge Files
Copy the YAML and adjacent Jinja files into a working directory so you do not edit the maintained example in place:
Open configs/cc_extract_judge/text_extraction_qwen_judge.yaml. The serving portion has four layers. Edit the marked values for your model and allocation:
model tells vLLM what to load. served_model_name is the API-facing name that Data Designer requests and can differ from a local weights path. dynamo_model controls server-side GPU use, while inference_parameters controls client requests. max_parallel_requests does not create model replicas. execution.stages[].num_workers controls Ray/Data Designer workers and is independent of both settings.
Add model-specific engine_kwargs, such as trust_remote_code, only when the selected architecture requires them. The supplied DYN_SYSTEM_PORT: "0" lets the Dynamo subprocess choose an available system port. The longer health-check timeout allows a large model to finish loading before Curator treats startup as failed.
As supplied, the YAML requests four GPUs: four replicas multiplied by one GPU per replica. Reduce num_replicas when fewer replicas fit, or increase tensor_parallel_size when one model replica must span multiple GPUs. Approximate GPU use for one model is:
The example disables Qwen thinking with extra_body.chat_template_kwargs.enable_thinking: false. This setting prevents private deliberation from consuming the completion budget needed for structured JSON. Remove the setting when the selected model does not support it. Do not lower max_model_len below the rendered prompt plus the requested completion budget. The Jinja character limits do not automatically adapt to this setting.
Do not change the Jinja field names for this tutorial. They match the dataset created in step 1. When adapting the workflow to another task, use the llm-judge-config authoring skill to design new YAML and Jinja files from the actual input schema.
3. Run the Judge Workflow
Run the two Qwen judge stages over the prepared JSONL parts:
The command crosses two lifecycle boundaries: the CLI owns Ray, while LLMJudgeWorkflow owns the local inference server. The complete flow is:
run_llm_judge.pyparses the flags and constructsLLMJudgeWorkflow. Construction loads the YAML, resolves it to an absolute path, and validates judge and filter references before model startup.- The runner starts
RayClient.LLMJudgeWorkflowdeliberately does not own Ray because it can be embedded in a longer-lived Curator application. workflow.run()convertsmodels[].dynamo_modelintoDynamoVLLMModelConfig, convertsdynamo_serverintoDynamoServerConfig, and startsInferenceServer.- The workflow passes
InferenceServer.endpointto NeMo Data Designer, loads the adjacent Jinja files, and builds the extraction-quality and semantic-disagreementDataDesignerStageobjects. RayDataExecutorreads each JSONL partition. Data Designer renders the row into each judge prompt and sends structured requests to the shared Dynamo/vLLM endpoint.- The writer preserves the input fields, adds both judge columns, and writes JSONL parts beneath
data/qwen_judgements. LLMJudgeWorkflowstopsInferenceServerin afinallyblock. Afterworkflow.run()returns or raises, the runner stopsRayClientin its ownfinallyblock.
The active data path for the supplied configuration is:
If you enable a workflow FastText language gate, it is inserted after JsonlReader and before the first Data Designer stage. If you enable a score filter, it is inserted directly after the stage that creates the referenced judge column, not deferred until after both judges.
There are no active score filters in the supplied YAML. The workflow therefore does not intentionally remove rows based on the three rubrics. NeMo Data Designer can still reject rows whose model response does not satisfy the structured output contract, so compare input and output counts.
4. Inspect Structured Judgments
List the output files and inspect the first record:
Each output record preserves the input fields and adds two judge columns. The extraction-quality column has this shape:
Check at least these conditions on a manually reviewed sample:
best_extraction.scoreselects the representation you would choose from the rendered evidence.- The 1–5 scores follow their written option anchors rather than an unstated interpretation of “quality.”
- Reasoning cites relevant evidence and does not follow instruction-like text embedded in the source HTML.
- Null or empty candidates receive sensible outcomes.
- The
last_messagetrace on the extraction-quality judge shows the expected rendered evidence and structured response. - The number of written records matches expectations.
5. Analyze Judge Decisions
Load every non-empty JSONL part and summarize the categorical decision:
A writer can create an empty part when an upstream filter or structured-output rejection removes every row in a batch. The guard skips empty and whitespace-only files before calling pd.read_json() and stops with a clear message when the output directory contains no judged records.
Map candidate_a back to justext_text and candidate_b back to trafilatura_text during analysis. Read the score reasoning for unexpected winners, low-fidelity results, and records with a high semantic_disagreement score.
Do not choose a production filter threshold from the option labels alone. Review the score distribution against human decisions, then decide whether errors near the proposed threshold are acceptable.
6. Optionally Filter High-Fidelity Records
After calibration, uncomment and adjust the example’s top-level filters block. For example, the supplied commented filter retains rows with a content-fidelity score of at least 4:
Write this filtered run to a new output directory. Because the configuration changed, also use a new checkpoint directory rather than reusing the checkpoint from the unfiltered run.
7. Optionally Compare Two Judge Models
text_extraction_qwen_gemma_judges.yaml applies the same two rubrics with Qwen and Gemma. As supplied, it requests eight GPUs: four single-GPU replicas for each model.
Before running it:
- Set both model paths and API-facing names.
- Adjust replicas and tensor parallelism for each model.
- Confirm that both models can share the configured Dynamo worker environment.
- Use a new output and checkpoint directory.
The result contains one extraction-quality column and one semantic-disagreement column per model. Extract corresponding nested scores into separate analysis columns and inspect records where the models disagree. A disagreement can reveal ambiguous source evidence, an underspecified rubric, or model-specific behavior. Do not automatically average the scores.
Summary
After completing this tutorial, you can perform the following tasks:
- Prepare Common Crawl records that contain raw HTML and two extraction candidates.
- Configure model serving, judge prompts, rubrics, execution stages, and optional filters.
- Run
LLMJudgeWorkflowwhile managing the separate Ray and inference-server lifecycles. - Inspect structured scores and safely analyze output directories that contain empty JSONL parts.
- Calibrate filter thresholds or compare two judge models before scaling the evaluation.
What Is Next
After you validate the example, use the following options to extend the evaluation:
- Test position bias by adding a second judge whose prompt swaps
candidate_aandcandidate_b, then map its result back to the original extractor identities before comparing decisions. - Increase
--record-limit, WARC count, client concurrency, or model replicas only after the small run produces valid prompts and stable structured results. - Split larger inputs across multiple JSONL files so Curator stages and Slurm job-array elements have multiple source tasks to schedule.
- Review the generic LLM Judge Workflow guide for execution stages, filters, traces, configuration fields, and scaling controls.
- Read the maintained tutorial source and examples when adapting the workflow.