nemo_curator.eval.llm_judge.workflow

View as Markdown

Config-driven text LLM judge workflow, run through a NeMo Curator pipeline.

The input records may have any text schema. The Jinja templates and score rubrics in the judge config define which fields are evaluated, what the judge returns, and how judges are grouped into execution.stages, each of which runs as its own NDD stage.

Module Contents

Classes

NameDescription
LLMJudgeWorkflowEnd-to-end config-driven LLM judge workflow.

Functions

NameDescription
_build_filter_stagesBuild Curator filters that retain rows satisfying every configured condition.
_build_language_filter_stageBuild an optional FastText language gate without retaining its score column.
_keep_judge_scoreReturn whether one NDD judge result satisfies a declarative comparison.
_load_yaml-
_place_filtersPlace top-level filters after the NDD stage that produces their judge column.
_read_template-
_start_inference_serverStart all configured Dynamo models behind one OpenAI-compatible endpoint.
_validate_filter_referencesEnsure filters refer to a configured judge output column and rubric score.
build_config_builderBuild one NDD configuration for a selected group of judge columns.
build_pipelineBuild a streaming pipeline with an optional language gate, NDD stages, filters, and writer.

Data

DataFormat

FilterOperator

_FILTER_OPERATORS

API

class nemo_curator.eval.llm_judge.LLMJudgeWorkflow(
judge_config: str | pathlib.Path,
input_path: str,
output_path: str,
files_per_partition: int | None = None,
language: str | None = None,
fasttext_langid_model_path: str | None = None,
min_langid_score: float = 0.3,
language_text_field: str = 'raw_text',
checkpoint_path: str | None = None
)
Dataclass

Bases: WorkflowBase

End-to-end config-driven LLM judge workflow.

Loads a judge config YAML (models, Jinja prompt templates, score rubrics, and execution.stages), starts a Dynamo/vLLM inference server hosting the configured judge models, then runs one Curator pipeline containing: reader -> optional FastText language gate -> one NDD DataDesignerStage (+ its filters) per judge stage -> writer.

checkpoint_path
str | None = None
config
dict[str, object] = field(init=False)
config_path
Path = field(init=False)
fasttext_langid_model_path
str | None = None
files_per_partition
int | None = None
input_format
DataFormat = 'jsonl'
input_path
str
judge_config
str | Path
language
str | None = None
language_text_field
str = 'raw_text'
min_langid_score
float = 0.3
output_format
DataFormat = 'jsonl'
output_path
str
nemo_curator.eval.llm_judge.LLMJudgeWorkflow.__post_init__() -> None
nemo_curator.eval.llm_judge.LLMJudgeWorkflow._build_judge_stages(
endpoint: str
) -> list[tuple[str, data_designer.config.DataDesignerConfigBuilder, list[data_designer.config.ModelProvider], dict[str, object] | None, int | None, list[dict[str, object]]]]
nemo_curator.eval.llm_judge.LLMJudgeWorkflow.run() -> nemo_curator.pipeline.workflow.WorkflowRunResult

Run the complete LLM judge pipeline.

Returns: WorkflowRunResult

WorkflowRunResult containing the pipeline output tasks and timing metadata.

nemo_curator.eval.llm_judge.workflow._build_filter_stages(
filters: list[dict[str, object]],
name_prefix: str
) -> list[nemo_curator.stages.text.filters.Filter]

Build Curator filters that retain rows satisfying every configured condition.

nemo_curator.eval.llm_judge.workflow._build_language_filter_stage(
language: str | None,
model_path: str | None,
min_score: float,
text_field: str
) -> nemo_curator.stages.text.filters.ScoreFilter | None

Build an optional FastText language gate without retaining its score column.

nemo_curator.eval.llm_judge.workflow._keep_judge_score(
judge_result: object,
score_name: str,
expected: object
) -> bool

Return whether one NDD judge result satisfies a declarative comparison.

nemo_curator.eval.llm_judge.workflow._load_yaml(
path: pathlib.Path
) -> dict[str, object]
nemo_curator.eval.llm_judge.workflow._place_filters(
config: dict[str, object],
stages: list[dict[str, object]]
) -> list[list[dict[str, object]]]

Place top-level filters after the NDD stage that produces their judge column.

nemo_curator.eval.llm_judge.workflow._read_template(
path: str,
config_path: pathlib.Path
) -> str
nemo_curator.eval.llm_judge.workflow._start_inference_server(
config: dict[str, object],
models: list[dict[str, object]],
config_path: pathlib.Path

Start all configured Dynamo models behind one OpenAI-compatible endpoint.

nemo_curator.eval.llm_judge.workflow._validate_filter_references(
config: dict[str, object],
stages: list[dict[str, object]]
) -> None

Ensure filters refer to a configured judge output column and rubric score.

nemo_curator.eval.llm_judge.workflow.build_config_builder(
config_path: str | pathlib.Path,
endpoint: str,
models: list[dict[str, object]],
judges: list[dict[str, object]]
) -> tuple[data_designer.config.DataDesignerConfigBuilder, list[data_designer.config.ModelProvider]]

Build one NDD configuration for a selected group of judge columns.

nemo_curator.eval.llm_judge.workflow.build_pipeline(
input_path: str,
output_path: str,
judge_stages: list[tuple[str, data_designer.config.DataDesignerConfigBuilder, list[data_designer.config.ModelProvider], dict[str, object] | None, int | None, list[dict[str, object]]]],
language_filter_stage: nemo_curator.stages.text.filters.ScoreFilter | None,
files_per_partition: int | None

Build a streaming pipeline with an optional language gate, NDD stages, filters, and writer.

nemo_curator.eval.llm_judge.workflow.DataFormat = Literal['jsonl', 'parquet']
nemo_curator.eval.llm_judge.workflow.FilterOperator = Literal['eq', 'ne', 'gt', 'gte', 'lt', 'lte', 'in', 'not_in']
nemo_curator.eval.llm_judge.workflow._FILTER_OPERATORS = {'eq', 'ne', 'gt', 'gte', 'lt', 'lte', 'in', 'not_in'}