nemo_curator.eval.llm_judge.workflow
nemo_curator.eval.llm_judge.workflow
Config-driven text LLM judge workflow, run through a NeMo Curator pipeline.
The input records may have any text schema. The Jinja templates and score
rubrics in the judge config define which fields are evaluated, what the
judge returns, and how judges are grouped into execution.stages, each of
which runs as its own NDD stage.
Module Contents
Classes
Functions
Data
API
Bases: WorkflowBase
End-to-end config-driven LLM judge workflow.
Loads a judge config YAML (models, Jinja prompt templates, score rubrics,
and execution.stages), starts a Dynamo/vLLM inference server hosting
the configured judge models, then runs one Curator pipeline containing:
reader -> optional FastText language gate -> one NDD DataDesignerStage
(+ its filters) per judge stage -> writer.
Run the complete LLM judge pipeline.
Returns: WorkflowRunResult
WorkflowRunResult containing the pipeline output tasks and timing metadata.
Build Curator filters that retain rows satisfying every configured condition.
Build an optional FastText language gate without retaining its score column.
Return whether one NDD judge result satisfies a declarative comparison.
Place top-level filters after the NDD stage that produces their judge column.
Start all configured Dynamo models behind one OpenAI-compatible endpoint.
Ensure filters refer to a configured judge output column and rubric score.
Build one NDD configuration for a selected group of judge columns.
Build a streaming pipeline with an optional language gate, NDD stages, filters, and writer.