nemo_automodel.recipes.retrieval

View as Markdown

Submodules

Package Contents

Classes

NameDescription
EmbeddingDistillRecipeRecipe for Stage-1 embedding distillation on bi-encoder backbones.
MineHardNegativesRecipeRecipe for mining hard negatives for encoder training.
TrainBiEncoderRecipeRecipe for training encoder models with contrastive learning.
TrainCrossEncoderRecipe-

API

class nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe()

Bases: TrainBiEncoderRecipe

Recipe for Stage-1 embedding distillation on bi-encoder backbones.

nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._build_optimizer_param_groups() -> list[dict[str, typing.Any]]

Build optimizer groups with projection params isolated before checkpoint restore.

nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._extract_scoring_reps(
model_output
)

Select the pooled student embedding for validation scoring.

RetrieverStudentWithProjection.forward returns (pooled, projected, intermediate_outputs). The pooled embedding is the student’s native retrieval representation (and what training’s InfoNCE terms score with), so it is what the inherited validation loop should compare against.

nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._forward_backward_step(
idx,
batch,
loss_buffer,
num_batches,
is_train: bool = True
)
nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._projection_parameters() -> list[torch.nn.Parameter]
nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._run_train_optim_step(
batches,
max_grad_norm = None
)
nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._sync_projection_gradients() -> None

Average projection gradients that are outside FSDP/DDP wrapping.

nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe._sync_projection_parameters() -> None

Keep the rank-local projection head replicated across DP ranks.

nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe.save_checkpoint(
epoch: int,
step: int,
train_loss: float,
val_loss: dict[str, float] | None = None,
best_metric_key: str = 'default'
)
nemo_automodel.recipes.retrieval.distill_bi_encoder.EmbeddingDistillRecipe.setup()
class nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe(
cfg
)

Recipe for mining hard negatives for encoder training.

This class orchestrates hard negative mining, from setup to mining execution. Hard negatives are documents that are semantically similar to the query but are not relevant, making them valuable for training more discriminative models.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._build_document_mappings()

Build bidirectional mappings between document IDs and indices.

Creates doc_to_idx and idx_to_doc dictionaries for efficient lookup during the mining process. Documents are sorted by ID for deterministic ordering across runs.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._build_negative_docs_by_question_id() -> typing.Dict[str, typing.List[typing.Dict[str, typing.Any]]]

Build mapping from question_id to mined negative documents with scores.

If use_negatives_from_file is True, supplied negatives are prepended to mined negatives (with score=-1 since we don’t compute their scores).

Returns: Dict[str, List[Dict[str, Any]]]

Dict mapping question_id to list of {“id”: doc_id, “score”: score} dicts.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._build_positive_scores_by_question_id() -> typing.Dict[str, typing.List[float]]

Build mapping from question_id to positive document scores.

Scores are in the same order as pos_doc in the original data.

Returns: Dict[str, List[float]]

Dict mapping question_id to list of positive scores.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._configure_tokenizer()

Load and configure tokenizer with appropriate settings.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_all_documents() -> numpy.ndarray

Encode all documents in corpus, chunk by chunk.

Returns: np.ndarray

numpy array of document embeddings [num_docs, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_chunk_distributed(
doc_indices: list[int],
cache_path: pathlib.Path
) -> numpy.ndarray

Encode a chunk of documents in distributed mode.

Shards the documents within the chunk across ranks, encodes each shard, and assembles on rank0.

Parameters:

doc_indices
list[int]

Corpus indices to encode.

cache_path
Path

Path for caching the assembled chunk.

Returns: np.ndarray

Assembled document embeddings for this chunk (rank0 only).

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_chunk_local(
doc_indices: list[int],
cache_path: pathlib.Path | None
) -> numpy.ndarray

Encode a chunk of documents locally (single-process).

Parameters:

doc_indices
list[int]

Corpus indices to encode.

cache_path
Path | None

Optional path for caching.

Returns: np.ndarray

Document embeddings for this chunk.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_document_indices(
doc_indices: list[int]
) -> numpy.ndarray

Fetch and encode corpus documents in bounded batches and stable order.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_documents_chunk(
doc_indices: typing.List[int],
cache_path: pathlib.Path | None = None
) -> numpy.ndarray

Encode a chunk of documents into embeddings.

Parameters:

doc_indices
List[int]

List of document indices to encode.

cache_path
Path | NoneDefaults to None

Optional path to save/load chunk cache.

Returns: np.ndarray

numpy array of document embeddings [num_docs, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_queries() -> numpy.ndarray

Encode all queries into embeddings.

Uses self.query_prefix and self.query_max_length from mining config.

Returns: np.ndarray

numpy array of query embeddings [num_queries, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_queries_sharded() -> numpy.ndarray

Encode queries sharded across ranks and assemble on rank0.

This is useful when the number of queries is large (e.g., 100k+), and we want to utilize multiple GPUs for query embedding generation without sharding mining/scoring.

Requires cache_embeddings_dir so ranks can write shard files and rank0 can assemble.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_texts(
texts: typing.List[str],
batch_size: int,
max_length: int,
prefix: str = ''
) -> numpy.ndarray

Encode texts into embeddings.

Parameters:

texts
List[str]

List of text strings to encode.

batch_size
int

Batch size for encoding.

max_length
int

Maximum sequence length for tokenization.

prefix
strDefaults to ''

Optional prefix to prepend to each text.

Returns: np.ndarray

numpy array of embeddings [num_texts, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._extract_mining_params()

Extract all mining parameters from configuration.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._generate_embeddings() -> typing.Tuple[numpy.ndarray, numpy.ndarray]

Generate embeddings for queries and documents.

Handles caching and orchestrates the encoding process.

Returns: Tuple[np.ndarray, np.ndarray]

Tuple of (query_embeddings, document_embeddings).

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._get_document_text(
doc: dict
) -> str

Extract text from document dict, checking common field names.

Parameters:

doc
dict

Document dictionary from corpus.

Returns: str

Document text string.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._get_mining_args_dict() -> typing.Dict[str, typing.Any]

Get dictionary of mining arguments for output metadata.

Returns: Dict[str, Any]

Dict containing all mining parameters for reproducibility.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._get_mining_param(
name,
default = None
)

Get mining parameter from config, with fallback to defaults.

Parameters:

name

Parameter name.

default
Defaults to None

Default value if not in config or MINING_DEFAULTS.

Returns:

Parameter value.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._has_full_embeddings_cache() -> bool

Check if the consolidated (rank0) embedding cache exists.

This is intentionally a lightweight existence check (no file reads), used to avoid redundant IO on non-main ranks in distributed runs.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._load_cached_chunk(
cache_path: pathlib.Path
) -> numpy.ndarray | None

Load a fully-assembled chunk cache if it exists.

In distributed mode, only rank0 loads the cache to avoid redundant IO.

Parameters:

cache_path
Path

Path to cached chunk file.

Returns: np.ndarray | None

Cached embeddings array, or None if cache doesn’t exist.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._load_data()

Load dataset and corpus from the input QA file.

Uses load_datasets() from retrieval_dataset.py to load the questions dataset and corpus dictionary. Validates that only a single corpus is referenced.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._load_embeddings_from_cache() -> typing.Tuple[numpy.ndarray | None, numpy.ndarray | None]

Load query and document embeddings from cache.

Returns: Tuple[np.ndarray | None, np.ndarray | None]

Tuple of (query_embeddings, document_embeddings), or (None, None) if not found.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._mine_hard_negatives(
query_embeddings: numpy.ndarray,
document_embeddings: numpy.ndarray,
pos_doc_indices: typing.List[typing.List[int]],
batch_size: int,
num_negs: int,
hard_neg_margin: float | None = None,
hard_neg_margin_type: str | None = None
) -> typing.Tuple[typing.List[typing.List[int]], typing.List[typing.List[float]], typing.List[typing.List[float]]]

Mine hard negatives for each query.

This implementation uses the following key behaviors:

  • Deduplicates positive indices before masking (avoids double-masking)
  • Preserves original order of positive scores (matches pos_doc order in input)
  • Uses vectorized batch-level margin filtering for efficiency
  • Uses batch-level topk for efficiency

Parameters:

query_embeddings
np.ndarray

Query embeddings [num_queries, embedding_dim].

document_embeddings
np.ndarray

Document embeddings [num_docs, embedding_dim].

pos_doc_indices
List[List[int]]

List of positive document indices for each query.

batch_size
int

Number of queries to process per batch.

num_negs
int

Number of hard negatives to mine per query.

hard_neg_margin
float | NoneDefaults to None

Optional margin for filtering false negatives.

hard_neg_margin_type
str | NoneDefaults to None

“perc” (percentage) or “abs” (absolute).

Returns: Tuple[List[List[int]], List[List[float]], List[List[float]]]

Tuple of:

  • neg_indices: List of hard negative indices per query
  • neg_scores: Similarity scores for each hard negative
  • pos_scores: Similarity scores for each positive document
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._prepare_data()

Extract query texts and document indices from questions dataset.

Iterates through the questions dataset and extracts:

  • Query texts (with EMPTY_QUESTION placeholder for empty queries)
  • Question IDs
  • Corpus IDs
  • Positive document indices (mapped via doc_to_idx)
  • Supplied negative document indices (if use_negatives_from_file is True)
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._print_configuration()

Print mining configuration summary.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._save_embeddings_to_cache(
query_embeddings: numpy.ndarray,
document_embeddings: numpy.ndarray
) -> None

Save query and document embeddings to cache.

Parameters:

query_embeddings
np.ndarray

Query embeddings array.

document_embeddings
np.ndarray

Document embeddings array.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._synchronize_ranks()

Synchronize all distributed ranks with a barrier.

Handles device-specific barrier calls for CUDA vs CPU.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._unload_model()

Unload the model from GPU memory after embedding generation.

This frees up GPU memory for the mining phase, which only operates on embeddings and doesn’t need the model parameters.

Model metadata (pooling, l2_normalize) is extracted before unloading to ensure it’s available for output generation.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._validate_mining_params()

Validate required mining parameters.

Raises:

  • ValueError: If any required parameter is missing or invalid.
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._write_output() -> None

Write the output JSON file with mined hard negatives.

The output format:

  • Preserves all top-level keys from input (corpus, etc.)
  • Adds mining metadata section with parameters used
  • Replaces neg_doc with newly mined negatives
  • Adds similarity scores to all documents (pos_doc and neg_doc)
  • Removes legacy score fields (pos_score, neg_scores) if present
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe.run()

Run the hard negative mining pipeline.

Generates query and document embeddings, mines hard negatives using similarity scores with margin filtering, and writes the output file.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe.setup()

Build all components needed for hard negative mining.

class nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe(
cfg
)

Bases: BaseRecipe

Recipe for training encoder models with contrastive learning.

cfg
temperature
= self.cfg.get('temperature', 1.0)
nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe._build_optimizer_param_groups() -> list[dict[str, typing.Any]]

Build optimizer parameter groups for trainable model parameters.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe._extract_scoring_reps(
model_output
)

Return the embedding tensor used for validation scoring from a forward output.

The base bi-encoder forward returns an embedding tensor directly. Subclasses whose forward returns a richer structure (e.g. the distillation student, which returns (pooled, projected, intermediate_outputs)) should override this to select the tensor to score with.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe._forward_backward_step(
idx,
batch,
loss_buffer,
num_batches,
is_train: bool = True
)

Forward and backward pass for a single micro-batch.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe._run_train_optim_step(
batches,
max_grad_norm = None
)

Run one optimization step with gradient accumulation.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe._run_validation_epoch(
val_dataloader
)

Run validation for one epoch and compute loss, accuracy@1, and MRR.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe._validate_model(
model: torch.nn.Module
) -> None

Validate recipe-specific model settings before constructing the optimizer.

Parameters:

model
torch.nn.Module

Constructed retrieval model, including infrastructure wrappers.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe.log_train_metrics(
)
nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe.log_val_metrics(
)
nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe.run_train_validation_loop()

Run the training loop over all epochs and batches.

nemo_automodel.recipes.retrieval.train_bi_encoder.TrainBiEncoderRecipe.setup()

Build all components needed for training/validation/logging/checkpointing.

class nemo_automodel.recipes.retrieval.train_cross_encoder.TrainCrossEncoderRecipe()

Bases: TrainBiEncoderRecipe

nemo_automodel.recipes.retrieval.train_cross_encoder.TrainCrossEncoderRecipe._forward_backward_step(
idx,
batch,
loss_buffer,
num_batches,
is_train: bool = True,
modality_loss_buffers = None
)

Forward and backward pass for a single micro-batch.

nemo_automodel.recipes.retrieval.train_cross_encoder.TrainCrossEncoderRecipe._run_train_optim_step(
batches,
max_grad_norm = None
)
nemo_automodel.recipes.retrieval.train_cross_encoder.TrainCrossEncoderRecipe._run_validation_epoch(
val_dataloader
)

Run validation for one epoch and compute loss, accuracy@1, and MRR.

nemo_automodel.recipes.retrieval.train_cross_encoder.TrainCrossEncoderRecipe._validate_model(
model: torch.nn.Module
) -> None

Validate the effective temperature applied by the constructed model.

nemo_automodel.recipes.retrieval.train_cross_encoder.TrainCrossEncoderRecipe.log_train_metrics(
)