nemo_automodel.recipes.retrieval
nemo_automodel.recipes.retrieval
Submodules
nemo_automodel.recipes.retrieval.distill_bi_encodernemo_automodel.recipes.retrieval.mine_hard_negativesnemo_automodel.recipes.retrieval.mining_encodernemo_automodel.recipes.retrieval.train_bi_encodernemo_automodel.recipes.retrieval.train_cross_encoder
Package Contents
Classes
API
Bases: TrainBiEncoderRecipe
Recipe for Stage-1 embedding distillation on bi-encoder backbones.
Build optimizer groups with projection params isolated before checkpoint restore.
Select the pooled student embedding for validation scoring.
RetrieverStudentWithProjection.forward returns
(pooled, projected, intermediate_outputs). The pooled embedding is the student’s
native retrieval representation (and what training’s InfoNCE terms score with), so it
is what the inherited validation loop should compare against.
Average projection gradients that are outside FSDP/DDP wrapping.
Keep the rank-local projection head replicated across DP ranks.
Recipe for mining hard negatives for encoder training.
This class orchestrates hard negative mining, from setup to mining execution. Hard negatives are documents that are semantically similar to the query but are not relevant, making them valuable for training more discriminative models.
Build bidirectional mappings between document IDs and indices.
Creates doc_to_idx and idx_to_doc dictionaries for efficient lookup during the mining process. Documents are sorted by ID for deterministic ordering across runs.
Build mapping from question_id to mined negative documents with scores.
If use_negatives_from_file is True, supplied negatives are prepended to mined negatives (with score=-1 since we don’t compute their scores).
Returns: Dict[str, List[Dict[str, Any]]]
Dict mapping question_id to list of {“id”: doc_id, “score”: score} dicts.
Build mapping from question_id to positive document scores.
Scores are in the same order as pos_doc in the original data.
Returns: Dict[str, List[float]]
Dict mapping question_id to list of positive scores.
Load and configure tokenizer with appropriate settings.
Encode all documents in corpus, chunk by chunk.
Returns: np.ndarray
numpy array of document embeddings [num_docs, embedding_dim].
Encode a chunk of documents in distributed mode.
Shards the documents within the chunk across ranks, encodes each shard, and assembles on rank0.
Parameters:
Corpus indices to encode.
Path for caching the assembled chunk.
Returns: np.ndarray
Assembled document embeddings for this chunk (rank0 only).
Encode a chunk of documents locally (single-process).
Parameters:
Corpus indices to encode.
Optional path for caching.
Returns: np.ndarray
Document embeddings for this chunk.
Fetch and encode corpus documents in bounded batches and stable order.
Encode a chunk of documents into embeddings.
Parameters:
List of document indices to encode.
Optional path to save/load chunk cache.
Returns: np.ndarray
numpy array of document embeddings [num_docs, embedding_dim].
Encode all queries into embeddings.
Uses self.query_prefix and self.query_max_length from mining config.
Returns: np.ndarray
numpy array of query embeddings [num_queries, embedding_dim].
Encode queries sharded across ranks and assemble on rank0.
This is useful when the number of queries is large (e.g., 100k+), and we want to utilize multiple GPUs for query embedding generation without sharding mining/scoring.
Requires cache_embeddings_dir so ranks can write shard files and rank0 can assemble.
Encode texts into embeddings.
Parameters:
List of text strings to encode.
Batch size for encoding.
Maximum sequence length for tokenization.
Optional prefix to prepend to each text.
Returns: np.ndarray
numpy array of embeddings [num_texts, embedding_dim].
Extract all mining parameters from configuration.
Generate embeddings for queries and documents.
Handles caching and orchestrates the encoding process.
Returns: Tuple[np.ndarray, np.ndarray]
Tuple of (query_embeddings, document_embeddings).
Extract text from document dict, checking common field names.
Parameters:
Document dictionary from corpus.
Returns: str
Document text string.
Get dictionary of mining arguments for output metadata.
Returns: Dict[str, Any]
Dict containing all mining parameters for reproducibility.
Get mining parameter from config, with fallback to defaults.
Parameters:
Parameter name.
Default value if not in config or MINING_DEFAULTS.
Returns:
Parameter value.
Check if the consolidated (rank0) embedding cache exists.
This is intentionally a lightweight existence check (no file reads), used to avoid redundant IO on non-main ranks in distributed runs.
Load a fully-assembled chunk cache if it exists.
In distributed mode, only rank0 loads the cache to avoid redundant IO.
Parameters:
Path to cached chunk file.
Returns: np.ndarray | None
Cached embeddings array, or None if cache doesn’t exist.
Load dataset and corpus from the input QA file.
Uses load_datasets() from retrieval_dataset.py to load the questions dataset and corpus dictionary. Validates that only a single corpus is referenced.
Load query and document embeddings from cache.
Returns: Tuple[np.ndarray | None, np.ndarray | None]
Tuple of (query_embeddings, document_embeddings), or (None, None) if not found.
Mine hard negatives for each query.
This implementation uses the following key behaviors:
- Deduplicates positive indices before masking (avoids double-masking)
- Preserves original order of positive scores (matches pos_doc order in input)
- Uses vectorized batch-level margin filtering for efficiency
- Uses batch-level topk for efficiency
Parameters:
Query embeddings [num_queries, embedding_dim].
Document embeddings [num_docs, embedding_dim].
List of positive document indices for each query.
Number of queries to process per batch.
Number of hard negatives to mine per query.
Optional margin for filtering false negatives.
“perc” (percentage) or “abs” (absolute).
Returns: Tuple[List[List[int]], List[List[float]], List[List[float]]]
Tuple of:
- neg_indices: List of hard negative indices per query
- neg_scores: Similarity scores for each hard negative
- pos_scores: Similarity scores for each positive document
Extract query texts and document indices from questions dataset.
Iterates through the questions dataset and extracts:
- Query texts (with EMPTY_QUESTION placeholder for empty queries)
- Question IDs
- Corpus IDs
- Positive document indices (mapped via doc_to_idx)
- Supplied negative document indices (if use_negatives_from_file is True)
Print mining configuration summary.
Save query and document embeddings to cache.
Parameters:
Query embeddings array.
Document embeddings array.
Synchronize all distributed ranks with a barrier.
Handles device-specific barrier calls for CUDA vs CPU.
Unload the model from GPU memory after embedding generation.
This frees up GPU memory for the mining phase, which only operates on embeddings and doesn’t need the model parameters.
Model metadata (pooling, l2_normalize) is extracted before unloading to ensure it’s available for output generation.
Validate required mining parameters.
Raises:
ValueError: If any required parameter is missing or invalid.
Write the output JSON file with mined hard negatives.
The output format:
- Preserves all top-level keys from input (corpus, etc.)
- Adds mining metadata section with parameters used
- Replaces neg_doc with newly mined negatives
- Adds similarity scores to all documents (pos_doc and neg_doc)
- Removes legacy score fields (pos_score, neg_scores) if present
Run the hard negative mining pipeline.
Generates query and document embeddings, mines hard negatives using similarity scores with margin filtering, and writes the output file.
Build all components needed for hard negative mining.
Bases: BaseRecipe
Recipe for training encoder models with contrastive learning.
Build optimizer parameter groups for trainable model parameters.
Return the embedding tensor used for validation scoring from a forward output.
The base bi-encoder forward returns an embedding tensor directly. Subclasses whose
forward returns a richer structure (e.g. the distillation student, which returns
(pooled, projected, intermediate_outputs)) should override this to select the
tensor to score with.
Forward and backward pass for a single micro-batch.
Run one optimization step with gradient accumulation.
Run validation for one epoch and compute loss, accuracy@1, and MRR.
Validate recipe-specific model settings before constructing the optimizer.
Parameters:
Constructed retrieval model, including infrastructure wrappers.
Run the training loop over all epochs and batches.
Build all components needed for training/validation/logging/checkpointing.
Bases: TrainBiEncoderRecipe
Forward and backward pass for a single micro-batch.
Run validation for one epoch and compute loss, accuracy@1, and MRR.
Validate the effective temperature applied by the constructed model.