nemo_automodel.recipes.retrieval.mine_hard_negatives
nemo_automodel.recipes.retrieval.mine_hard_negatives
Hard negative mining recipe for encoder models.
Module Contents
Classes
Functions
Data
API
Recipe for mining hard negatives for encoder training.
This class orchestrates hard negative mining, from setup to mining execution. Hard negatives are documents that are semantically similar to the query but are not relevant, making them valuable for training more discriminative models.
Build bidirectional mappings between document IDs and indices.
Creates doc_to_idx and idx_to_doc dictionaries for efficient lookup during the mining process. Documents are sorted by ID for deterministic ordering across runs.
Build mapping from question_id to mined negative documents with scores.
If use_negatives_from_file is True, supplied negatives are prepended to mined negatives (with score=-1 since we don’t compute their scores).
Returns: Dict[str, List[Dict[str, Any]]]
Dict mapping question_id to list of {“id”: doc_id, “score”: score} dicts.
Build mapping from question_id to positive document scores.
Scores are in the same order as pos_doc in the original data.
Returns: Dict[str, List[float]]
Dict mapping question_id to list of positive scores.
Load and configure tokenizer with appropriate settings.
Encode all documents in corpus, chunk by chunk.
Returns: np.ndarray
numpy array of document embeddings [num_docs, embedding_dim].
Encode a chunk of documents in distributed mode.
Shards the documents within the chunk across ranks, encodes each shard, and assembles on rank0.
Parameters:
Corpus indices to encode.
Path for caching the assembled chunk.
Returns: np.ndarray
Assembled document embeddings for this chunk (rank0 only).
Encode a chunk of documents locally (single-process).
Parameters:
Corpus indices to encode.
Optional path for caching.
Returns: np.ndarray
Document embeddings for this chunk.
Fetch and encode corpus documents in bounded batches and stable order.
Encode a chunk of documents into embeddings.
Parameters:
List of document indices to encode.
Optional path to save/load chunk cache.
Returns: np.ndarray
numpy array of document embeddings [num_docs, embedding_dim].
Encode all queries into embeddings.
Uses self.query_prefix and self.query_max_length from mining config.
Returns: np.ndarray
numpy array of query embeddings [num_queries, embedding_dim].
Encode queries sharded across ranks and assemble on rank0.
This is useful when the number of queries is large (e.g., 100k+), and we want to utilize multiple GPUs for query embedding generation without sharding mining/scoring.
Requires cache_embeddings_dir so ranks can write shard files and rank0 can assemble.
Encode texts into embeddings.
Parameters:
List of text strings to encode.
Batch size for encoding.
Maximum sequence length for tokenization.
Optional prefix to prepend to each text.
Returns: np.ndarray
numpy array of embeddings [num_texts, embedding_dim].
Extract all mining parameters from configuration.
Generate embeddings for queries and documents.
Handles caching and orchestrates the encoding process.
Returns: Tuple[np.ndarray, np.ndarray]
Tuple of (query_embeddings, document_embeddings).
Extract text from document dict, checking common field names.
Parameters:
Document dictionary from corpus.
Returns: str
Document text string.
Get dictionary of mining arguments for output metadata.
Returns: Dict[str, Any]
Dict containing all mining parameters for reproducibility.
Get mining parameter from config, with fallback to defaults.
Parameters:
Parameter name.
Default value if not in config or MINING_DEFAULTS.
Returns:
Parameter value.
Check if the consolidated (rank0) embedding cache exists.
This is intentionally a lightweight existence check (no file reads), used to avoid redundant IO on non-main ranks in distributed runs.
Load a fully-assembled chunk cache if it exists.
In distributed mode, only rank0 loads the cache to avoid redundant IO.
Parameters:
Path to cached chunk file.
Returns: np.ndarray | None
Cached embeddings array, or None if cache doesn’t exist.
Load dataset and corpus from the input QA file.
Uses load_datasets() from retrieval_dataset.py to load the questions dataset and corpus dictionary. Validates that only a single corpus is referenced.
Load query and document embeddings from cache.
Returns: Tuple[np.ndarray | None, np.ndarray | None]
Tuple of (query_embeddings, document_embeddings), or (None, None) if not found.
Mine hard negatives for each query.
This implementation uses the following key behaviors:
- Deduplicates positive indices before masking (avoids double-masking)
- Preserves original order of positive scores (matches pos_doc order in input)
- Uses vectorized batch-level margin filtering for efficiency
- Uses batch-level topk for efficiency
Parameters:
Query embeddings [num_queries, embedding_dim].
Document embeddings [num_docs, embedding_dim].
List of positive document indices for each query.
Number of queries to process per batch.
Number of hard negatives to mine per query.
Optional margin for filtering false negatives.
“perc” (percentage) or “abs” (absolute).
Returns: Tuple[List[List[int]], List[List[float]], List[List[float]]]
Tuple of:
- neg_indices: List of hard negative indices per query
- neg_scores: Similarity scores for each hard negative
- pos_scores: Similarity scores for each positive document
Extract query texts and document indices from questions dataset.
Iterates through the questions dataset and extracts:
- Query texts (with EMPTY_QUESTION placeholder for empty queries)
- Question IDs
- Corpus IDs
- Positive document indices (mapped via doc_to_idx)
- Supplied negative document indices (if use_negatives_from_file is True)
Print mining configuration summary.
Save query and document embeddings to cache.
Parameters:
Query embeddings array.
Document embeddings array.
Synchronize all distributed ranks with a barrier.
Handles device-specific barrier calls for CUDA vs CPU.
Unload the model from GPU memory after embedding generation.
This frees up GPU memory for the mining phase, which only operates on embeddings and doesn’t need the model parameters.
Model metadata (pooling, l2_normalize) is extracted before unloading to ensure it’s available for output generation.
Validate required mining parameters.
Raises:
ValueError: If any required parameter is missing or invalid.
Write the output JSON file with mined hard negatives.
The output format:
- Preserves all top-level keys from input (corpus, etc.)
- Adds mining metadata section with parameters used
- Replaces neg_doc with newly mined negatives
- Adds similarity scores to all documents (pos_doc and neg_doc)
- Removes legacy score fields (pos_score, neg_scores) if present
Run the hard negative mining pipeline.
Generates query and document embeddings, mines hard negatives using similarity scores with margin filtering, and writes the output file.
Build all components needed for hard negative mining.
Compute contiguous partition boundaries for a given rank.
Distributes total_size items across world_size ranks as evenly as possible,
with remainder items distributed to lower ranks.
Parameters:
Total number of items to partition.
Number of ranks.
Current rank (0-indexed).
Returns: Tuple[int, int]
Tuple of (start_idx, end_idx) for this rank’s partition.
Concatenate embedding shards while normalizing zero-row metadata shards.
Load numpy array from NPZ archive.
Parameters:
Path to NPZ file.
Returns: np.ndarray
Loaded numpy array.
Atomically save an embeddings array to an NPZ archive.
Validate that a shard has the expected number of items.
Parameters:
Path to the shard file (for error reporting).
Expected number of items.
Actual number of items.
Raises:
ValueError: If sizes don’t match.
Build and initialize distributed resources.
Parameters:
Configuration for distributed environment.
Returns: DistInfo
Distributed environment information from initialize_distributed.