ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelRecipesRetrievalnemo_automodel.recipes.retrieval.mine_hard_negatives

nemo_automodel.recipes.retrieval.mine_hard_negatives

View as Markdown

Hard negative mining recipe for encoder models.

Module Contents

Classes

NameDescription
MineHardNegativesRecipeRecipe for mining hard negatives for encoder training.

Functions

NameDescription
_compute_rank_partitionCompute contiguous partition boundaries for a given rank.
_concatenate_embedding_shardsConcatenate embedding shards while normalizing zero-row metadata shards.
_load_npz_arrayLoad numpy array from NPZ archive.
_save_npz_arrayAtomically save an embeddings array to an NPZ archive.
_validate_shard_shapeValidate that a shard has the expected number of items.
build_distributedBuild and initialize distributed resources.

Data

CORPUS_CHUNKS_DIR

DOCUMENT_EMBEDDINGS_FNAME

EMPTY_QUESTION

MINING_DEFAULTS

MULTIMODAL_SCRATCH_DIR

QUERY_EMBEDDINGS_FNAME

QUERY_SHARDS_DIR

TOPK_BUFFER_MULTIPLIER

logger

API

class nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe(
cfg
)

Recipe for mining hard negatives for encoder training.

This class orchestrates hard negative mining, from setup to mining execution. Hard negatives are documents that are semantically similar to the query but are not relevant, making them valuable for training more discriminative models.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._build_document_mappings()

Build bidirectional mappings between document IDs and indices.

Creates doc_to_idx and idx_to_doc dictionaries for efficient lookup during the mining process. Documents are sorted by ID for deterministic ordering across runs.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._build_negative_docs_by_question_id() -> typing.Dict[str, typing.List[typing.Dict[str, typing.Any]]]

Build mapping from question_id to mined negative documents with scores.

If use_negatives_from_file is True, supplied negatives are prepended to mined negatives (with score=-1 since we don’t compute their scores).

Returns: Dict[str, List[Dict[str, Any]]]

Dict mapping question_id to list of {“id”: doc_id, “score”: score} dicts.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._build_positive_scores_by_question_id() -> typing.Dict[str, typing.List[float]]

Build mapping from question_id to positive document scores.

Scores are in the same order as pos_doc in the original data.

Returns: Dict[str, List[float]]

Dict mapping question_id to list of positive scores.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._configure_tokenizer()

Load and configure tokenizer with appropriate settings.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_all_documents() -> numpy.ndarray

Encode all documents in corpus, chunk by chunk.

Returns: np.ndarray

numpy array of document embeddings [num_docs, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_chunk_distributed(
doc_indices: list[int],
cache_path: pathlib.Path
) -> numpy.ndarray

Encode a chunk of documents in distributed mode.

Shards the documents within the chunk across ranks, encodes each shard, and assembles on rank0.

Parameters:

doc_indices
list[int]

Corpus indices to encode.

cache_path
Path

Path for caching the assembled chunk.

Returns: np.ndarray

Assembled document embeddings for this chunk (rank0 only).

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_chunk_local(
doc_indices: list[int],
cache_path: pathlib.Path | None
) -> numpy.ndarray

Encode a chunk of documents locally (single-process).

Parameters:

doc_indices
list[int]

Corpus indices to encode.

cache_path
Path | None

Optional path for caching.

Returns: np.ndarray

Document embeddings for this chunk.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_document_indices(
doc_indices: list[int]
) -> numpy.ndarray

Fetch and encode corpus documents in bounded batches and stable order.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_documents_chunk(
doc_indices: typing.List[int],
cache_path: pathlib.Path | None = None
) -> numpy.ndarray

Encode a chunk of documents into embeddings.

Parameters:

doc_indices
List[int]

List of document indices to encode.

cache_path
Path | NoneDefaults to None

Optional path to save/load chunk cache.

Returns: np.ndarray

numpy array of document embeddings [num_docs, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_queries() -> numpy.ndarray

Encode all queries into embeddings.

Uses self.query_prefix and self.query_max_length from mining config.

Returns: np.ndarray

numpy array of query embeddings [num_queries, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_queries_sharded() -> numpy.ndarray

Encode queries sharded across ranks and assemble on rank0.

This is useful when the number of queries is large (e.g., 100k+), and we want to utilize multiple GPUs for query embedding generation without sharding mining/scoring.

Requires cache_embeddings_dir so ranks can write shard files and rank0 can assemble.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._encode_texts(
texts: typing.List[str],
batch_size: int,
max_length: int,
prefix: str = ''
) -> numpy.ndarray

Encode texts into embeddings.

Parameters:

texts
List[str]

List of text strings to encode.

batch_size
int

Batch size for encoding.

max_length
int

Maximum sequence length for tokenization.

prefix
strDefaults to ''

Optional prefix to prepend to each text.

Returns: np.ndarray

numpy array of embeddings [num_texts, embedding_dim].

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._extract_mining_params()

Extract all mining parameters from configuration.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._generate_embeddings() -> typing.Tuple[numpy.ndarray, numpy.ndarray]

Generate embeddings for queries and documents.

Handles caching and orchestrates the encoding process.

Returns: Tuple[np.ndarray, np.ndarray]

Tuple of (query_embeddings, document_embeddings).

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._get_document_text(
doc: dict
) -> str

Extract text from document dict, checking common field names.

Parameters:

doc
dict

Document dictionary from corpus.

Returns: str

Document text string.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._get_mining_args_dict() -> typing.Dict[str, typing.Any]

Get dictionary of mining arguments for output metadata.

Returns: Dict[str, Any]

Dict containing all mining parameters for reproducibility.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._get_mining_param(
name,
default = None
)

Get mining parameter from config, with fallback to defaults.

Parameters:

name

Parameter name.

default
Defaults to None

Default value if not in config or MINING_DEFAULTS.

Returns:

Parameter value.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._has_full_embeddings_cache() -> bool

Check if the consolidated (rank0) embedding cache exists.

This is intentionally a lightweight existence check (no file reads), used to avoid redundant IO on non-main ranks in distributed runs.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._load_cached_chunk(
cache_path: pathlib.Path
) -> numpy.ndarray | None

Load a fully-assembled chunk cache if it exists.

In distributed mode, only rank0 loads the cache to avoid redundant IO.

Parameters:

cache_path
Path

Path to cached chunk file.

Returns: np.ndarray | None

Cached embeddings array, or None if cache doesn’t exist.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._load_data()

Load dataset and corpus from the input QA file.

Uses load_datasets() from retrieval_dataset.py to load the questions dataset and corpus dictionary. Validates that only a single corpus is referenced.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._load_embeddings_from_cache() -> typing.Tuple[numpy.ndarray | None, numpy.ndarray | None]

Load query and document embeddings from cache.

Returns: Tuple[np.ndarray | None, np.ndarray | None]

Tuple of (query_embeddings, document_embeddings), or (None, None) if not found.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._mine_hard_negatives(
query_embeddings: numpy.ndarray,
document_embeddings: numpy.ndarray,
pos_doc_indices: typing.List[typing.List[int]],
batch_size: int,
num_negs: int,
hard_neg_margin: float | None = None,
hard_neg_margin_type: str | None = None
) -> typing.Tuple[typing.List[typing.List[int]], typing.List[typing.List[float]], typing.List[typing.List[float]]]

Mine hard negatives for each query.

This implementation uses the following key behaviors:

  • Deduplicates positive indices before masking (avoids double-masking)
  • Preserves original order of positive scores (matches pos_doc order in input)
  • Uses vectorized batch-level margin filtering for efficiency
  • Uses batch-level topk for efficiency

Parameters:

query_embeddings
np.ndarray

Query embeddings [num_queries, embedding_dim].

document_embeddings
np.ndarray

Document embeddings [num_docs, embedding_dim].

pos_doc_indices
List[List[int]]

List of positive document indices for each query.

batch_size
int

Number of queries to process per batch.

num_negs
int

Number of hard negatives to mine per query.

hard_neg_margin
float | NoneDefaults to None

Optional margin for filtering false negatives.

hard_neg_margin_type
str | NoneDefaults to None

“perc” (percentage) or “abs” (absolute).

Returns: Tuple[List[List[int]], List[List[float]], List[List[float]]]

Tuple of:

  • neg_indices: List of hard negative indices per query
  • neg_scores: Similarity scores for each hard negative
  • pos_scores: Similarity scores for each positive document
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._prepare_data()

Extract query texts and document indices from questions dataset.

Iterates through the questions dataset and extracts:

  • Query texts (with EMPTY_QUESTION placeholder for empty queries)
  • Question IDs
  • Corpus IDs
  • Positive document indices (mapped via doc_to_idx)
  • Supplied negative document indices (if use_negatives_from_file is True)
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._print_configuration()

Print mining configuration summary.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._save_embeddings_to_cache(
query_embeddings: numpy.ndarray,
document_embeddings: numpy.ndarray
) -> None

Save query and document embeddings to cache.

Parameters:

query_embeddings
np.ndarray

Query embeddings array.

document_embeddings
np.ndarray

Document embeddings array.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._synchronize_ranks()

Synchronize all distributed ranks with a barrier.

Handles device-specific barrier calls for CUDA vs CPU.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._unload_model()

Unload the model from GPU memory after embedding generation.

This frees up GPU memory for the mining phase, which only operates on embeddings and doesn’t need the model parameters.

Model metadata (pooling, l2_normalize) is extracted before unloading to ensure it’s available for output generation.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._validate_mining_params()

Validate required mining parameters.

Raises:

  • ValueError: If any required parameter is missing or invalid.
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe._write_output() -> None

Write the output JSON file with mined hard negatives.

The output format:

  • Preserves all top-level keys from input (corpus, etc.)
  • Adds mining metadata section with parameters used
  • Replaces neg_doc with newly mined negatives
  • Adds similarity scores to all documents (pos_doc and neg_doc)
  • Removes legacy score fields (pos_score, neg_scores) if present
nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe.run()

Run the hard negative mining pipeline.

Generates query and document embeddings, mines hard negatives using similarity scores with margin filtering, and writes the output file.

nemo_automodel.recipes.retrieval.mine_hard_negatives.MineHardNegativesRecipe.setup()

Build all components needed for hard negative mining.

nemo_automodel.recipes.retrieval.mine_hard_negatives._compute_rank_partition(
total_size: int,
world_size: int,
rank: int
) -> typing.Tuple[int, int]

Compute contiguous partition boundaries for a given rank.

Distributes total_size items across world_size ranks as evenly as possible, with remainder items distributed to lower ranks.

Parameters:

total_size
int

Total number of items to partition.

world_size
int

Number of ranks.

rank
int

Current rank (0-indexed).

Returns: Tuple[int, int]

Tuple of (start_idx, end_idx) for this rank’s partition.

nemo_automodel.recipes.retrieval.mine_hard_negatives._concatenate_embedding_shards(
parts: list[numpy.ndarray]
) -> numpy.ndarray

Concatenate embedding shards while normalizing zero-row metadata shards.

nemo_automodel.recipes.retrieval.mine_hard_negatives._load_npz_array(
path: pathlib.Path
) -> numpy.ndarray

Load numpy array from NPZ archive.

Parameters:

path
Path

Path to NPZ file.

Returns: np.ndarray

Loaded numpy array.

nemo_automodel.recipes.retrieval.mine_hard_negatives._save_npz_array(
path: pathlib.Path,
embeddings: numpy.ndarray
) -> None

Atomically save an embeddings array to an NPZ archive.

nemo_automodel.recipes.retrieval.mine_hard_negatives._validate_shard_shape(
shard_path: pathlib.Path,
expected_size: int,
actual_size: int
) -> None

Validate that a shard has the expected number of items.

Parameters:

shard_path
Path

Path to the shard file (for error reporting).

expected_size
int

Expected number of items.

actual_size
int

Actual number of items.

Raises:

  • ValueError: If sizes don’t match.
nemo_automodel.recipes.retrieval.mine_hard_negatives.build_distributed(
cfg_dist: typing.Dict[str, typing.Any]

Build and initialize distributed resources.

Parameters:

cfg_dist
Dict[str, Any]

Configuration for distributed environment.

Returns: DistInfo

Distributed environment information from initialize_distributed.

nemo_automodel.recipes.retrieval.mine_hard_negatives.CORPUS_CHUNKS_DIR = 'corpus_chunks'
nemo_automodel.recipes.retrieval.mine_hard_negatives.DOCUMENT_EMBEDDINGS_FNAME = 'passage_embeddings.npz'
nemo_automodel.recipes.retrieval.mine_hard_negatives.EMPTY_QUESTION = '##### keep empty questions #####'
nemo_automodel.recipes.retrieval.mine_hard_negatives.MINING_DEFAULTS = {'hard_negatives_to_mine': 20, 'hard_neg_margin': 0.95, 'hard_neg_margin_type': ...
nemo_automodel.recipes.retrieval.mine_hard_negatives.MULTIMODAL_SCRATCH_DIR = 'multimodal_scratch'
nemo_automodel.recipes.retrieval.mine_hard_negatives.QUERY_EMBEDDINGS_FNAME = 'query_embeddings.npz'
nemo_automodel.recipes.retrieval.mine_hard_negatives.QUERY_SHARDS_DIR = 'query_shards'
nemo_automodel.recipes.retrieval.mine_hard_negatives.TOPK_BUFFER_MULTIPLIER = 2
nemo_automodel.recipes.retrieval.mine_hard_negatives.logger = logging.getLogger(__name__)