> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.recipes.retrieval.mining_encoder

Checkpoint inference adapters for the hard-negative mining recipe.

## Module Contents

### Classes

| Name                                                                                                                    | Description                                                                       |
| ----------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- |
| [`CheckpointMiningEncoder`](#nemo_automodel-recipes-retrieval-mining_encoder-CheckpointMiningEncoder)                   | Encode mining queries and documents through the checkpoint's retrieval processor. |
| [`CheckpointMiningEncoderConfig`](#nemo_automodel-recipes-retrieval-mining_encoder-CheckpointMiningEncoderConfig)       | Build mining inference using the checkpoint's saved prompts and preprocessing.    |
| [`SentenceTransformerMiningEncoder`](#nemo_automodel-recipes-retrieval-mining_encoder-SentenceTransformerMiningEncoder) | Adapt Sentence Transformers query/document inference to the mining corpus.        |
| [`_RetrievalProcessor`](#nemo_automodel-recipes-retrieval-mining_encoder-_RetrievalProcessor)                           | -                                                                                 |

### Functions

| Name                                                                                      | Description                                                                   |
| ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- |
| [`_prepare_document`](#nemo_automodel-recipes-retrieval-mining_encoder-_prepare_document) | Use all available document content, normalizing absent images and blank text. |

### Data

[`logger`](#nemo_automodel-recipes-retrieval-mining_encoder-logger)

### API

```python
class nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoder(
    model: torch.nn.Module,
    processor: nemo_automodel.recipes.retrieval.mining_encoder._RetrievalProcessor,
    device: torch.device
)
```

Encode mining queries and documents through the checkpoint's retrieval processor.

**`l2_normalize`** `bool`

Whether the checkpoint normalizes embeddings.

---

**`pooling`** `str`

Pooling mode used by the checkpoint.

---

```python
nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoder._encode_batch(
    inputs: dict[str, typing.Any]
) -> numpy.ndarray
```

Return finite embeddings of shape \[batch, hidden] for one processor batch.

```python
nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoder.encode_documents(
    documents: list[dict[str, typing.Any]],
    batch_size: int
) -> numpy.ndarray
```

Encode corpus documents in input order with bounded processor batches.

```python
nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoder.encode_queries(
    queries: list[str],
    batch_size: int
) -> numpy.ndarray
```

Encode queries in input order with bounded processor batches.

```python
nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoder.release_model() -> None
```

Move the model to CPU and release it after embedding generation; safe to repeat.

```python
class nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoderConfig()
```

Dataclass

Build mining inference using the checkpoint's saved prompts and preprocessing.

```python
nemo_automodel.recipes.retrieval.mining_encoder.CheckpointMiningEncoderConfig.build(
    device: torch.device,
    model_name_or_path: str,
    trust_remote_code: bool = False,
    attn_implementation: str | None = None
) -> 'CheckpointMiningEncoder | SentenceTransformerMiningEncoder'
```

Build an encoder using Sentence Transformers metadata or the checkpoint processor.

**Parameters:**

**`device`** `torch.device`

Device on which model inputs and embeddings are computed.

---

**`model_name_or_path`** `str`

Local checkpoint directory or Hugging Face model ID.

---

**`trust_remote_code`** `bool` — default: False

Whether model loading may execute remote code.

---

**`attn_implementation`** `str | None` — default: None

Optional attention backend for model loading.

---

**Returns:** `'CheckpointMiningEncoder | SentenceTransformerMiningEncoder'`

A configured mining encoder.

```python
class nemo_automodel.recipes.retrieval.mining_encoder.SentenceTransformerMiningEncoder(
    model: typing.Any
)
```

Adapt Sentence Transformers query/document inference to the mining corpus.

**`l2_normalize`**

---

**`pooling`** `= 'avg' if pooling == 'mean' else pooling`

---

```python
nemo_automodel.recipes.retrieval.mining_encoder.SentenceTransformerMiningEncoder.encode_documents(
    documents: list[dict[str, typing.Any]],
    batch_size: int
) -> numpy.ndarray
```

Encode text, image, and image-text corpus records in input order.

```python
nemo_automodel.recipes.retrieval.mining_encoder.SentenceTransformerMiningEncoder.encode_queries(
    queries: list[str],
    batch_size: int
) -> numpy.ndarray
```

Encode query strings with the checkpoint's saved prompts and sequence limits.

```python
nemo_automodel.recipes.retrieval.mining_encoder.SentenceTransformerMiningEncoder.release_model() -> None
```

Move the model to CPU and release it after embedding generation; safe to repeat.

```python
class nemo_automodel.recipes.retrieval.mining_encoder._RetrievalProcessor()
```

Protocol

```python
nemo_automodel.recipes.retrieval.mining_encoder._RetrievalProcessor.process_documents(
    documents: list[dict[str, typing.Any]],
    return_tensors: typing.Literal['pt']
) -> dict[str, typing.Any]
```

```python
nemo_automodel.recipes.retrieval.mining_encoder._RetrievalProcessor.process_queries(
    queries: list[str],
    return_tensors: typing.Literal['pt']
) -> dict[str, typing.Any]
```

```python
nemo_automodel.recipes.retrieval.mining_encoder._prepare_document(
    document: dict[str, typing.Any]
) -> tuple[typing.Any, str]
```

Use all available document content, normalizing absent images and blank text.

```python
nemo_automodel.recipes.retrieval.mining_encoder.logger = logging.getLogger(__name__)
```