Curate TextProcess DataEmbeddings

vLLM Embedder

View as Markdown

Generate text embeddings using vLLM’s optimized inference engine. The VLLMEmbeddingModelStage provides high-throughput embedding generation, particularly for large embedding models where vLLM’s batching and GPU memory management provide significant performance advantages over Sentence Transformers.

Installation: The vLLM embedder is included in the text_cuda12 installation. Install it with:

uv pip install --override text_cuda12-overrides.txt "nemo-curator[text_cuda12]"

First download the override file and add the CUDA indexes as shown in the installation guide. Standard pip is not supported for this extra.

How It Works

VLLMEmbeddingModelStage is a single-stage embedder that handles both tokenization and embedding generation within one stage. Unlike EmbeddingCreatorStage (which splits tokenization and model inference into separate stages), the vLLM embedder delegates all GPU operations to vLLM’s inference engine.

Key features:

  • Optional pretokenization: When pretokenize=True, the stage tokenizes text on CPU before passing tokens to vLLM, reducing GPU idle time and improving throughput
  • vLLM-managed batching: Leverages vLLM’s built-in request scheduling for optimal GPU utilization
  • Model download caching: Automatically downloads and caches models from Hugging Face Hub
  • Character truncation: Optional max_chars parameter to limit input length before tokenization

Quick Start

from nemo_curator.backends.xenna import XennaExecutor
from nemo_curator.stages.text.embedders.vllm import VLLMEmbeddingModelStage
from nemo_curator.pipeline import Pipeline
from nemo_curator.stages.text.io.reader import ParquetReader
from nemo_curator.stages.text.io.writer import ParquetWriter
pipeline = Pipeline(
name="vllm_embeddings",
stages=[
ParquetReader(file_paths="input_data/", files_per_partition=1, fields=["text"]),
VLLMEmbeddingModelStage(
model_identifier="google/embeddinggemma-300m",
text_field="text",
embedding_field="embeddings",
metadata_fields=["text"],
),
ParquetWriter(path="output/", fields=["text", "embeddings"]),
],
)
executor = XennaExecutor()
pipeline.run(executor)

Configuration

Parameters

ParameterTypeDefaultDescription
model_identifierstrRequiredHugging Face model name or path for the embedding model
vllm_init_kwargsdictNoneAdditional keyword arguments passed to vllm.LLM() for engine configuration
text_fieldstr"text"Name of the input text column in the data
pretokenizeboolTrueTokenize text on CPU before passing to vLLM. Whether this improves throughput is model-dependent
embedding_fieldstr"embeddings"Name of the output embedding column
max_charsintNoneMaximum characters per document (truncates before tokenization)
model_inference_batch_sizeint | None8192Maximum number of texts sent to vLLM in each embedding call. Set to 0 or None to process the entire input task in one call. Negative values are rejected.
cache_dirstrNoneDirectory for caching downloaded model files
hf_tokenstrNoneHugging Face token for accessing gated models
verboseboolFalseEnable verbose logging and progress bars

Use a smaller model_inference_batch_size to reduce memory pressure while preparing each embedding call. The setting does not impose a hard GPU-memory limit or change the size of the input task.

vLLM Engine Options

Pass additional vLLM configuration through vllm_init_kwargs:

VLLMEmbeddingModelStage(
model_identifier="google/embeddinggemma-300m",
pretokenize=True,
vllm_init_kwargs={
"enforce_eager": True, # Disable CUDA graph for debugging
"tensor_parallel_size": 2, # Distribute across 2 GPUs
"gpu_memory_utilization": 0.9,
"max_model_len": 512,
},
)

Default vLLM settings applied by the stage (can be overridden):

  • enforce_eager=False — Uses CUDA graphs for faster inference
  • runner="pooling" — Configures vLLM for embedding (pooling) tasks
  • model_impl="vllm" — Uses vLLM’s native model implementation
  • disable_log_stats=True — Suppresses stats logging when verbose=False

Pretokenization

When pretokenize=True, the stage:

  1. Loads a Hugging Face Auto Tokenizer for the specified model
  2. Tokenizes the input text batch on CPU with truncation to max_model_len
  3. Passes token IDs directly to vLLM using TokensPrompt

The stage defaults to pretokenize=True because benchmarks show that CPU tokenization can improve per-task throughput by reducing GPU idle time. The benefit is model-dependent, so set pretokenize=False to let vLLM tokenize internally when that performs better for your workload. SemanticDeduplicationStage currently overrides this stage default with embedding_pretokenize=False; set that option to True to enable pretokenization in semantic deduplication pipelines.

# Direct text mode (opt out of the default CPU pretokenization)
VLLMEmbeddingModelStage(
model_identifier="google/embeddinggemma-300m",
pretokenize=False, # vLLM handles tokenization internally
)
# Pretokenize mode (the default)
VLLMEmbeddingModelStage(
model_identifier="intfloat/e5-large-v2",
pretokenize=True, # Tokenize on CPU, embed on GPU
)

Resources

The VLLMEmbeddingModelStage requests 1 CPU and 1 GPU per worker by default. For multi-GPU models, configure tensor_parallel_size in vllm_init_kwargs.