Semantic Deduplication
Detect and remove semantically redundant data from your large text datasets using NeMo Curator.
Unlike exact or fuzzy deduplication, which focus on textual similarity, semantic deduplication leverages the meaning of content to identify duplicates. This approach can significantly reduce dataset size while maintaining or even improving model performance.
The technique uses embeddings to identify “semantic duplicates” - content pairs that convey similar meaning despite using different words.
GPU Acceleration: Semantic deduplication requires GPU acceleration for both embedding generation and clustering operations. This method uses cuDF for GPU-accelerated dataframe operations and PyTorch models on GPU for optimal performance.
How It Works
Semantic deduplication identifies meaning-based duplicates using embeddings:
- Generates embeddings for each document using transformer models
- Clusters embeddings using K-means
- Computes pairwise cosine similarities within clusters
- Identifies semantic duplicates based on similarity threshold
- Removes duplicates, keeping one representative per group
Based on SemDeDup: Data-efficient learning at web-scale through semantic deduplication by Abbas et al.
Before You Start
Prerequisites:
- GPU acceleration (required for embedding generation and clustering)
- Stable document identifiers for removal (either existing IDs or IDs managed by the workflow and removal stages)
- Access to the gated default model: accept the EmbeddingGemma usage license, then authenticate with a Hugging Face read token through
HF_TOKENor thehf_tokenparameter
Running in Docker: When running semantic deduplication inside the NeMo Curator container, ensure the container is started with --gpus all so that CUDA GPUs are available. Without this flag, you will see RuntimeError: No CUDA GPUs are available. Also activate the virtual environment with source /opt/venv/env.sh after entering the container.
Quick Start
Get started with semantic deduplication using the following example of identifying duplicates, then remove them in one step:
Configuration
Configure semantic deduplication using these key parameters:
Step-by-Step Workflow
For fine-grained control, break semantic deduplication into separate stages:
This approach enables analysis of intermediate results and parameter tuning.
Disable Weighted K-Means Task Assignment
By default, the Ray Actor Pool executor uses each input partition’s file size to balance K-means work across RAFT actors. To restore equal task-count assignment, override the decomposed KMeansStage with USE_TASK_WEIGHTS=False:
This override is available when using KMeansStage directly. The high-level semantic deduplication workflows construct their K-means stage internally and keep weighted assignment enabled. Disabling weights is mainly useful for comparisons or troubleshooting; it can leave actors with uneven amounts of input data.
Comparison with Other Deduplication Methods
Compare semantic deduplication with other methods:
Key Parameters
For a remote output_path, configure storage_options in both cache_kwargs and write_kwargs, even when cache_path is local. The options in cache_kwargs must also allow access to cache_path if it is remote.
The same discovery rules apply to TextSemanticDeduplicationWorkflow, the pre-computed-embedding SemanticDeduplicationWorkflow, and TextDuplicatesRemovalWorkflow. The extension filter does not infer the reader format: keep input_filetype="parquet" when reading a custom Parquet suffix such as .pq. See Input File Discovery for examples and recursion details.
TextSemanticDeduplicationWorkflow always writes Parquet embeddings before clustering, regardless of
the original input_filetype. The lower-level SemanticDeduplicationWorkflow and KMeansStage accept
precomputed Parquet or JSONL embeddings and use the parameter name fit_data_fraction.
Control Pairwise memory and precision
Pairwise comparison keeps a reusable N x B similarity workspace, where N is the cluster size and B is pairwise_batch_size. The default is 1024; a smaller positive value lowers peak memory at the cost of more matrix multiplications. A larger value also computes more self/later-row products before masking them, so the largest batch that fits is not necessarily the fastest. Only earlier-ranked rows are valid neighbors, and equal similarity scores deterministically choose the earliest-ranked row.
How Pairwise chooses duplicates
Input order is the preservation preference: A is preferred over B, B over C, and so on. Consider these normalized embeddings:
Because they have unit length, X @ X.T is their cosine-similarity matrix:
Only candidates to the left of the query’s diagonal are eligible. Pairwise therefore selects B -> A (0.96), C -> B (0.28), D -> A (0.00), and E -> C (0.96). With eps=0.1, the duplicate threshold is 1 - eps = 0.9, so the next stage removes query rows B and E while preserving A, C, and D. An earlier-ranked candidate remains eligible even if that candidate is itself later marked for removal.
KMeans fit, prediction, and distance calculations remain FP32. By default, KMeans casts normalized embeddings to FP16 and stores their bit patterns in uint16 list leaves because cuDF does not support numeric FP16 columns. Set kmeans_embedding_output_dtype="float32" to retain FP32 embeddings. The uint16 representation does not guarantee half-sized Parquet files because Parquet stores uint16 values using an INT32 physical type.
Pairwise infers storage precision from the embedding list leaves: uint16 leaves are decoded as FP16 bit patterns, FP32 leaves keep their precision, and legacy FP64 leaves are normalized to FP32. With pairwise_compute_dtype="auto", FP16 embeddings use FP16 multiplication and FP32 embeddings use FP32. FP32 embeddings can instead use "float16" multiplication; upcasting stored FP16 embeddings is rejected because it cannot restore the discarded precision.
FP16 scores remain FP16 through multiplication and reduction, then are promoted to FP32 because cuDF does not support an FP16 numeric output column. Pairwise records cumulative footer scan, read, rank, precision conversion, compute, and write time.
Use FP32 storage and compute when exact or near-exact similarity thresholds make FP16 rounding material, or when the selected nearest-neighbor ID is part of the required output contract.
Control K-means fitting memory
For the lower-level workflow and stage, the behavior of fit_data_fraction=None depends on the
embedding input format:
- Parquet: Selects as many complete files as fit the actor’s live GPU-memory budget. Only one contiguous float32 embedding array remains resident during fitting; each input frame is released after its embeddings are copied into that array.
- JSONL: Fits all input files in one pass because row and embedding counts are not available without reading the data.
Parquet is recommended for large semantic-deduplication workloads. Its footer metadata sizes the
fit array, while cuDF’s chunked reader bounds fit-pass memory and output column sizes. After fitting, Parquet
rereads all files in bounded frames to predict and write every row, including when
fit_data_fraction=1.0 or automatic sizing selects every file. The fit array is released before
prediction and writing, so it does not compete with their temporary allocations. Prediction groups
respect cuDF’s embedding child-column limit, but this is not a GPU-memory limit: a large group or
metadata-heavy input can still exhaust memory during reading, prediction, or writing.
An explicit JSONL fit fraction also rereads the input, but loads all files together for prediction.
Set a fraction in (0, 1] to choose the number of sampled files explicitly. Both formats sample
round(fit_data_fraction * number_of_actor_files) complete files per actor, with a minimum of one,
but their I/O differs:
- Parquet: Exact footer counts size one contiguous float32 fit array. cuDF reads the sampled files in chunks that share the remaining GPU memory between output and decompression. The sampled frames are released during the first pass; the second pass rereads all files to predict and write every row.
- JSONL: The first pass reads the sampled files and fits K-means. The second pass reads all files, including the sampled files, to predict and write every row.
Because sampling is file-based rather than row-based, the realized row fraction can differ from the configured value when file sizes vary. Every actor contributes at least one file, so very small fractions can also sample more data than expected on highly parallel runs. Choose a fraction that leaves enough representative rows for at least n_clusters centroids, and use random_state or kmeans_random_state to make file selection repeatable.
All input rows receive cluster assignments; the fraction affects only centroid fitting. For Parquet input, both fitting and prediction use exact footer element counts to keep each read below cuDF’s column-size limit.
Text workflow
Embeddings workflow
KMeansStage with saved centroids
KMeansStage.cache_path controls only centroid persistence. When it is set, actor 0 writes kmeans_centroids.npy after fitting; when it is None, centroids are not saved. This differs from SemanticDeduplicationWorkflow.cache_path, which stores the workflow’s K-means and pairwise intermediate results. The workflow intentionally does not forward its cache path to KMeansStage.cache_path, so use KMeansStage directly when you need the centroid array.
Similarity Threshold
Control deduplication aggressiveness with eps:
- Lower values (such as 0.001): More strict, less deduplication, higher confidence
- Higher values (such as 0.1): Less strict, more aggressive deduplication
Experiment with different values to balance data reduction and dataset diversity.
Embedding Models
Embedding generation uses vLLM as the inference backend. The default model is google/embeddinggemma-300m.
Default (vLLM):
Custom model with vLLM options:
vLLM Embedder (recommended for large models):
For large embedding models, you can generate embeddings separately using VLLMEmbeddingModelStage before running the deduplication workflow. This provides better GPU utilization and throughput for models with 500M+ parameters. See vLLM Embedder for details.
Generate embeddings with VLLMEmbeddingModelStage using the vLLM Embedder pipeline, then pass the output to SemanticDeduplicationWorkflow:
When choosing a model:
- Use models that support vLLM pooling (embedding) mode
- Choose models appropriate for your language or domain
- Prefer models trained for sentence embeddings (for example, EmbeddingGemma, E5, BGE, or SBERT)
- Use
embedding_pretokenize=Truefor models that benefit from explicit tokenization control - Pass additional vLLM configuration through
embedding_vllm_init_kwargs - For more control over the embedding process, consider using VLLMEmbeddingModelStage separately
Advanced Configuration
GPU write tuning:
NeMo Curator defaults KVIKIO_AUTO_DIRECT_IO_WRITE=0 because KvikIO direct
writes can reduce throughput on distributed filesystems such as Lustre. If the
workflow writes to a local SSD, setting the variable to 1 may improve
throughput:
Set the variable before importing NeMo Curator or starting the job, ensure every distributed worker inherits it, and benchmark both values against the target filesystem. The setting affects the workflow’s GPU-backed cuDF intermediate and output writes, not general CPU-based writes. See Tune GPU Writes for the Target Filesystem for details.
Output Format
The semantic deduplication process produces the following directory structure in your configured cache_path:
File Formats
The workflow produces these output files:
-
Document Embeddings (
embeddings/*.parquet):- Contains document IDs and their vector embeddings
- Format: Parquet files with columns:
[id_column, embedding_column]
-
Cluster Assignments (
semantic_dedup/kmeans_results/):embs_by_nearest_center/: Parquet files containing cluster members- Format: Parquet files with columns:
[id_column, embedding_column, cluster_id]
A direct
KMeansStage(cache_path="kmeans_cache/")additionally writes the fitted cluster centers tokmeans_cache/kmeans_centroids.npy. The workflow wrappers do not save this file. -
Duplicate IDs (
output_path/duplicates/*.parquet):- IDs of documents identified as duplicates for removal
- Format: Parquet file with columns:
["id"] - Important: Contains only the IDs of documents to remove, not the full document content
- When
perform_removal=True, clean dataset is saved tooutput_path/deduplicated/
Performance Considerations
Performance characteristics:
- Computationally intensive, especially for large datasets
- GPU acceleration required for embedding generation and clustering
- Benefits often outweigh upfront cost (reduced training time, improved model performance)
GPU requirements:
- NVIDIA GPU with CUDA support
- Sufficient GPU memory (recommended: >8GB for medium datasets)
- RAPIDS libraries (cuDF) for GPU-accelerated dataframe operations
- CPU-only processing not supported
Performance tuning:
- Adjust
n_clustersbased on dataset size and available resources - Use batched cosine similarity to reduce memory requirements
- Consider distributed processing for very large datasets
For more details, see the SemDeDup paper by Abbas et al.
Advanced Configuration
ID Generator for large-scale operations:
Critical requirements:
- Use the same input configuration (file paths, partitioning) across all stages
- ID consistency maintained by hashing filenames in each task
- Mismatched partitioning causes ID lookup failures
Ray backend configuration:
Provides distributed processing, memory management, and fault tolerance.