Language Identification
Large unlabeled text corpora often contain a variety of languages. NVIDIA NeMo Curator provides tools to accurately identify the language of each document, which is essential for language-specific curation tasks and building high-quality monolingual datasets.
How it Works
NeMo Curator’s language identification system works through a three-step process:
-
Text Preprocessing: For FastText classification, normalize input text by stripping whitespace and converting newlines to spaces.
-
FastText Language Detection: A FastText-compatible language identification model, such as
lid.176.binor GlotLID, analyzes the preprocessed text and returns:- A confidence score (0.0 to 1.0) indicating certainty of the prediction
- The model’s language label (for example,
en) or language-script label (for example,eng_Latn)
-
Filtering and Scoring: The pipeline filters documents based on a configurable confidence threshold (
min_langid_score) and stores both the confidence score and language code as metadata.
Language Detection Process
The FastTextLangId filter implements this workflow by:
- Loading the FastText language identification model on worker initialization
- Processing text through
model.predict()withk=1to get the top language prediction - Removing the
__label__prefix while preserving the model label’s original casing (for example,__label__enbecomesenand__label__eng_Latnbecomeseng_Latn) - Comparing confidence scores against the threshold to determine document retention
- Returning results as
[confidence_score, language_code]for downstream processing
The standard FastText model supports 176 languages. GlotLID can be used when broader language coverage or script identification is required.
Usage
The following example demonstrates how to create a language identification pipeline using Curator with distributed processing.
Python
Using GlotLID
Download the GlotLID FastText model from Hugging Face, then pass its local path to FastTextLangId in the same way as the standard FastText model:
Language matching is case-insensitive. A filter without an underscore matches the language portion of a GlotLID label, while a filter containing an underscore matches the complete language-script label.
Understanding Results
The language identification process adds a score field to each document batch:
-
languagefield: Contains the FastText language identification results as a string representation of a list with two elements (for backend compatibility):- Element 0: The confidence score (between 0 and 1)
- Element 1: The model label without the
__label__prefix and with its original casing preserved (for example,enoreng_Latn)
-
Task-based processing: Curator processes documents in batches (tasks), and results are available through the task’s Pandas DataFrame:
For quick exploratory inspection, converting a DocumentBatch to a Pandas DataFrame is fine. For performance and scalability, write transformations as ProcessingStages (or with the @processing_stage decorator) and run them inside a Pipeline with an executor. Curator’s parallelism and resource scheduling apply when code runs as pipeline stages; ad‑hoc Pandas code executes on the driver and will not scale.