NeMo ASR Models
Use NeMo Framework’s automatic speech recognition models for transcription in your audio curation pipelines. This guide covers basic usage and configuration.
Model Selection
NeMo Framework provides pre-trained ASR models through the Hugging Face model hub. For the complete list of available models and their specifications, refer to the NeMo Framework ASR documentation.
Example Model Usage
Basic Usage
Simple ASR Inference
Custom Configuration
Model Caching
The executor downloads and caches model weights once per node, then loads one adapter in each worker:
Resource Configuration
Configure GPU and CPU resources based on your hardware:
NeMoASRAdapter uses zero or one GPU per worker. Scale throughput by allowing the executor to run more one-GPU workers instead of assigning multiple GPUs to one stage instance.
Resource requirements vary by model. Test with your specific model to determine optimal settings.