nemo_curator.models.asr.qwen_asr
nemo_curator.models.asr.qwen_asr
Qwen3-ASR vLLM implementation of the shared ASR adapter.
This uses the same Qwen3ASRModel.LLM construction and vLLM engine settings
as the nkoluguri reference. ASRStage owns mono conversion and resampling;
the adapter hands one prepared batch to one transcribe call and maps
results back to ASRResult positions.
Module Contents
Classes
Functions
Data
API
Run vLLM-backed Qwen3-ASR over Curator waveform items.
Every valid item in one adapter call goes to a single transcribe call,
so the caller’s batch boundary is the model’s batch boundary.
max_inference_batch_size is the library’s own internal cap and is passed
through at construction.
revision is an adapter-owned Hugging Face option and is forwarded to
both weight prefetch and the vLLM model loader.
vllm_kwargs exposes additional engine settings, following the existing
Qwen-Omni adapter convention. Adapter-owned settings cannot be overridden
through this mapping. Its default is empty, so normal construction exactly
matches the nkoluguri reference engine arguments.
Return the qwen-asr constructor arguments owned by this adapter.
Populate the local Hugging Face cache without allocating a GPU.
Load one worker-local Qwen3-ASR model through its vLLM backend.
Transcribe one adapter call while preserving input order.
Release the worker-local model and CUDA cache state.