nemo_curator.models.asr.qwen_asr

View as Markdown

Qwen3-ASR vLLM implementation of the shared ASR adapter.

This uses the same Qwen3ASRModel.LLM construction and vLLM engine settings as the nkoluguri reference. ASRStage owns mono conversion and resampling; the adapter hands one prepared batch to one transcribe call and maps results back to ASRResult positions.

Module Contents

Classes

NameDescription
QwenASRAdapterRun vLLM-backed Qwen3-ASR over Curator waveform items.

Functions

NameDescription
_qwen_asr_model_cls-

Data

_DEFAULT_QWEN3_ASR_MODEL

_MIN_SAMPLES

API

class nemo_curator.models.asr.qwen_asr.QwenASRAdapter(
model_id: str = _DEFAULT_QWEN3_ASR_MODEL,
revision: str | None = None,
gpu_memory_utilization: float = 0.7,
max_new_tokens: int = 4096,
max_inference_batch_size: int = 128,
vllm_kwargs: dict[str, typing.Any] = dict()
)
Dataclass

Run vLLM-backed Qwen3-ASR over Curator waveform items.

Every valid item in one adapter call goes to a single transcribe call, so the caller’s batch boundary is the model’s batch boundary. max_inference_batch_size is the library’s own internal cap and is passed through at construction.

revision is an adapter-owned Hugging Face option and is forwarded to both weight prefetch and the vLLM model loader.

vllm_kwargs exposes additional engine settings, following the existing Qwen-Omni adapter convention. Adapter-owned settings cannot be overridden through this mapping. Its default is empty, so normal construction exactly matches the nkoluguri reference engine arguments.

_model
Any = field(default=None, init=False, repr=False)
gpu_memory_utilization
float = 0.7
max_inference_batch_size
int = 128
max_new_tokens
int = 4096
model_id
str = _DEFAULT_QWEN3_ASR_MODEL
revision
str | None = None
vllm_kwargs
dict[str, Any] = field(default_factory=dict)
nemo_curator.models.asr.qwen_asr.QwenASRAdapter.__post_init__() -> None
nemo_curator.models.asr.qwen_asr.QwenASRAdapter._adapter_owned_model_kwargs() -> dict[str, typing.Any]

Return the qwen-asr constructor arguments owned by this adapter.

nemo_curator.models.asr.qwen_asr.QwenASRAdapter._waveform(
item: dict[str, typing.Any]
) -> numpy.ndarray
staticmethod
nemo_curator.models.asr.qwen_asr.QwenASRAdapter.download_weights_on_node() -> None

Populate the local Hugging Face cache without allocating a GPU.

nemo_curator.models.asr.qwen_asr.QwenASRAdapter.load_model(
num_gpus: int
) -> None

Load one worker-local Qwen3-ASR model through its vLLM backend.

nemo_curator.models.asr.qwen_asr.QwenASRAdapter.transcribe_batch(
items: list[dict[str, typing.Any]]

Transcribe one adapter call while preserving input order.

nemo_curator.models.asr.qwen_asr.QwenASRAdapter.unload_model() -> None

Release the worker-local model and CUDA cache state.

nemo_curator.models.asr.qwen_asr._qwen_asr_model_cls() -> typing.Any
nemo_curator.models.asr.qwen_asr._DEFAULT_QWEN3_ASR_MODEL = 'Qwen/Qwen3-ASR-0.6B'
nemo_curator.models.asr.qwen_asr._MIN_SAMPLES = 1600