Text-to-Audio
Text-to-Audio
Synthesize speech with vLLM-Omni through the /v1/audio/speech endpoint
Text-to-audio (TTS) generation runs a vLLM-Omni worker with --output-modalities audio. See the Diffusion Overview for installation and shared configuration.
Tested Models
Launch
Launch using the provided script with Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice:
The same script serves Nemotron Audex, which runs a two-stage pipeline: stage 0 is the thinker that emits speech codec tokens, and stage 1 decodes those tokens to a 16 kHz mono waveform.
Both Audex checkpoints report the same model type, so the 30B-A3B needs an explicit stage configuration. Otherwise auto-detection selects the 2B-tuned file. vLLM-Omni ships the configuration, so resolve it from the installed package and pass it through to the worker:
Generate Speech
CustomVoice (predefined speaker)
CustomVoice + style
VoiceDesign (describe a voice)
Audex (guided speech)
Parameters
The /v1/audio/speech endpoint follows the vLLM-Omni API format. TTS-specific parameters are top-level fields, except single-model knobs, which ride in nvext:
Text to synthesize.
TTS model name.
Speaker name (e.g., vivian, ryan). Validated against model config.
Audio output format.
Speed factor (0.25–4.0).
Synthesis task type (Qwen3-TTS).
Language code. Validated against model config.
Voice style/emotion description. Required for VoiceDesign.
Reference audio URL or base64 data URI. Required for Base.
Transcript of reference audio (Base task).
Maximum tokens to generate (1–4096).
Classifier-free guidance scale for Nemotron Audex (1.0–10.0). Model-specific, so it rides in nvext rather than at the top level (vLLM-Omni likewise takes it under extra_params). 1.0 disables guidance. Omitting the field decodes unguided on an Audex speech deployment, but an Audex text-to-audio (TTA) deployment falls back to its default of 3.0. Ignored by other audio models.
Available voices and languages are loaded dynamically from the model’s config.json at startup. Nemotron Audex builds its own prompt and accepts only input, model, response_format, speed, max_new_tokens, and nvext.cfg_scale; speech models additionally tolerate voice when it is omitted or default, while an Audex TTA deployment rejects voice outright. Other non-Qwen3-TTS audio models (e.g., MiMo-Audio) use a generic text prompt and ignore TTS-specific parameters.
stream: true) and the Base task (voice cloning) are not yet supported.