Text-to-Audio

Synthesize speech with vLLM-Omni through the /v1/audio/speech endpoint

以 Markdown 格式查看

Text-to-audio (TTS) generation runs a vLLM-Omni worker with --output-modalities audio. See the Diffusion Overview for installation and shared configuration.

Tested Models

ModelNotes
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoiceDefault model; predefined speakers
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesignDescribe a voice via instructions
nvidia/Nemotron-Labs-Audex-2BSingle built-in voice; optional nvext.cfg_scale guidance
nvidia/Nemotron-Labs-Audex-30B-A3BSame contract as the 2B; needs an explicit stage config

Launch

Launch using the provided script with Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice:

bash examples/backends/vllm/launch/agg_omni_audio.sh

The same script serves Nemotron Audex, which runs a two-stage pipeline: stage 0 is the thinker that emits speech codec tokens, and stage 1 decodes those tokens to a 16 kHz mono waveform.

bash examples/backends/vllm/launch/agg_omni_audio.sh --model nvidia/Nemotron-Labs-Audex-2B

Both Audex checkpoints report the same model type, so the 30B-A3B needs an explicit stage configuration. Otherwise auto-detection selects the 2B-tuned file. vLLM-Omni ships the configuration, so resolve it from the installed package and pass it through to the worker:

AUDEX_30B_CFG=$(python3 -c 'import pathlib, vllm_omni; print(pathlib.Path(vllm_omni.__file__).parent / "deploy/audex_tts_30b.yaml")')
bash examples/backends/vllm/launch/agg_omni_audio.sh \
--model nvidia/Nemotron-Labs-Audex-30B-A3B \
--stage-configs-path "$AUDEX_30B_CFG"

Generate Speech

curl -X POST http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Hello, how are you?",
"voice": "vivian",
"language": "English"
}' --output output.wav

Parameters

The /v1/audio/speech endpoint follows the vLLM-Omni API format. TTS-specific parameters are top-level fields, except single-model knobs, which ride in nvext:

input
stringRequired

Text to synthesize.

model
stringDefaults to auto-detected

TTS model name.

voice
stringDefaults to Vivian

Speaker name (e.g., vivian, ryan). Validated against model config.

response_format
wav | mp3 | pcm | flac | aac | opusDefaults to wav

Audio output format.

speed
floatDefaults to 1.0

Speed factor (0.25–4.0).

task_type
CustomVoice | VoiceDesign | BaseDefaults to CustomVoice

Synthesis task type (Qwen3-TTS).

language
stringDefaults to Auto

Language code. Validated against model config.

instructions
string

Voice style/emotion description. Required for VoiceDesign.

ref_audio
string

Reference audio URL or base64 data URI. Required for Base.

ref_text
string

Transcript of reference audio (Base task).

max_new_tokens
intDefaults to 2048

Maximum tokens to generate (1–4096).

nvext.cfg_scale
float

Classifier-free guidance scale for Nemotron Audex (1.0–10.0). Model-specific, so it rides in nvext rather than at the top level (vLLM-Omni likewise takes it under extra_params). 1.0 disables guidance. Omitting the field decodes unguided on an Audex speech deployment, but an Audex text-to-audio (TTA) deployment falls back to its default of 3.0. Ignored by other audio models.

Available voices and languages are loaded dynamically from the model’s config.json at startup. Nemotron Audex builds its own prompt and accepts only input, model, response_format, speed, max_new_tokens, and nvext.cfg_scale; speech models additionally tolerate voice when it is omitted or default, while an Audex TTA deployment rejects voice outright. Other non-Qwen3-TTS audio models (e.g., MiMo-Audio) use a generic text prompt and ignore TTS-specific parameters.

Audio streaming (stream: true) and the Base task (voice cloning) are not yet supported.

See Also