Skip to content
English

Audio Models

Audio models run on the same unified POST /v1/tasks endpoint and API key as image and video models.

ModelBest forNotes
ElevenLabs Text to Dialogue v3Multi-speaker dialogue: audio dramas, podcasts, game NPCs67 preset voices, 70+ languages, per-line voice control; billed per character

The following pages are pre-launch contracts. Dev acceptance has passed; production access remains closed until the release gate opens.

ModelPlanned roleReadiness
Qwen-Audio 3.0 TTS PlusHigh-quality voiceovers, video narration, and spoken contentTwo system voices, PCM/WAV/MP3/Opus, and complete speech controls; dev acceptance passed
Qwen-Audio 3.0 TTS FlashAI assistants, voice agents, notifications, and responsive speech workflowsFour system voices, PCM/WAV/MP3/Opus, and complete speech controls; dev acceptance passed

More audio models (music generation, sound effects) are on the way. Check HiAPI Pricing for the live model list.