Audio Models
Audio models run on the same unified POST /v1/tasks endpoint and API key as image and video models.
| Model | Best for | Notes |
|---|---|---|
| ElevenLabs Text to Dialogue v3 | Multi-speaker dialogue: audio dramas, podcasts, game NPCs | 67 preset voices, 70+ languages, per-line voice control; billed per character |
Coming soon
Section titled “Coming soon”The following pages are pre-launch contracts. Dev acceptance has passed; production access remains closed until the release gate opens.
| Model | Planned role | Readiness |
|---|---|---|
| Qwen-Audio 3.0 TTS Plus | High-quality voiceovers, video narration, and spoken content | Two system voices, PCM/WAV/MP3/Opus, and complete speech controls; dev acceptance passed |
| Qwen-Audio 3.0 TTS Flash | AI assistants, voice agents, notifications, and responsive speech workflows | Four system voices, PCM/WAV/MP3/Opus, and complete speech controls; dev acceptance passed |
More audio models (music generation, sound effects) are on the way. Check HiAPI Pricing for the live model list.