AI Models Market: Text, Image, Video & Audio APIs
The HiAPI models market brings together text LLMs, image, video, music, and speech models. Available models include one-key access, usage-based pricing, an online Playground, and request examples; upcoming models include capability and launch previews. Filter by provider or task type.
- codex-gpt-5.4-mini API — text generation, usage-based pricing; see live pricing. GPT-5.4 mini is an efficient reasoning model for coding, tool use, and high-throughput agent workflows, with a 400K-token context window, image input, and configurable reasoning effort.
- DeepSeek V4 Flash API — text generation, usage-based pricing; see live pricing. DeepSeek V4 Flash is the efficiency-focused open-weight MoE text model in the DeepSeek V4 family, with a 1M-token context window, high/max reasoning, structured output, and tool calling. HiAPI exposes this model through the Chat Completions API.
- DeepSeek V4 Pro API — text generation, usage-based pricing; see live pricing. DeepSeek V4 Pro (0813) targets complex reasoning, coding, and agent workflows. It natively supports the Responses API and remains compatible with Chat Completions.
- ElevenLabs Text to Dialogue API — music generation, usage-based pricing; see live pricing. Create multi-speaker dialogue with ElevenLabs Text to Dialogue. Hear a real audio drama, map voices to lines, review pricing, and copy API code.
- FLUX 1.1 Pro API — image generation, usage-based pricing; see live pricing. Black Forest Labs' most advanced image generation model with exceptional photorealism and prompt adherence
- FLUX.2 Image to Image API — image generation, usage-based pricing; see live pricing. Edit and compose images with FLUX.2 Image to Image using up to eight references, auto ratio, 1K/2K output, a matched watercolor example, and API quickstart code.
- FLUX.2 [klein] 4B Image to Image API — image generation, usage-based pricing; see live pricing. Compact, fast single-reference FLUX.2 editing with natural-language changes and output up to 4 MP.
- FLUX.2 [klein] 4B Text to Image API — image generation, usage-based pricing; see live pricing. Try FLUX.2 [klein] 4B Text to Image for low-cost drafts and commercial visuals with eleven ratios, five megapixel tiers, seed control, and API code.
- FLUX.2 [klein] 9B Image to Image API — image generation, usage-based pricing; see live pricing. Edit with FLUX.2 [klein] 9B Image to Image using one reference, real cleanup and text-replacement examples, five ratios, four to eight inference steps, and API code.
- FLUX.2 [klein] 9B Text to Image API — image generation, usage-based pricing; see live pricing. Try FLUX.2 [klein] 9B Text to Image with real event, product, magazine, and neon examples, five aspect ratios, four to eight inference steps, and API code.
- FLUX.2 Pro API — image generation, usage-based pricing; see live pricing. Black Forest Labs FLUX.2 Pro: high-fidelity text-to-image with strong prompt adherence and crisp in-image text.
- FLUX.3 Video API — video generation, usage-based pricing; see live pricing. Try the FLUX.3 Video API with text prompts, 1 to 10 keyframes, source-video continuation, synchronized audio, Draft previews, and per-second pricing.
- FLUX.1 Schnell API — image generation, usage-based pricing; see live pricing. FLUX.1 Schnell, Black Forest Labs' open-source ultra-fast text-to-image model: 1-4 inference steps with second-level output, Apache 2.0 commercial license, and HiAPI's lowest per-image price tier — built for high-volume generation and rapid iteration.
- gemini-3.1-flash-tts API — music generation, usage-based pricing; see live pricing. Google Gemini 3.1 Flash TTS generates natural single- or multi-speaker dialogue with per-speaker voice, accent, style, and pace controls plus scene-level direction.
- GPT-5.6 Luna API — text generation, usage-based pricing; see live pricing. GPT-5.6 Luna targets cost-sensitive, high-throughput workloads such as bulk classification, extraction, rewriting, and lightweight coding. HiAPI currently exposes text, multi-turn, and function tool calling through the streaming Responses API.
- GPT-5.6 Sol API — text generation, usage-based pricing; see live pricing. GPT-5.6 Sol is the flagship reasoning model in OpenAI's GPT-5.6 family for complex coding, professional analysis, and demanding agent workflows. HiAPI currently exposes text, multi-turn, and function tool calling through the streaming Responses API.
- GPT-5.6 Terra API — text generation, usage-based pricing; see live pricing. GPT-5.6 Terra balances intelligence, speed, and cost for everyday coding, document work, and general agent workflows. HiAPI currently exposes text, multi-turn, and function tool calling through the streaming Responses API.
- GPT Image 2 API — image generation, usage-based pricing; see live pricing. Use the GPT Image 2 text-to-image model on HiAPI for posters, knowledge cards, product visuals, and other generated images. Try it online or integrate the async API.
- GPT Image 2 Image-to-Image API — image generation, usage-based pricing; see live pricing. OpenAI image-to-image model for fast, high-quality image generation and editing with flexible aspect ratios and resolutions.
- GPT Image 2 Multi-ratio 4K I2I API — image generation, usage-based pricing; see live pricing. GPT Image 2 image-to-image high-resolution multi-ratio line: blend and edit with up to 6 reference images, across 16 aspect ratios x 1K/2K/4K x low/medium/high quality tiers. Precise control of output framing and finish for editing, style transfer and multi-image fusion.
- gpt-image-2/image-to-image@official API — image generation, usage-based pricing; see live pricing. GPT Image 2 direct image-to-image editing with one reference image and low, medium, or high quality output.
- gpt-image-2/image-to-image@pro API — image generation, usage-based pricing; see live pricing. GPT Image 2 Pro image-to-image: high-quality stable tier, upload 1-5 reference images to edit/compose/repaint, stability-first for production use.
- GPT Image 2 Beta API — image generation, usage-based pricing; see live pricing. GPT Image 2 Beta is the preview of OpenAI's next-generation image model. It works with the standard OpenAI Images API format and produces high-quality images, ideal for creative design, posters, illustrations, product concepts, and social media assets.
- GPT Image 2 Multi-ratio 4K API — image generation, usage-based pricing; see live pricing. GPT Image 2 high-resolution multi-ratio line: freely combine 16 aspect ratios (incl. 5:4, 4:5, 2:1, 21:9) x 1K/2K/4K resolutions x low/medium/high quality tiers, up to 3840x2160 output. Built for posters, banners and print assets that demand exact framing and finish.
- gpt-image-2/text-to-image@official API — image generation, usage-based pricing; see live pricing. GPT Image 2 direct text-to-image with low, medium, and high quality across common landscape and portrait sizes.
- gpt-image-2/text-to-image@pro API — image generation, usage-based pricing; see live pricing. GPT Image 2 high-quality stable tier: stability-first, designed for production scenarios with strict requirements on success rate and latency consistency.
- Grok Imagine 1.5 Image to Video API — video generation, usage-based pricing; see live pricing. xAI Grok Imagine 1.5 image-to-video: animate a still image with upgraded motion quality, 1-15 seconds, 480p/720p, same price as 1.0.
- grok-imagine-image-2.0/image-to-image API — image generation, usage-based pricing; see live pricing. xAI Grok Imagine Image 2.0 image editing: provide one reference image and describe the edit in natural language, with 1k/2k output.
- grok-imagine-image-2.0/text-to-image API — image generation, usage-based pricing; see live pricing. xAI Grok Imagine Image 2.0 text-to-image: generate images with low or medium quality, 1k/2k resolution, and flexible square-to-ultrawide aspect ratios.
- Grok Imagine Image to Image API — image generation, usage-based pricing; see live pricing. Edit with Grok Imagine Image to Image using up to three references, a matched cabin-to-night example, fourteen ratios, flat 1k/2k pricing, and API code.
- Grok Imagine Image to Video API — video generation, usage-based pricing; see live pricing. xAI Grok Imagine image-to-video animates one or more reference images into a cinematic short clip, with an optional motion prompt, selectable aspect ratio, duration (6-30s), and 480p/720p resolution.
- Grok Imagine Quality Image to Image API — image generation, usage-based pricing; see live pricing. Use Grok Imagine Quality Image to Image with up to three references, a matched night-scene edit, fourteen ratios, 1k/2k quality pricing, and API code.
- Grok Imagine Quality Text to Image API — image generation, usage-based pricing; see live pricing. Create hero-grade cinematic images with Grok Imagine Quality Text to Image, thirteen ratios, 1k/2k tiers, real API examples, and quickstart code.
- Grok Imagine Text to Image API — image generation, usage-based pricing; see live pricing. Generate cinematic landscapes and night scenes with Grok Imagine Text to Image, thirteen aspect ratios, flat 1k/2k pricing, real API examples, and quickstart code.
- Grok Imagine Text to Video API — video generation, usage-based pricing; see live pricing. xAI Grok Imagine text-to-video generates cinematic short clips from a text prompt, with selectable motion mode, aspect ratio, duration (6-30s), and 480p/720p resolution.
- hailuo-2.3-fast/image-to-video API — video generation, usage-based pricing; see live pricing. MiniMax Hailuo 2.3 Fast image-to-video: the speed tier at the lowest price, 6s/10s per-video billing.
- hailuo-2.3/image-to-video API — video generation, usage-based pricing; see live pricing. MiniMax Hailuo 2.3 image-to-video: first-frame driven with natural motion, 6s/10s clips billed per video.
- hailuo-2.3/text-to-video API — video generation, usage-based pricing; see live pricing. MiniMax Hailuo 2.3 text-to-video: the standard tier with excellent motion physics, 6s/10s clips billed per video.
- HappyHorse 1.0 API — video generation, usage-based pricing; see live pricing. Alibaba HappyHorse 1.0 generates short, cinematic videos from a single prompt, with smooth motion and strong scene consistency. Choose 720p or 1080p, set the duration, and create realistic AI video clips in seconds.
- HappyHorse 1.1 Image-to-Video API — video generation, usage-based pricing; see live pricing. Alibaba HappyHorse 1.1 image-to-video: drive generation from a first frame while keeping subject and style consistent, native audio, 720p/1080p.
- HappyHorse 1.1 Reference-to-Video API — video generation, usage-based pricing; see live pricing. Alibaba HappyHorse 1.1 reference-to-video: generate video from up to 9 reference images while keeping subject, scene, and style consistent, native audio.
- HappyHorse 1.1 Text-to-Video API — video generation, usage-based pricing; see live pricing. Alibaba HappyHorse 1.1 text-to-video: cinematic motion, strong prompt adherence, native audio, 720p/1080p.
- Ideogram V4 API — image generation, usage-based pricing; see live pricing. Try Ideogram V4 for accurate in-image text, logos, posters, and signage with TURBO, BALANCED, and QUALITY tiers plus real API examples.
- Kling 3.0 Omni Image-to-Video API — video generation, usage-based pricing; see live pricing. Kuaishou Kling 3.0 Omni image-to-video: drive generation from first/last frames, native synced audio, up to 4K.
- kling-3.0-omni/reference-to-video API — video generation, usage-based pricing; see live pricing. Kuaishou Kling 3.0 Omni reference-to-video: lock subjects/props with named reference elements (cite via @name in the prompt), native audio, up to 4K.
- Kling 3.0 Omni Text-to-Video API — video generation, usage-based pricing; see live pricing. Kuaishou Kling 3.0 Omni text-to-video: cinematic motion, multi-shot storytelling, native synced audio, up to 4K.
- Kling 3.0 Turbo Image-to-Video API — video generation, usage-based pricing; see live pricing. Kling 3.0 Turbo image-to-video: first-frame driven generation at speed, 3-15s, 720p/1080p, billed per second.
- Kling 3.0 Turbo Text to Video API — video generation, usage-based pricing; see live pricing. Try Kling 3.0 Turbo text-to-video in the HiAPI Playground. Compare real 720p results, copy tested prompts, review 3–15 second parameters and pricing, then call POST /v1/tasks.
- kling-4-0-text-to-video API — video generation, coming soon. Kuaishou Kling 4.0 text-to-video: the next-generation cinematic video model with upgraded prompt fidelity, scene coherence, and motion control. Coming soon.
- lyria-3-pro API — music generation, usage-based pricing; see live pricing. Google Lyria 3 Pro generates full songs up to approximately three minutes from a prompt and optional inspiration images. Images guide the mood, theme, and arrangement; they are not video frames or audio source media. Lyrics, timestamps, and multilingual prompts are supported.
- minimax-h3 API — video generation, usage-based pricing; see live pricing. MiniMax H3 is a native 2K multimodal video model for text-to-video, first/last-frame control, and reference-driven generation with images, video, and audio across 4 to 15 seconds.
- MiniMax Music 1.5 API — music generation, usage-based pricing; see live pricing. Generate complete Chinese or English songs with MiniMax Music 1.5. Test real tracks, structured lyrics, audio settings, pricing, and API code.
- MiniMax Music 2.6 API — music generation, usage-based pricing; see live pricing. Generate full songs with MiniMax Music 2.6 using custom lyrics, automatic lyrics, or instrumental mode. Hear real tracks and copy API code.
- minimax-music-3 API — music generation, usage-based pricing; see live pricing. MiniMax Music 3 generates complete songs up to five minutes from a music description and lyrics, with detailed control over genre, mood, vocals, instrumentation, and arrangement, output as 44.1 kHz stereo WAV.
- Nano Banana API — image generation, usage-based pricing; see live pricing. Google's powerful image generation model with stunning quality and fast generation times
- Nano Banana 2 API — image generation, usage-based pricing; see live pricing. Nano Banana 2 — the Google Gemini 3.1 Flash Image model. Built for developers, it pairs lightning speed with Pro-grade quality: precise text rendering, strong character consistency, and up to 4K output. The best balance of speed, quality, and price for large-scale image generation and editing workflows. Commercial license supported.
- Nano Banana 2 Lite API — image generation, usage-based pricing; see live pricing. Google Nano Banana 2 Lite: low-latency, ultra low-cost 1K image generation with optional reference images (up to 10) for editing and remixing.
- Nano Banana Pro API — image generation, usage-based pricing; see live pricing. Nano Banana Pro — the Gemini 3 Pro Image model from Google DeepMind and the flagship of the Nano Banana series. Delivers top-tier output: sharper 2K imagery, smart upscaling, advanced text rendering, and outstanding character consistency. Built for high-end creative work, brand asset generation, and API-driven production workflows. Commercial license supported.
- Qwen-Audio 3.0 TTS Flash API — music generation, usage-based pricing; see live pricing. Try Qwen-Audio 3.0 TTS Flash for low latency text-to-speech. Hear a real multilingual sample, compare voices and controls, and copy API code.
- Qwen-Audio 3.0 TTS Plus API — music generation, usage-based pricing; see live pricing. Try Qwen-Audio 3.0 TTS Plus for high quality text-to-speech. Hear a real multilingual sample, compare voices and controls, and copy API code.
- Qwen Image 2.0 API — image generation, usage-based pricing; see live pricing. Alibaba Qwen Image 2.0 - cost-effective image generation with excellent Chinese text rendering, multi-style output, up to 2K resolution.
- Qwen Image 2.0 Pro API — image generation, usage-based pricing; see live pricing. Try Qwen Image 2.0 Pro with real Chinese typography, portrait, product, and concept examples, five 2K-class sizes, prompt extension, and API code.
- Qwen Image 3.0 Image to Image API — image generation, usage-based pricing; see live pricing. Use Qwen Image 3.0 Image to Image with up to three references, a verified product recolor example, eight ratios, 1K/2K output, and API code.
- Qwen Image 3.0 Pro Image to Image API — image generation, usage-based pricing; see live pricing. Edit images with Qwen Image 3.0 Pro using up to three references, exact cover-text replacement, 1K/2K pricing, eight ratios, and API quickstart code.
- qwen-image-3.0-pro/text-to-image API — image generation, usage-based pricing; see live pricing. Qwen Image 3.0 Pro text-to-image for high-fidelity final images, complex materials, and refined layouts.
- qwen-image-3.0/text-to-image API — image generation, usage-based pricing; see live pricing. Qwen Image 3.0 text-to-image for Chinese text, posters, concepts, and everyday visual creation.
- Seedance 2.0 API — video generation, usage-based pricing; see live pricing. ByteDance Seedance 2.0 - ByteDance's latest video generation model with cinematic quality, exceptional motion, and native audio.
- Seedance 2.0 Ext API — video generation, usage-based pricing; see live pricing. Seedance 2.0 extended route for cinematic video generation up to 4K, with native audio plus image and audio references.
- Seedance 2.0 Fast API — video generation, usage-based pricing; see live pricing. ByteDance Seedance 2.0 Fast: high-speed video generation with native audio. Text-to-video, image-to-video (first/last frame), and multimodal reference-to-video (image/video/audio). 480p/720p, 4-15s.
- Seedance 2.0 Mini API — video generation, usage-based pricing; see live pricing. Seedance 2.0 Mini by ByteDance is a cost-efficient video generation model supporting both text-to-video and image-to-video, with native synced audio, first/last-frame control, and image/video/audio multimodal references. Up to 720P, flexible 4–15s clips.
- Seedance 2.5 Image to Video API — video generation, usage-based pricing; see live pricing. Try Seedance 2.5 image-to-video with first-frame and first/last-frame control. Compare real 720p results, copy tested prompts, review 4–30 second pricing, and call POST /v1/tasks.
- Seedance 2.5 Reference to Video API — video generation, usage-based pricing; see live pricing. Try Seedance 2.5 reference-to-video with video, image, and audio references. Learn @video1 mappings, review 4–30 second pricing, and call POST /v1/tasks.
- Seedance 2.5 Text to Video API — video generation, usage-based pricing; see live pricing. ByteDance Seedance 2.5 text-to-video turns a single prompt into one continuous shot of up to 30 seconds at 720p or 1080p, across seven aspect ratios, with natively synchronised audio and prompts in 11 languages.
- Seedream 4.5 Image to Image API — image generation, usage-based pricing; see live pricing. ByteDance Seedream 4.5 image editing: unified generation-editing architecture with up to 14 reference images for edits and composites, consistent subjects, 2K/4K output.
- Seedream 4.5 Text to Image API — image generation, usage-based pricing; see live pricing. Try Seedream 4.5 Text to Image with real landscape, food, sci-fi, and Chinese-aesthetic examples, eight ratios, 2K/4K output, and API code.
- Seedream 5.0 Lite Image to Image API — image generation, usage-based pricing; see live pricing. Edit with Seedream 5.0 Lite Image to Image using up to 14 references, real recolor, text replacement, and composite examples, 2K or 4K output, and API code.
- Seedream 5.0 Lite Text to Image API — image generation, usage-based pricing; see live pricing. Try Seedream 5.0 Lite Text to Image with real poster, infographic, banner, and photography examples, eight aspect ratios, 2K or 4K output, and API code.
- Seedream 5.0 Pro Image to Image API — image generation, usage-based pricing; see live pricing. Edit with Seedream 5.0 Pro Image to Image using one to ten references, real recolor, weather, and multi-image examples, 1K or 2K pricing, and API code.
- Seedream 5.0 Pro Text to Image API — image generation, usage-based pricing; see live pricing. Try Seedream 5.0 Pro Text to Image with real product, portrait, poster, and cinematic examples, 1K or 2K pricing, eight aspect ratios, and API code.
- Veo 3.1 Fast Image to Video API — image generation, usage-based pricing; see live pricing. Google Veo 3.1 Fast image-to-video: fast, cost-efficient image animation with native audio, up to 4K, 4/6/8-second clips.
- Veo 3.1 Fast Text to Video API — video generation, usage-based pricing; see live pricing. Google Veo 3.1 Fast: high-speed text-to-video with native audio, up to 4K, supports 4/6/8-second clips.
- Veo 3.1 Image to Video API — image generation, usage-based pricing; see live pricing. Google Veo 3.1 image-to-video: animate a still image into a cinematic clip with native audio, up to 4K, 4/6/8 seconds.
- veo-3.1-lite/image-to-video API — video generation, usage-based pricing; see live pricing. Veo 3.1 Lite image-to-video with native audio, 720p or 1080p output, and 4, 6, or 8 second clips.
- veo-3.1-lite/reference-to-video API — video generation, coming soon. Google Veo 3.1 Lite reference-to-video: generate 8-second clips from 1–3 reference images and a prompt with native audio, portrait output, and 720p/1080p/4K tiers.
- veo-3.1-lite/text-to-video API — video generation, usage-based pricing; see live pricing. Veo 3.1 Lite text-to-video with native audio, 720p or 1080p output, and 4, 6, or 8 second clips.
- Veo 3.1 Text to Video API — video generation, usage-based pricing; see live pricing. Google Veo 3.1: flagship text-to-video with native audio, cinematic realism, up to 4K, 4/6/8-second clips.
- Wan 2.7 Image Text to Image API — image generation, usage-based pricing; see live pricing. Try Wan 2.7 Image Text to Image with Chinese prompts, 1K/2K/4K output, extreme wide and tall ratios, optional thinking mode, and API code.
- Wan 2.7 Image-to-Video API — video generation, usage-based pricing; see live pricing. Animate any still image into a high-quality video with natural motion, supporting first-frame, first+last-frame, and video continuation modes
- Wan 2.7 Text-to-Video API — video generation, usage-based pricing; see live pricing. Alibaba's latest video generation model with cinematic quality, native audio support, and up to 1080P 15-second output
- wan3.0-video API — video generation, usage-based pricing; see live pricing. Wan 3.0 all-purpose video generation from text, first/last frames, reference media, files, or webpages, with 480P, 720P, or 1080P output and optional synchronized audio.
- Z-Image API — image generation, usage-based pricing; see live pricing. Try Z-Image for low-cost Chinese-friendly portraits, signage, city scenes, and food photography with five ratios, real examples, and API code.
Image to Image & Editing APIs
Black Forest Labs Model APIs