Skip to content
English

Changelog

August 2026

  1. New model docs: qwen-image-3.0/text-to-image, qwen-image-3.0/image-to-image, qwen-image-3.0-pro/text-to-image, and qwen-image-3.0-pro/image-to-image. All use POST /v1/tasks; image-to-image requires 1–3 reference images, while every model supports 8 aspect ratios, 1K/2K output, PNG/JPEG delivery, prompt controls, and task callbacks.

  2. New model: minimax-music-3 — MiniMax high-fidelity song generation with controllable length. Required lyrics with structure tags, duration caps output at 1-300 seconds (default 60), and seed reproduces results. Lossless WAV output billed per second of requested duration.

  3. deepseek-v4-pro is now documented for native POST /v1/responses and compatible Chat Completions, with a 1M-token context, 384K maximum output, none/low/high/max reasoning, streaming, JSON output, and tools.

  4. Remote MCP v0.4.0 now exposes 9 live-catalog tools. New generate_text support selects Chat Completions or Responses by model, get_task_status retrieves long-running media tasks, and generic parameters let newly listed task models work without a server release. The Agent Skills guide now includes the latest Seedance 2.5 and Seedream 5.0 Pro integrations.

  5. The seedance-2.0 ext route now exposes only its supported 720p, 1080p, and 4k tiers. The unsupported 2k option has been removed from pricing and request documentation.

  6. New models: seedance-2.5/text-to-video, seedance-2.5/image-to-video, and seedance-2.5/reference-to-video. The default is 480p for 5 seconds; you can choose 720p and 4–30 seconds. All support synchronized audio; image-to-video supports first/last-frame control, while reference-to-video requires 1–10 reference videos and accepts optional image and audio references.

  7. deepseek-v4-flash is available through POST /v1/chat/completions, with streaming, thinking controls, high/max reasoning effort, JSON mode, tools, and Token usage. Media models continue to use POST /v1/tasks.

  8. New model: flux-3 by Black Forest Labs for text-to-video, 1-10 image keyframe storyboards, and source-video continuation. It generates 5-20 second clips at 720p/1080p with optional synchronized audio and a lower-cost 720p Draft tier.

July 2026

  1. New model: minimax-h3 — native 2K video generation through POST /v1/tasks, with 4-15 second duration, text-to-video, first/last-frame control, and image/video/audio multimodal references. Billed per output second.

  2. New image-editing models: flux-2-klein-4b/image-to-image for low-cost single-reference edits with 0.25-4 MP output tiers, and flux-2-klein-9b/image-to-image for fast four-step edits with stronger detail preservation and text rendering.

  3. New models: flux-2-klein-4b/text-to-image for cost-efficient 0.25-4 MP generation, and flux-2-klein-9b/text-to-image for fast four-step output with stronger realistic detail and readable in-image text.

  4. seedance-2.0 now documents standard and ext routes on one page. Keep the base model name and pass top-level route: "ext"; the ext route supports 720p/1080p/4k plus image and audio references.

  5. Go and Java joined the official SDK lineup — go get github.com/HiAPIAI/hiapi-go and ai.hiapi:hiapi on Maven Central, alongside the existing Python SDK (pip install hiapi). All three now ship v0.2.1 with model route selection and Idempotency-Key retries built in. See the SDKs page.

  6. New models: Seedream 5.0 Pro Text to Image and Image to Image — HiAPI flagship quality tier. Upgraded photorealism with sharp in-image text rendering (brush calligraphy and signage verified), resolution 1K/2K output 1K/2K images billed per image; editing takes 1-10 reference images.

  7. New model family: MiniMax Hailuo 2.3 — standard hailuo-2.3/text-to-video and image-to-video (6/10s per-video tiers, motion-physics leader), plus hailuo-2.3-fast image-to-video — the price floor of the family.

  8. New models: Kling 3.0 Turbo Text to Video and Image to Video — the speed tier of the Kling 3.0 family. 3-15 second clips at 720p/1080p with per-second tiered billing and lip-synced quoted dialogue; image-to-video is first-frame driven.

  9. New models: Grok Imagine Text to Image and Image to Image — xAI's image generation line with 13 aspect ratios (ultra-wide 2:1/20:9 to extra-tall 9:20) and reference edits with up to 3 images. Both bill one flat rate per image; the grok-imagine-quality family offers hero-grade output tiered by 1k/2k resolution.

  10. New models: Seedream 4.5 Text to Image and Image to Image — ByteDance's community-favorite quality tier. Photorealistic detail with accurate in-image text rendering (English and Chinese), 8 aspect ratios, 2K/4K at one flat per-image price; editing takes up to 14 reference images with subject and text consistency.

  11. New model: FLUX.2 Image to Image — pro-tier image editing with up to 8 reference images per request. Multi-image composition, background replacement, style and material edits with strong subject and logo consistency; aspect_ratio: auto follows the first input, 1K/2K output billed per image.

  12. New model: MiniMax Music 2.6 — full-song music generation with natural vocals in English or Chinese. Lyrics up to 3,500 characters with 14 structure tags, instrumental mode via is_instrumental, and auto-generated lyrics when lyrics is left empty. Billed per song.

  13. POST /v1/tasks now accepts an optional top-level route parameter for models with multiple routes — passing the bare model name with route: "pro" is the preferred spelling of the legacy @pro suffix (the old suffix keeps working). An unknown route returns a 400 listing the available routes, and the task detail echoes route plus the resolved full model name. See Model routes on the Create Task page.

  14. New models: five additions across image, video and audio — Nano Banana 2 Lite (entry-tier 1K image generation), Qwen Image 2.0 Pro (pro-grade text rendering and photorealism), Ideogram V4 (best-in-class typography with TURBO/BALANCED/QUALITY tiers), Grok Imagine 1.5 Image to Video preview (upgraded motion, 1–15s), and MiniMax Music 1.5 (full songs up to 4 minutes with vocals, HiAPI's first music model).

  15. New models: Seedream 5.0 Lite Text to Image and Image to Image — ByteDance's reasoning-capable image model with real-time web knowledge and accurate multilingual text rendering. Editing supports up to 14 reference images and in-image text rewriting. 2K/4K output at the same per-image price.

  16. POST /v1/tasks now accepts an optional Idempotency-Key header — retries with the same key under the same account create the task only once, and replays return the original taskId, preventing duplicate tasks and duplicate charges from timeout retries. See Idempotency key in the Unified Async API intro.

  17. New models: Veo 3.1 Text to Video and Image to Video — Google's flagship video tier with cinematic quality, native audio, up to 4K, and 4/6/8-second clips. Also added Veo 3.1 Fast Image to Video, the cost-efficient way to animate stills.

  18. New model: ElevenLabs Text to Dialogue v3 — HiAPI's first audio model. Multi-speaker dialogue speech with a voice per line (67 presets), 70+ languages, tunable stability, billed per character.

  19. New models: Grok Imagine Text to Video and Image to Video — xAI Grok Imagine video generation with selectable motion mode, aspect ratio, 6–30s duration, and 480p/720p resolution.

  20. seedance-2.0 now supports 4k resolution — pick 480p, 720p, 1080p, or 4k for ultra HD output.

  21. New model: Seedance 2.0 Fast — ByteDance's high-speed video model for text-to-video, image-to-video (first/last frame), and multimodal reference-to-video (image/video/audio), with native audio. 480p/720p, 4–15s, billed per second.

June 2026

  1. New model: Seedance 2.0 Mini — ByteDance's cost-efficient video model for both text-to-video and image-to-video, with native audio, first/last-frame control, and image/video/audio multimodal references. 480p/720p, 4–15s, billed per second.

  2. Output Storage is live — keep the images and videos you create with POST /v1/tasks past the default 7 days. Pick a storage tier when you create a task, or promote existing outputs to long-term storage, then list and delete them from the API or your dashboard. Results are still fetched the same way, via GET /v1/tasks/:id or a callback.url.

May 2026

  1. New model: HappyHorse 1.0 text-to-video — 720p/1080p, 3–15s, multiple aspect ratios.

  2. Agent Skills — call a single model straight from Codex, Claude Code, OpenClaw and other agents with one install.

April 2026

  1. gpt-image-2 added 1K / 2K / 4K resolution tiers.

  2. Remote MCP server — connect any MCP client and generate with image and video tools over a single endpoint.

  3. New models: Qwen Image 2.0 and Seedance 2.0.