HiAPI
  • Models
  • Pricing
Search

Search HiAPI models, tools, and resources.

  • Models
  • Pricing
HiAPI

One API, All AI Models

Generate images, video, and audio with leading models through one production-ready API.

Get a free API key

AI Image API

  • All image models
  • GPT Image 2
  • Nano Banana 2
  • Seedream 5.0 Pro
  • Qwen Image 2.0 Pro
  • FLUX 1.1 Pro

AI Video API

  • All video models
  • Seedance 2.5
  • FLUX.3 Video
  • Seedance 2.0
  • Veo 3.1
  • Kling 3.0

AI Audio API

  • All audio models
  • MiniMax Music 2.6
  • MiniMax Music 1.5
  • ElevenLabs v3
  • Text to music
  • Text to speech

Product

  • Model marketplace
  • Playground
  • Pricing
  • Image API Cost Calculator
  • Free GPT Image 2 Generator
  • Free Nano Banana Image Generator
  • Outfit Preview
  • Product Photo Lab

Developers

  • Documentation
  • API Reference
  • Agent Skills
  • LLM integration index
  • Blog

Company

  • About
  • Contact support
  • Terms of Service
  • Privacy Policy

© 2026 hiapi. All rights reserved.

Open source on GitHubPython SDK on PyPI
  • Summary
  • What MiniMax H3 Actually Is
  • Text-to-Video: A Real Clip From a Real Prompt
  • Image-to-Video: Starting From a Locked Frame
  • Reference-Driven Generation: Images, Video, and Audio Together
  • How Requests Work
  • Pricing and Budgeting for Short-Form Batches
  • FAQ
  • Takeaways
GuideAug 14, 2026

MiniMax H3 for Short-Form Video: A Hands-On Guide to the hiapi API

hiapiMiniMax H3Text-to-VideoImage-to-VideoShort-Form Video

Latest models

Explore models

Contents
  • Summary
  • What MiniMax H3 Actually Is
  • Text-to-Video: A Real Clip From a Real Prompt
  • Image-to-Video: Starting From a Locked Frame
  • Reference-Driven Generation: Images, Video, and Audio Together
  • How Requests Work
  • Pricing and Budgeting for Short-Form Batches
  • FAQ
  • Takeaways

Generate it with HiAPI

Choose a model, enter your prompt, and see the result.

HiAPI Blog

Related articles

HiAPI

Generate it with HiAPI

MiniMax H3 is hiapi's native 2K multimodal video model, and the part that matters most for short-form content — TikTok, Reels, Shorts — is that it renders directly in vertical 9:16 without any letterboxing tricks. We ran a real text-to-video generation through the MiniMax H3 model page on hiapi and embedded the actual output below, alongside the exact prompt we used and a practical breakdown of its three generation modes.

Summary

  • MiniMax H3 generates natively at 2K and renders clean vertical 9:16 output — no cropping or upscaling needed for TikTok/Reels/Shorts delivery.
  • It supports three generation modes: text-to-video, image-to-video (with first/last-frame control), and reference-driven generation from images, video, or audio.
  • Clips run 4 to 15 seconds at $0.1189/second (as of 2026-08 data from hiapi's pricing page) — a 4-second clip costs roughly $0.48, a 15-second clip roughly $1.78.
  • Specific, physical prompt language (materials, light direction, camera framing) produced a clean, on-model result in our test — vague prompts are the most common cause of drifting or generic-looking output.
  • All requests go through hiapi's single asynchronous task endpoint, the same submit-and-poll shape used across every video model on the platform.

What MiniMax H3 Actually Is

MiniMax H3 is a native 2K multimodal video model built for text-to-video, first/last-frame image-to-video control, and reference-driven generation using images, video, and audio, across clip lengths of 4 to 15 seconds. For short-form content specifically, the native 2K output matters more than it might seem: most vertical social formats top out around 1080×1920, so a 2K source gives you headroom to crop for a thumbnail or a slightly tighter reframe without visible quality loss — you're downscaling, not upscaling.

Text-to-Video: A Real Clip From a Real Prompt

We generated this clip directly from a text prompt, no reference image or video involved. Here's the actual output:

Prompt:

A barista's hand pours steamed milk from a stainless steel pitcher into a black espresso cup held at a slight angle on a light grey marble countertop. The milk stream settles and blooms into a full rosetta (fern-leaf) latte-art pattern by the end of the pour. Warm, directional side light rakes across the cup and pitcher. Vertical 9:16 framing, macro-close perspective, shallow depth of field, no camera movement, no text, no logos, no additional hands or objects entering frame.

This is the kind of shot that works well for short-form: a single, physically grounded action (a pour, a reveal, a transformation) shot close and vertical, with enough specific detail — the material of the pitcher, the direction of the light, the exact pattern forming — that the model has almost nothing left to guess at. Naming the camera behavior explicitly (we asked for "no camera movement") is also worth doing even when you want a static shot; without it, models will sometimes add a slow, unrequested drift. We cover camera-movement prompting in more depth in our guide to cinematic camera movements in AI video, which applies the same principle to models that support explicit pans, orbits, and pushes.

Image-to-Video: Starting From a Locked Frame

MiniMax H3's image-to-video mode takes a reference image through the image_urls input field and uses it as a starting frame, with first/last-frame control letting you optionally pin both the opening and closing composition of the clip. This is the better choice for short-form work when the opening frame is the hook — a product shot, a face, a specific piece of packaging — and you need the video to start from that exact composition rather than let the model interpret it from a text description. Text-to-video is faster to iterate on wording; image-to-video is what you reach for once you've already nailed the still frame you want the clip to open on.

Reference-Driven Generation: Images, Video, and Audio Together

The third mode is reference-driven generation, where MiniMax H3 can take cues from a reference image, a reference video, or a reference audio track — or a combination of the three — rather than generating from a text prompt alone. In practice this is the mode to reach for when you already have an asset whose visual style, motion, or timing you want to carry into a new clip: matching a new subject to an established brand look, or syncing a generated clip's pacing to an existing audio track for a short-form cut. It's a heavier, more deliberate workflow than text-to-video, so it's worth prototyping the shot with a plain text prompt first and only moving to reference-driven generation once you know what you're trying to match.

How Requests Work

Every video model on hiapi, including MiniMax H3, goes through the same asynchronous task flow: you submit a generation request and get back a task ID, then poll that task until it reports a finished status and returns an output URL. There's no long-lived connection to manage and no model-specific request shape to relearn — the same submit-and-poll pattern in hiapi's docs covers every model on the platform, so switching from testing one video model to another is a matter of changing the model name and its input fields, not rebuilding your integration.

Pricing and Budgeting for Short-Form Batches

MiniMax H3 is billed at $0.1189 per second of output (as of 2026-08 data from hiapi's pricing page), with clip lengths from 4 to 15 seconds. That puts a single clip at roughly:

  • 4 seconds: ~$0.48
  • 8 seconds: ~$0.95
  • 15 seconds: ~$1.78

For short-form content specifically, this favors a "shoot short, cut together" approach over one long generation: two or three 4-6 second clips edited together in a caption tool cost less combined than a single 15-second render, give you more cut points to hold viewer attention, and let you re-roll just the one clip that didn't land instead of the whole sequence.

FAQ

Does MiniMax H3 output true vertical video, or does it need cropping for TikTok/Reels? It renders natively at 2K and accepts a 9:16 aspect ratio directly — the clip above was generated vertical from the start, no cropping or reframing needed for short-form platforms.

How long can a MiniMax H3 clip be? 4 to 15 seconds per generation. For short-form content, stitching several shorter clips (4-6 seconds each) usually produces a more watchable result than one long single-shot render.

Can I start a MiniMax H3 video from my own image? Yes — image-to-video mode accepts a reference image through the image_urls field and uses it as the starting frame, with first/last-frame control available if you also want to pin the closing frame.

What's the difference between image-to-video and reference-driven generation on MiniMax H3? Image-to-video locks a starting (and optionally ending) frame from a still image. Reference-driven generation goes further, letting the model take style, motion, or timing cues from a reference image, video, or audio track — useful when you're matching a new clip to an existing asset rather than starting from a blank frame.

How much does a typical short-form clip cost? At $0.1189/second, a 4-second clip runs about $0.48 and an 8-second clip about $0.95 (pricing as of 2026-08 — check hiapi's live pricing page for current rates). Budgeting a handful of short clips per concept is usually cheaper than one long generation you might need to re-roll entirely.

Is MiniMax H3 a good fit if I'm already using another hiapi video model, like Hailuo, for short-form clips? They're complementary rather than competing choices — see our guide on using Hailuo 2.3 for short-form video for a text-to-video comparison point, and pick based on which mode (plain text-to-video vs. image-anchored or reference-driven generation) your shot actually needs.

Takeaways

  • MiniMax H3 renders natively in 2K vertical, so short-form clips need no cropping or upscaling for TikTok/Reels/Shorts.
  • Specific, physically-grounded prompts (materials, light direction, explicit camera behavior) produced a clean result in our real test — vague prompts leave too much for the model to guess.
  • Use text-to-video to iterate on an idea quickly, image-to-video when the opening frame is locked in advance, and reference-driven generation when you're matching an existing image, video, or audio asset.
  • At $0.1189/second across a 4-15s range, several short clips cut together is usually more cost-efficient and more watchable than one long single-shot render.
  • All requests go through the same task-submit-and-poll flow as every other video model on hiapi, so testing MiniMax H3 alongside models like Hailuo 2.3 doesn't require separate integration work.

Latest models

View all models
  • GPT Image 2From $0.007/image
  • Nano Banana 2From $0.051/image
  • Seedream 5.0 ProFrom $0.050/image
  • Seedance 2.5From $0.121/s

Explore models

TextImageVideoAudio
Back to blog
GPT Image 2From $0.007/image
Nano Banana 2From $0.051/image
Seedream 5.0 ProFrom $0.050/image
Seedance 2.5From $0.121/s
View all models
TextChat and reasoning
ImageGenerate and edit
VideoText and image to video
AudioSpeech and music
Start generating
View model pricing
View all articles
Seedance 2.5 Text-to-Video: Build Short-Form Clips with the hiapi API

Seedance 2.5 Text-to-Video: Build Short-Form Clips with the hiapi API

Seedance 2.5 Reference-to-Video for Short-Form TikTok and Reels Clips

Seedance 2.5 Reference-to-Video for Short-Form TikTok and Reels Clips

Grok Imagine Image 2.0 Image-to-Image Prompts: 4 Recipes With Real Outputs

Grok Imagine Image 2.0 Image-to-Image Prompts: 4 Recipes With Real Outputs

Grok Imagine 2.0 Text-to-Image Prompt Recipes: Copy-Paste Prompts With Real Outputs

Grok Imagine 2.0 Text-to-Image Prompt Recipes: Copy-Paste Prompts With Real Outputs

Using flux-2-klein-9b/text-to-image for E-Commerce Product Images via the hiapi API

Using flux-2-klein-9b/text-to-image for E-Commerce Product Images via the hiapi API

Flux-2-Klein-9b Image-to-Image for E-Commerce Product Photos

Flux-2-Klein-9b Image-to-Image for E-Commerce Product Photos

Start generating