All models

MiniMax-H3 ★

MiniMaxVideo
Get your API key
MiniMax-H3

Native audio short-film model integrating text, image, audio, and video references

MiniMax-H3 is a general-purpose multimodal audio-video model for short-film creation. It can combine text, images, video, and audio to understand creative intent and generate videos with native stereo sound. It is suitable both for building shots from a script and for using first and last frames to control transitions, or for guiding characters, products, actions, and sound with various reference materials. It is suited for advertising, character storytelling, and content with audiovisual rhythm.

MiniMaxModel brand
VideoModel type
VideoTask capability

Specifications and API features

Clarify capacity, inputs and outputs, and invocation methods before selecting a model.

Video duration
4–15 seconds, whole-second durations
Output resolution
768P, 2K
Aspect ratio control
21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and adaptive
Native audiovisual specifications
24 FPS video; 32 kHz stereo audio
Creation methods
Text-to-video, first frame/last frame/first-and-last frames, multimodal reference generation
Number of references
Up to 9 images, 3 video clips, and 3 audio clips; up to 12 files in total
Text and delivery
Text entries up to 7000 characters; supports waiting for results, asynchronous queries, and callbacks

Frame rate and audio sampling rate are native model specifications; this platform entry provides controls for resolution, duration, aspect ratio, and asset roles.

Core Capabilities

Learn what MiniMax-H3 can bring to your work.

Create visuals and sound together

H3 jointly generates video and stereo sound, so you do not have to completely separate visual concepts from sound design. Prompts can simultaneously describe subject actions, camera changes, ambient sounds, dialogue, and musical atmosphere, allowing sound to serve the scene's expression rather than simply adding a background track after the video is finished.

Use first and last frames to guide camera direction

Providing a first frame can bring existing product images, posters, or character visuals to life; providing a last frame can specify the camera's final destination; providing both can generate the transition in between. Focus on describing the transitional action, and keep the composition, proportions, and lighting of the two images as consistent as possible.

Give different assets their respective roles

Reference images guide people, clothing, products, or style; reference videos guide actions and camera movement; reference audio guides timbre, music, or rhythm. Clearly defining the role of each asset is better suited to organizing complex creations and makes it easier to determine whether the final video retains key characteristics.

Use Cases

Start with specific tasks to find where the model can make an impact.

Product and brand shorts

Input product images, model images, and scene references, then describe product close-ups, how people use the product, and how it should be presented at the end to generate horizontal or vertical advertising assets. Specify the person's appearance and the product structure separately, and before delivery, focus on checking outlines, how the product is worn, reflections, and brand text.

Character storytelling and storyboard previews

Input character and spatial references, along with descriptions of character relationships, emotional changes, shot sizes, and dialogue, to create short drama clips or storyboard previews. Use close-ups to convey eyes and expressions, use reference images to guide clothing and appearance, then incorporate suitable short clips into subsequent editing workflows.

Action and music-rhythm content

Use action videos, character images, and audio as references, specify who performs the action, what camera approach to use, and how sound should complement the visuals to generate dance or fashion rhythm shorts. First trim assets into short segments relevant to the target clip to prevent unrelated content from distracting from the creative focus.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

How to Choose Between H3 and H3 Max

Choose H3 when you need 2K output, short shots starting at 4 seconds, or complete audiovisual creation based on multimodal materials. H3 Max is another speed-optimized model, with native output at 480P or 768P and durations of 5–15 seconds; consider it when faster generation is more important, but do not mix up the specifications of the two.

Choose a Creation Method by Control Goal

When you have no existing materials, use text to clearly specify the subject, actions, shots, and sound; when you already have a defined opening or ending frame, choose first-and-last-frame mode; when you need to incorporate a person's appearance, actions, and voice, choose reference mode. First-and-last-frame materials and reference materials cannot be included in the same request, so determine the most important control goal first.

Get Started

From a small-scale task to full integration.

01

Prepare Tasks and Materials

Clarify goals, required inputs, and output requirements, and use real business examples as a starting point.

02

Try It in the API Debugging Area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate According to the API Documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage Limits

Before formal use, understand the output quality and capability scope.

  • Each generation is limited to 4–15 seconds, with a resolution of 768P or 2K. Longer narratives need to be split into shots and then edited; continuity of people, spaces, and sound across segments should be checked segment by segment. Do not treat a single short-video generation as long-form video production.
  • Each reference video and audio clip must be 2–15 seconds, with total duration across each type not exceeding 15 seconds; mixed references may include up to 12 files in total. More reference materials are not necessarily better: clearly define what to retain, and avoid conflicting requirements for actions, styling, or sound.
  • Material references and first-and-last-frame controls are used to guide generation and do not guarantee pixel-for-pixel replication. Product details, text, complex actions, and dialogue transitions still need to be checked; when involving portraits, voices, and brand assets, use materials with legal authorization.

Frequently Asked Questions

Answers to common questions about using MiniMax-H3.

Can MiniMax-H3 generate videos using only text?

Yes. Provide a non-empty text item in content to start text-to-video generation, and set resolution, duration, and a fixed ratio. Prompts are recommended to describe the subject, action, scene, camera, lighting, and sound in that order; adaptive aspect ratios cannot be used in text-only mode.

What is the difference between a first-frame image and a reference image?

first_frame specifies the opening frame of the video, while reference_image guides the subject appearance, product, or style and is not required to become the beginning. Use last_frame to control the final frame; first-and-last-frame mode cannot be used together with reference_* assets.

Can audio be used to guide sound and rhythm?

reference_audio can be used in reference mode, with text specifying whether it is for voice timbre, dialogue, music, or rhythm. Audio assets support WAV and MP3; H3 natively supports stereo generation, and both Chinese and English are stably supported dialogue languages.

Is 2K output just standard upscaling?

H3's native 2K workflow uses regeneration, combining the base video with the original context to restore details, rather than relying solely on a conventional super-resolution module. Select 2K for resolution when calling it to express the output requirement; you do not need to organize these internal generation steps yourself.

How do I submit a task and obtain the finished video?

Submit MiniMax-H3 and content assets to /minimax/videos, which waits for completion and returns a task by default. Pass async=true or callback_url to immediately obtain a task identifier, then query through /minimax/tasks or receive a callback; after success, obtain the finished video from task.content.url.

Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.

Use MiniMax-H3 for your next task

Start with a clear goal and judge from real results whether it suits your work.