All models

minimax-t2v

MiniMaxVideo
Get your API key
minimax-t2v

Generate previewable video assets from text-based ideas

minimax-t2v is MiniMax's text-to-video entry point, designed for creative tasks without a first-frame image that aim to turn text concepts directly into dynamic visuals. Enter descriptions of the subject, environment, actions, and visual atmosphere to start video generation and obtain a finished video link. It is suitable for concept previews, creative proposals, and asset exploration; compared with reference-image-driven minimax-i2v, the creative starting point is text rather than existing images.

MiniMaxModel brand
VideoModel type
VideoTask capability

Specifications and API features

Clarify capacity, inputs and outputs, and invocation methods before selecting a model.

Creation method
Text-to-video, no first-frame reference image required
Content input
prompt text prompt
Generation request
POST /hailuo/videos;model=minimax-t2v;action=generate
Task execution
Supports async=true, returns task_id first
Completion notification
Supports callback_url for receiving POST JSON results
Result delivery
video_url video link and task status information

This lists the actual creation and invocation methods on this platform and does not treat native parameters from other Hailuo versions as specifications for this model.

Core capabilities

Learn what minimax-t2v can bring to your work.

No need to create an image first—start with a description

minimax-t2v uses text as the starting point for creation, making it suitable for the concept stage when image assets are not yet available. You can structure descriptions around “who is doing what in which environment,” then add lighting, color tones, and visual atmosphere. Turn ideas into watchable dynamic drafts first, then decide whether to create reference images or proceed with further editing.

Iterate on visuals through prompts

Creative input is centered on the prompt, making it easy to retain the same theme while progressively adjusting scene or action descriptions. It is recommended to specify one main change each time, such as changing a sunset coast to an overcast coast, and then compare the results. This iterative approach is suited to exploring visual directions, but it is not equivalent to precisely locking camera movement or controlling individual frames.

Separate generation from result management

An asynchronous task workflow is supported: save the task_id after submission, then query the status or receive a completion callback. The video_url in the result is used to obtain the video, while the task identifier corresponds to the original concept. For applications that need to submit multiple pieces of copy continuously, this approach makes it easy to separate generation from preview, review, and asset archiving.

Applicable Scenarios

Start with specific tasks to find where the model can be effective.

Dynamic Proposals for Advertising Creative

Rewrite a visual segment in an advertising script as descriptions of the subject, scene, and action—for example, a person walking down a street on a rainy night—then generate a video draft for team discussion. The key deliverable is to present the creative direction and atmosphere, rather than directly replacing a final ad that requires accurate trademarks, packaging text, and character identities.

Concept Previsualization for Story Scenes

Enter scene descriptions from a story into the model to explore how environments, actions, and visual atmosphere can be combined. It is suitable for creating dynamic references for individual plot points before formal filming or animation production. If the story requires the same character to appear consistently, it is recommended to first use this model to explore directions, then plan subsequent shots in combination with a reference-image-based creation approach.

Asset Exploration in Content Production

Write prompts around ocean waves, city streetscapes, or other themes, generate candidate videos, and preview them through links to select assets suitable for subsequent editing. Different prompts and task results can be archived together for easy comparison and reuse. Subtitles, music, voice-over, and full timeline arrangement should be planned as separate production steps.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

When You Do Not Have Images, Start with Text

If you only have a creative idea or a scene description, minimax-t2v better matches the current input conditions, so there is no need to prepare a first frame before submitting the task. If you already have a product image, character image, or a defined composition and want to use it as the starting point for a video, minimax-i2v is more suitable. The main difference between the two is the basis for creation; text-to-video should not be treated as animation processing for a specified image.

Choose Another Creation Method for Fine Control Needs

minimax-t2v is suitable for text-driven visual exploration and should not be used as Director mode. If the task starts with a reference image and focuses on enhanced creative control, consider minimax-i2v-director. When choosing, first consider whether image constraints are needed, then consider control requirements; do not assume that different models have the same camera capabilities simply because they all belong to the Hailuo series.

Get Started

From a small-scale task to full integration.

01

Prepare Tasks and Materials

Define the objective, required inputs, and output requirements, using real business examples as a starting point.

02

Try It in the API Testing Area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate According to the API Documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage Limitations

Before formal use, understand the output quality and capability scope.

  • Text descriptions do not strictly lock in a person's identity, product appearance, or composition. If a task requires faithfully continuing a specific image, use the reference-image creation entry point; this model is better suited for first exploring visual concepts and then manually selecting results, rather than handling precise reproduction tasks for specified images.
  • Do not treat camera descriptions in prompts as dedicated camera-control instructions. Complex continuous camera movements, multiple simultaneous actions, or strict frame-by-frame arrangements should be split into clearer independent creative ideas for testing; do not expect this model to have the control capabilities of the Director model.
  • The core deliverable is the generated video and its result link. Voice-over, sound effects, lip sync, or a complete editing timeline should not be treated as default deliverables. When these production steps are needed, arrange separate audio and editing workflows, and check before use whether the actual video meets project requirements.

Frequently Asked Questions

Answers to common questions about using minimax-t2v.

Does minimax-t2v require uploading an image?

No, it generates videos from text prompts. It is recommended to clearly describe the main subject, environment, and primary action first, then add atmosphere details. If your task must start from an existing image, choose minimax-i2v instead of treating a first-frame image as a required input for minimax-t2v.

What is the difference between it and minimax-i2v?

minimax-t2v starts from a text concept; minimax-i2v starts from a reference image and requires an image link. Choose the former when you have no visual assets and want to explore creative ideas; choose the latter when you already have a specific image and want to create dynamic content around it, as it better fits the task requirements.

Is minimax-t2v a Director version?

It should be used through the text-to-video entry point and should not be regarded as a Director model. For reference-image-driven Director creation, choose minimax-i2v-director. Even if you include camera movement requirements in the prompt, this does not mean you gain Director-exclusive preset shots or fine-grained motion control.

How do I write prompts suitable for it?

It is recommended to structure the prompt around a clear scene, describing the main subject, the action taking place, the environment, and the visual atmosphere. First avoid conflicting action requirements, then compare results by changing one element at a time. This approach makes iteration easier; it does not mean complex instructions can always be executed precisely.

How do I retrieve the generated video after submission?

When calling it, you can set async=true, first record the returned task_id, then query the task result; you can also configure callback_url to receive completion notifications. After obtaining the result, use the status information to determine whether the task succeeded, and retrieve the video through video_url in data; do not treat the task ID as the finished video URL.

Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.

Use minimax-t2v for your next task

Start with a clear goal and evaluate whether it suits your work based on actual results.