Create Dynamic Video Assets by Blending Multiple Reference Images
Veo 3.1 Fast Ingredients is the video entry point in the Google Veo series for creating with blended reference images, suitable for projects that already have product images, character images, or scene assets. It uses 1-3 images together with text prompts to generate dynamic visuals, focusing on organizing reference content into video rather than specifying first and last frames. Images are required; text-only generation is not supported. It is suitable for ad concepts, concept shots, and iterative asset planning.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation Method
Generate video by blending reference images; images are required
Image Input
1-3 image links, submitted via image_urls
Aspect Ratio Options
16:9 landscape, 9:16 portrait
Output Settings
720p by default; 1080p or gif available
Prompt Control
Use text to describe scenes, actions, and shots; supports automatic translation with translation
Tasks and Results
Supports asynchronous tasks and callbacks; returns video ID, video_url, and status after completion
The above are the input and output settings for this entry point on this platform and do not represent the complete specifications of all native Veo series versions.
Core Capabilities
Organize Reference Assets into Dynamic Scenes
Use images to provide visual content, then use prompts to explain their relationships and actions within the scene. You can develop videos around product, character, or background assets without having to describe every aspect of their appearance using text alone. Multi-image blending is suitable for exploring asset combinations; when submitting images, clearly specify which elements are the main subjects and which are environmental references.
Define Shot Intent with Prompts
Images provide visual references, while text provides narrative and motion instructions. Describe how the subject moves, how the camera observes it, and the atmosphere of the image, rather than simply writing “make the image move.” Chinese-language creation can enable automatic translation with translation, making it easier to use existing Chinese scripts for video tasks while keeping the prompt structure clear.
Integrate into Asynchronous Asset Production Workflows
Generation tasks can be submitted asynchronously: first save the task_id, then query the completed result; you can also set callback_url to receive result notifications. The completion response includes the video link, video ID, and status, making it easier to associate each generation with asset combinations and prompt versions, and organize subsequent downloading, review, and asset archiving.
Applicable Scenarios
Product Advertising Concept Shots
Input product images and scene references, then use text to describe the product's position, presentation action, and camera movement to generate advertising visuals for team discussion. Suitable for comparing different environments and presentation approaches before formal shooting or detailed retouching, delivering reviewable dynamic assets rather than simply assembling multiple static images into a carousel.
Character and Environment Combination Previsualization
When character designs and environment concept art are already available, use them as blend references and describe shots of the character entering the scene, observing the environment, or completing an action. The results are suitable for concept shorts and storyboard discussions, helping the team determine whether the visual combination works before deciding whether to continue producing more complete shots.
Landscape and Portrait Creative Asset Exploration
Plan landscape and portrait compositions separately around the same set of reference images, and generate asset options with different action or shot descriptions. Landscape can be used for scene presentation, while portrait can emphasize the subject. When delivering, record versions by aspect ratio, reference image, and prompt for easy selection of visuals suited to subsequent editing and publishing layouts.
How to Choose This Model
Blend References or Control the First and Last Frames
If the goal is to organize content from several assets into a dynamic scene, choose veo31-fast-ingredients. If the opening or ending visual is already determined and you want images to serve as the first frame or the first and last frames, standard veo31-fast is better suited to the task. Both use reference images, but their control semantics differ: blending references does not fix the start and end points of the video.
Choose the Relevant Version Based on Creation Conditions
If you already have images and want to explore combination options, prioritize the Ingredients entry point; if you only have a text script, choose veo31-fast or veo31, which support text-to-video. If the project explicitly requires 4K, consider veo31, which supports that setting, rather than adding 4k to an Ingredients request. First ensure the choice meets the input method and delivery requirements, then compare the actual visuals.
Getting Started
Understand the purpose of each reference image
Prepare 1–3 publicly accessible images, clearly describing how the product, person, or scene should be combined; these images are material references, not start and end frames.
Select the dedicated creation endpoint
Specify model=veo31-fast-ingredients and action=ingredients2video for /veo/videos, and provide image_urls and prompt. Start with aspect_ratio=16:9 and resolution=720p, then increase the output level within this model's supported range.
Save the task and video IDs separately
After setting async=true, save the task_id and retrieve the completed video through /veo/tasks or a callback; verify reference material integration and subject consistency; if the result includes an audio track, listen to it, then save the video ID and downloaded file for subsequent production.
Trial suggestion: product and environment reference integration
Input and objective
The first image is a teacup, the second image is a wooden table, and the third image provides a window-side environment; naturally place the teacup on the wooden table, with the camera slowly pushing in under warm light.
Acceptance criteria and next steps
You must provide 1–3 image_urls and select ingredients2video; the images are used as integration references and do not mean the first image will necessarily become the first frame.
Usage boundaries
You must provide 1-3 reference images; it cannot be used as a pure text-to-video model. Even a complete prompt cannot replace image input; when preparing materials, first determine the subjects and scenes that need to be referenced, and avoid submitting more images than the quantity limit.
Multi-image integration does not mean the start and end frames are locked. In this endpoint, two images still serve as integration references, and you should not assume that the first image is fixed as the opening or the second image as the ending. If the shot must transition from a specified image to another image, use a start-and-end-frame creation method.
This endpoint does not offer a 4K output option, so you cannot reuse veo31's 4K settings. For tasks requiring precise product appearance, logos, or composition, it is recommended to treat generated images as materials pending review, check them item by item before moving on to editing and delivery, and avoid treating integration references as pixel-perfect reproductions.
Frequently Asked Questions
Can I use Ingredients with just one image?
Yes, you can input 1–3 images, and one image also meets the requirements. However, creation here still follows the reference fusion approach and is not equivalent to a fixed first frame. It is recommended to specify in the prompt the main content you want to preserve, as well as the desired action or scene changes.
How do I submit a multi-image fusion task?
Send a POST request to /veo/videos, set model to veo31-fast-ingredients, use the ingredients2video operation, provide image links through image_urls, and use prompt to describe the fusion relationship and camera intent. Images must be provided.
What is the main difference between it and veo31-fast?
Ingredients requires images and focuses on fusing reference content; veo31-fast supports text-only generation, as well as a single-image first frame and two-image first and last frames. Choose the former when you want to combine assets, and the latter when you want to start from a specified image or connect specified first and last images.
Can I generate 1080p or vertical video assets?
You can select 1080p and choose 9:16 vertical or 16:9 horizontal through aspect_ratio. When output settings are not specified, 720p is used by default. When planning vertical video assets, it is recommended to also specify the subject position and composition in the prompt, to avoid changing only the aspect ratio without adjusting the camera design.
How do I get the video after asynchronous submission?
After setting async=true, first obtain the task_id, then retrieve the status and result through task queries; you can also use callback_url to receive completion notifications. After completion, read video_url from data to get the video, and save the video ID to facilitate linking generation records and subsequent processing.