Can OmniHuman 1.5 generate a talking video from text only?
This workflow requires a character image and driving audio; text alone cannot replace these two assets. You can first record or create the script as an mp3/wav, then submit the audio URL. The prompt is used to adjust expressions, emotion, and style; it does not directly read the lines aloud.
What conditions should the photo meet?
It is recommended to use a clear, front-facing portrait with good lighting and no obstructions, with the person taking up an appropriate portion of the frame. The photo needs a publicly accessible URL. Before formal production, you can test the photo with a short audio clip to check lip sync, expressions, and head-and-shoulder movement, then continue using the asset.
Can I make a specific person in a group photo speak?
You can submit an array of subject mask URLs through mask_url to specify the person to be driven in a multi-person photo; even if there is only one mask URL, it should still be submitted as an array. You should first identify the target and prepare the corresponding mask. This feature is for subject selection and should not be understood as allowing multiple characters to speak separately or complete a multi-person conversation in a single request.
How can I control tone and character movements?
The voice itself provides speaking speed, pauses, and delivery rhythm, while the prompt can add requirements such as gentle, calm, natural, or slight head movements. It is recommended to keep the voice-over and prompt consistent, avoiding an impassioned voice with text requesting calmness; the final performance still needs to be checked through the generated video.
How do I submit and retrieve videos in the application?
Submit image_url and audio_url to POST /dreamina/videos, and specify omnihuman-1.5. Results can be obtained synchronously for short tasks; for longer tasks, you can use callback_url, or set async:true and then query /dreamina/tasks. After completion, read video_url and save it.