Is gpt-4o-image the name of a standalone model released by OpenAI?
Here, gpt-4o-image is a compatible entry point for conversational image creation. It is used to organize GPT-4o image generation workflows and should not be treated as another standalone official model or confused with gpt-image-1. It is mainly selected to submit requirements through conversational messages and receive creation results.
Can I generate images without a reference image?
Yes. For text-to-image generation, simply clearly describe the image you want, such as a sunset scene in a futuristic city. It is recommended to specify the subject, environment, style, and lighting, and directly state that you want an image generated; if the image needs text, list the exact copy separately for easier subsequent review.
How can I use a character image and a prop image together?
Combine text blocks and multiple image_url blocks in the same user message in Chat Completions, describing the character, props, and action relationship. For example, request that a character hold a reference coffee cup and prepare to drink from it. Reference images provide visual information; the final image still needs to be checked for appearance and spatial relationships.
Which field should I read the generated result from?
When using /openai/chat/completions, read message.content from choices, which contains the image and download link. The client needs to identify and display the image link, then download and save the file; do not write the entire content string directly to a file as image data.
Should I send messages or input when integrating?
It depends on the entry point: Chat Completions uses model and messages, while Responses uses model and input. Hosted conversation entry points use methods such as question or structured message. Existing applications with image-and-text messages can prioritize continuing to use Chat Completions to avoid mixing request structures from different entry points.