What is the main difference between tts-1 and tts-1-hd?
The main difference is their optimization focus: tts-1 is designed for real-time use cases, while tts-1-hd prioritizes audio quality. For instant prompts and conversational announcements, consider tts-1 first; for polished narration, compare audio samples using the same script, as you cannot determine which model was used based on the file format alone.
Which voices can tts-1 use?
You can choose alloy, echo, fable, onyx, nova, or shimmer, with alloy as the default voice. It is recommended to test each one with actual business scripts before settling on a voice. The names themselves do not promise an accent, emotion, or character, nor do they indicate the ability to replicate any person's voice.
How do I explicitly call tts-1?
Submit the input to be spoken to POST /v1/audio/speech, and explicitly set model to tts-1; add voice, response_format, and speed as needed. Explicitly selecting the model makes the request intent clearer, and the returned result should be handled as binary audio data.
What formats can it output, and can the speed be adjusted?
Audio format options include mp3, opus, aac, flac, wav, and pcm, with mp3 as the default; speed defaults to 1.0. Choose a format based on your player and audio processing workflow, and listen after adjusting the speed rather than treating unspecified numeric ranges as usable limits.
Can tts-1 directly listen to recordings and respond?
You cannot treat tts-1 as a complete voice conversation model. It receives text and generates speech; if user input is a recording, it must first be transcribed, then response text generated, and finally passed to tts-1 for narration. This also makes it easier to review the content before playback and improve the spoken expression of the text.