OpenAI Text to Speech
OpenAI Text to Speech: voices, pricing and the 2027 shutdown
OpenAI Text to Speech turns written scripts into multilingual narration, with delivery prompts, stock voices and streamed audio. Its current Speech API models are deprecated and scheduled to shut down on 6 January 2027; OpenAI recommends a move to Realtime 2.1 Mini.
Paid plansAccount required

- Access
- Paid access
- Usage pricing
- USD 15 / 1,000,000 charactersTTS-1 — deprecated
- Available on
- Web, API
Scripted speech with a confirmed retirement date
OpenAI Text to Speech supplies narration for articles, learning materials and spoken application responses. Its Speech endpoint takes a script and returns playable audio or a stream. On 1 October 2026, OpenAI deprecated TTS-1, TTS-1 HD and both dated GPT-4o Mini TTS releases, scheduling their removal for 6 January 2027. Existing integrations need to account for that deadline. OpenAI recommends GPT-Realtime-2.1 Mini as the replacement, through the Realtime API.
GPT-4o Mini TTS accepts directions for accent, emotion, intonation and pace. The current API reference lists 13 built-in voice names, while the older TTS models use a smaller selection and do not accept those delivery instructions. Output supports common compressed and lossless formats. Custom voices are available to eligible customers after approval, with a matching speaker reference and a separate consent recording. Prompt-designed voices belong to the Live API rather than this Speech endpoint.
Speech pricing depends on the model: TTS-1 and HD charge by input characters, while Mini TTS charges separately for text input and generated audio tokens. Realtime 2.1 Mini has its own rates and connection model, including WebRTC, WebSocket and SIP. It should be treated as a migration requiring integration work. OpenAI requires clear disclosure that the voice is AI-generated. API content is not used for training by default; standard Speech abuse-monitoring logs can retain content for up to 30 days, with approved retention controls for eligible customers.
Best for
- Developers maintaining existing Speech API integrations
- Publishers generating narration before the retirement date
- Teams planning a move to OpenAI Realtime speech
Limitations
The listed Speech models are deprecated and scheduled to shut down on 6 January 2027. The recommended replacement uses Realtime rather than the same Speech endpoint. Voice selection and controls differ by model. Custom voices require approval and recorded speaker consent. Users must be told the voice is AI-generated.
OpenAI Speech API features
Narration from written scripts
Convert supplied text into spoken audio for articles, learning content and application responses while the retiring service remains available.
Directed vocal delivery
GPT-4o Mini TTS accepts style and delivery instructions; TTS-1 and HD do not support that parameter.
Stock voices and languages
The current Speech reference lists 13 built-in voices. Older models have a smaller set; voices are optimised for English with multilingual synthesis.
Audio files and streaming
Choose MP3, Opus, AAC, FLAC, WAV or PCM. Mini TTS supports SSE streaming; older models use audio streaming.
Approved custom voices
Eligible customers can create audio-sample voices with recorded speaker consent. Prompt-designed voices are limited to Live.
Published migration timeline
The January 2027 shutdown applies to the existing Speech models; OpenAI recommends Realtime 2.1 Mini for migration.
OpenAI Text to Speech pricing
Paid plans. Usage-based pricing. Visit OpenAI Text to Speech for full plan details and current offers.
Example plans and usage rates; additional charges may apply.
TTS-1 — deprecated
USD 15 / 1,000,000 characters
TTS-1
Plan details
Per million input characters; scheduled removal 6 January 2027.
TTS-1 HD — deprecated
USD 30 / 1,000,000 characters
TTS-1 HD
Plan details
Per million input characters; scheduled removal 6 January 2027.
Mini TTS text input — deprecated
USD 0.6 / 1,000,000 input tokens
GPT-4o Mini TTS
Plan details
Per million text input tokens; generated audio charged separately; scheduled removal 6 January 2027.
Technical specifications
Voice generation
| Feature | OpenAI Text to Speech |
|---|---|
| Speech languages | Multilingual speech with voices optimised for English. Voice selection differs between Mini TTS and the older TTS-1 models.Deprecated models; retirement 6 January 2027 |
| Voice and pronunciation controls | Delivery instructions work with GPT-4o Mini TTS, not TTS-1 or HD. Speech speed accepts 0.25–4×; built-in voice selection varies by model.Deprecated; model-specific instruction support |
| Voice cloning | Audio-sample custom voices require customer eligibility, separate recorded consent and a matching speaker reference of at most 30 seconds. Maximum 20 voices per organisation; prompt-designed voices are Live-only.Approved customers; recorded speaker consent |
| Audio exports and streaming | MP3, Opus, AAC, FLAC, WAV or raw PCM. Mini TTS supports SSE or audio streaming; TTS-1 and HD do not support SSE.Deprecated models; output and streaming selection |
| Commercial-use terms | Clearly disclose to end users that the voice is AI-generated. Custom voice creation requires speaker consent and the applicable supplemental agreement.Required disclosure; custom voice permission |
| Output sample rate (Hz) | 24000Raw 16-bit little-endian PCM at 24 kHz; not a maximum across every encoding |
| API access | SupportedAPI account and authentication; models retire 6 January 2027 |
OpenAI Text to Speech alternatives
These services offer other routes for hosted narration and application speech.
Google Cloud Text-to-Speech
Chirp stock voices and Gemini delivery prompts, with character or token billing by model.
Compare with OpenAI Text to SpeechAmazon Polly
AWS speech synthesis with engine-specific voices, SSML and pronunciation lexicons.
Compare with OpenAI Text to SpeechCartesia
Sonic streaming voices and paid-plan custom voice creation through a managed speech API.
Compare with OpenAI Text to SpeechOpenAI Text to Speech FAQs
When will OpenAI’s current text-to-speech models shut down?
OpenAI deprecated TTS-1, TTS-1 HD and the March and December 2025 GPT-4o Mini TTS releases on 1 October 2026. Their scheduled API removal date is 6 January 2027. OpenAI recommends GPT-Realtime-2.1 Mini through the Realtime API.
How much does OpenAI Text to Speech cost?
TTS-1 costs USD 15 per million input characters and HD costs 30. GPT-4o Mini TTS charges 0.60 per million text input tokens plus 12 per million audio output tokens. These are the retiring models’ published rates; tokens and characters are different billing units.
Is Realtime 2.1 Mini a direct replacement for the Speech endpoint?
It is OpenAI’s recommended migration model, but it uses the Realtime API with different integration and billing. Published standard rates are USD 0.60 per million text input tokens, 2.40 per million text output tokens, 10 per million audio input tokens and 20 per million audio output tokens. Charges depend on the session’s inputs and outputs.
Can OpenAI Text to Speech use a custom voice?
Audio-sample custom voices are limited to eligible customers and require recorded consent from the same speaker. Voice creation is API-based; samples must be no longer than 30 seconds. Prompt-designed voices work only in Live. End users must receive clear disclosure that the audio is AI-generated.
Is this your product? Claim this page to update your listing.