Deepgram Aura
Deepgram Aura: multilingual speech for apps
Deepgram Aura is a text-to-speech API with seven languages, named voices, streamed audio and character-based pricing for apps, narration and voice agents.
Paid plansAccount required
- Access
- Paid access
- Usage pricing
- USD 0.03 / 1,000 charactersAura-2 Pay As You Go
- Available on
- API, Web
Deepgram Aura for narration and voice applications
Deepgram Aura turns text into generated speech through an API. It gives developers named voices and audio settings for narrated app content, notifications and voice-enabled services. Integration happens in code rather than through a voiceover editing timeline.
Aura-2 offers voices in seven languages. A voice’s model identifier selects its language and delivery characteristics. English and Spanish voices also support speed and IPA pronunciation controls, which can be used together for names, terminology and reading pace.
REST and WebSockets serve different audio needs. REST can produce compressed audio files and begin playback before a response is complete. WebSockets accept text progressively and return raw audio for interactive applications. Their encoding and sample-rate options differ.
Aura-2 pricing uses generated characters, making the speech-generation cost straightforward to estimate from the text. Introductory API credit provides a starting balance, while Growth uses prepaid annual usage. Voice Agent, speech recognition and Flux TTS have separate prices.
Aura-2 does not include Flux TTS’s explicit pause tags or expressivity parameter. Its strengths are named multilingual voices, speech controls in English and Spanish, and API delivery. Creators who want an editing studio or voice cloning can also explore ElevenLabs and Cartesia.
Best for
- Developers building voice-enabled applications
- Teams generating narration through an API
- Voice-agent projects needing streamed audio
Limitations
Aura-2 speed and pronunciation controls apply to English and Spanish. Explicit pause tags and expressivity belong to Flux TTS; REST and WebSocket audio formats also differ.
Deepgram Aura features
Seven languages and named voices
Aura-2 covers English, Spanish, Dutch, French, German, Italian and Japanese. Each named voice has its own language, accent and delivery characteristics; a model such as aura-2-thalia-en selects a specific English voice.
Speed and pronunciation controls
Aura-2 supports speed and pronunciation adjustments in English and Spanish. Speed ranges from 0.7 to 1.5, and pronunciation overrides use IPA notation. Both controls can be combined; pronunciation-controlled input is limited to 2,000 characters.
Audio files through REST
REST output includes WAV with Linear16 audio, MP3, Ogg Opus, FLAC and AAC, plus mu-law and A-law options. Linear16/WAV at 24 kHz is the default; container, sample-rate and bitrate choices depend on the encoding.
Raw audio through WebSockets
The WebSocket API accepts text for interactive speech generation and returns raw Linear16, mu-law or A-law audio. Compressed formats such as MP3 are REST options, so file playback and live audio connections use different output settings.
Playback before generation finishes
A REST response can start playing as audio bytes arrive, without waiting for the complete file. WebSockets provide a separate streaming connection for applications that send text progressively.
Concurrent requests for voice applications
Pay As You Go allows up to 15 concurrent Aura or Aura-2 REST requests and 45 WebSocket connections per project. Limits depend on the model, plan and regional endpoint; these figures describe Aura, not Flux TTS.
Deepgram Aura pricing
Aura-2 Pay As You Go costs $0.030 per 1,000 generated characters, or $3 for 100,000 characters. Deepgram offers $200 in introductory API credit. Growth lists $0.027 per 1,000 characters with prepaid plans starting at $4,000 per year. Aura-1, Flux TTS, speech recognition and Voice Agent usage have separate rates.
Example plans and usage rates; additional charges may apply.
Aura-2 Pay As You Go
USD 0.03 / 1,000 characters
Aura-2
Plan details
USD 0.030 per 1,000 generated characters, checked 9 October 2026. Aura-1, Flux TTS, Voice Agent and speech recognition have separate rates.
Technical specifications
Voice generation
| Feature | Deepgram Aura |
|---|---|
| Speech languages | Aura-2 voices cover English, Spanish, Dutch, French, German, Italian and Japanese. Language and accent depend on the selected voice.Aura-2 named voices |
| Voice and pronunciation controls | Aura-2 supports speed from 0.7 to 1.5 and IPA pronunciation controls in English and Spanish. Controls can be combined; pronunciation-controlled input is limited to 2,000 characters. Explicit pause tags and expressivity are not Aura-2 features.English and Spanish; pronunciation-controlled input up to 2,000 characters |
| Audio exports and streaming | REST supports WAV/Linear16, MP3, Ogg Opus, FLAC, AAC, mu-law and A-law. WebSocket output uses raw Linear16, mu-law or A-law; format settings differ by protocol.Aura; voice and protocol availability vary |
| Output sample rate (Hz) | 24000REST default Linear16/WAV; other encodings differ |
| API access | SupportedAura; voice and protocol availability vary |
Deepgram Aura alternatives
Cartesia and ElevenLabs combine speech APIs with voice-cloning options. OpenAI Text to Speech is another developer-focused option for adding generated speech to applications.
Deepgram Aura FAQs
What is Deepgram Aura?
Aura is Deepgram’s text-to-speech model family. An application sends text and a voice model identifier to the API and receives generated speech. It suits narrated app content and voice-enabled services; it does not provide a timeline-based voiceover editor.
Which languages does Aura-2 support?
Aura-2 has voices for English, Spanish, Dutch, French, German, Italian and Japanese. Language, accent and voice characteristics are tied to the selected model. Aura-1 is a separate English-language model family.
Can I change Aura’s speed and pronunciation?
Aura-2 supports speed and IPA pronunciation controls in English and Spanish. Speed ranges from 0.7 to 1.5; Deepgram recommends 0.9 to 1.5 for Spanish. Speed and pronunciation can be combined, with a 2,000-character maximum for pronunciation-controlled input.
Does Aura support pause tags or expressivity controls?
Aura-2 does not support explicit pause tags or the expressivity parameter. Those features belong to Flux TTS. Punctuation and phrasing can still influence Aura’s delivery, but they do not set an exact pause duration.
Can I download MP3 or WAV audio?
Yes. Aura’s REST API supports MP3 and WAV output, alongside Ogg Opus, FLAC, AAC and telephony encodings. WebSocket output is raw Linear16, mu-law or A-law; it does not use the same compressed file formats as REST.
How much does Deepgram Aura cost?
As of 9 October 2026, Aura-2 Pay As You Go costs $0.030 per 1,000 generated characters. That makes 100,000 characters $3 for Aura-2 speech generation. Growth lists $0.027 per 1,000 characters with prepaid plans starting at $4,000 per year. Aura-1 and other Deepgram APIs have separate prices.
Is there a free version of Deepgram Aura?
Deepgram offers $200 in introductory API credit to try its services. Aura then uses character-based paid pricing. The credit is an introductory balance, rather than a recurring free monthly allowance.
Is Aura the same as Deepgram’s Voice Agent API?
No. Aura generates speech from text. The Voice Agent API combines components for a conversational agent and has separate pricing. Flux TTS is another speech model family with its own controls and rates; neither product’s features or costs automatically apply to Aura.
Is this your product? Claim this page to update your listing.