AssemblyAI
AssemblyAI: speech-to-text APIs for recorded and live audio
AssemblyAI provides speech-to-text APIs for recorded files, live audio and short synchronous clips. Developers can add speaker labels, word timestamps, subtitle exports and speech analysis to their own products. A browser playground helps explore the models; production use centers on API integration, with language coverage and pricing varying by model.
Free trial availableAccount required

- Access
- Free trial available
- Usage pricing
- USD 0.15 / 60 audio minutesUniversal-2 — recorded
- Available on
- Web, API
Transcription infrastructure for applications and voice products
AssemblyAI is designed for developers who need transcription inside another application. Recorded audio goes through an asynchronous API; live audio uses a streaming connection; Sync returns short-clip transcription in a single request. SDKs and the browser playground make the service easier to explore, but a production captioning, meeting or interview product still needs an integration around the API.
Results can include word-level timestamps and speaker-labelled utterances. Completed recorded transcripts have subtitle and text export endpoints. Speech Understanding adds analysis such as summaries, translation and speaker identification; LLM Gateway provides a separate way to run language-model requests. These features have their own billing, so a base transcription rate does not represent every analysis operation.
The standard service processes audio in the cloud. Enterprise self-hosting is available for specific Realtime and Sync APIs, with deployment assistance. Paid-account owners and admins can opt out of AssemblyAI’s model-improvement program for subsequent requests and configure retention. Current recorded-transcript deletion begins after thirty days by default; uploaded audio has a shorter lifecycle. Settings, API type and deletion-processing delays affect retention, and billing/logging metadata is kept separately.
Best for
- Developers adding transcription to applications
- Teams building live captions or voice interfaces
- Services processing recorded interviews at scale
Limitations
This is a developer API rather than a finished collaborative transcript editor. Language coverage, add-ons and limits vary by endpoint. Streaming bills open connections including idle time; multichannel jobs multiply recorded usage. Trial credits do not include a recurring free allowance, and paid training opt-out applies to subsequent requests.
AssemblyAI speech-to-text and developer features
Recorded-file transcription API
Process long recordings asynchronously with model-specific language support and word timestamps.
Live streaming transcription
Build live captions and voice interfaces over WebSocket; open connection time is billable.
Short-clip Sync API
Return transcription in a single request for brief audio, with different limits from the long-file API.
Speaker-labelled results
Add diarization to separate voices; identifying actual names or roles is a distinct analysis feature.
Subtitle and transcript endpoints
Retrieve structured JSON, plain text or SRT/VTT subtitles from completed recorded jobs.
Account data controls
Paid owners/admins can configure future-request training opt-out and retention; enterprise self-hosting covers specific APIs.
AssemblyAI pricing
Free trial available. Usage-based pricing. Visit AssemblyAI for full plan details and current offers.
Example plans and usage rates; additional charges may apply.
Universal-2 — recorded
USD 0.15 / 60 audio minutes
Universal-2
Plan details
USD 0.15 per audio hour for recorded transcription in 99 languages. Add-ons and multichannel jobs cost extra.
Universal-3.5 Pro — recorded
USD 0.21 / 60 audio minutes
Universal-3.5 Pro
Plan details
USD 0.21 per audio hour for recorded transcription in 18 languages. Add-ons and multichannel jobs cost extra.
Recorded speaker diarization — add-on
USD 0.02 / 60 audio minutes
Recorded diarization
Plan details
Adds USD 0.02 per audio hour to recorded transcription. Realtime diarization has a different price.
Technical specifications
Transcription and meetings
| Feature | AssemblyAI |
|---|---|
| File and meeting capture | Recorded-file REST API, live WebSocket transcription and Sync for short clips. SDKs and a browser playground support integration; this is not a collaborative document editor.APIs have separate models, limits and billing |
| Transcription languages | Recorded Universal-2: 99 languages; Universal-3.5 Pro: 18. Realtime Universal-3.6 Pro: 32; Universal-Streaming Multilingual: English, Spanish, German, French, Portuguese and Italian.Recorded versus realtime models |
| Speaker-label workflow | Speaker diarization separates voices into labelled utterances; recorded add-on USD 0.02/audio hour. Speaker identification is a distinct analysis feature for names/roles.Diarization and identification are separately priced features |
| Transcript exports | JSON transcripts with words/timestamps and optional speakers; completed recorded transcripts can be exported as SRT, VTT or plain text via API.API export endpoints, not DOCX/PDF editorial files |
| Recording and usage limits | Recorded /v2/transcript supports files up to 5 GB and ten hours; /v2/upload local-file uploads are limited to 2.2 GB. Sync handles short clips separately. Streaming charges continue during idle connections.Recorded-file limits; Sync and realtime have separate limits |
| Maximum recording duration (minutes) | 600Pre-recorded /v2/transcript only; Sync and realtime have separate limits |
| API access | SupportedREST/WebSocket APIs; production integration requires developer setup |
AssemblyAI alternatives
For a finished transcription workspace, these products provide file editing and collaboration without building a custom API integration.
Sonix
A browser transcript editor with multilingual subtitles and workspace billing.
Compare with AssemblyAIRev
A transcript-review workspace with AI analysis and a separate human service.
Compare with AssemblyAITrint
Shared editorial transcripts, Story Builder and newsroom production exports.
Compare with AssemblyAIAssemblyAI FAQs
Is AssemblyAI free?
AssemblyAI offers USD 50 in trial credits. Continued API usage requires paid credits or an account agreement; it is not a recurring free transcription allowance. LLM Gateway usage has separate terms.
How much does AssemblyAI transcription cost?
Recorded Universal-2 is USD 0.15 per audio hour and Universal-3.5 Pro USD 0.21 per audio hour before add-ons. Streaming rates begin at USD 0.15 per hour of open connection time, including idle time.
Can AssemblyAI export subtitles?
Yes. Completed recorded transcripts can be retrieved as SRT or VTT through API endpoints, alongside JSON and plain text. These are developer endpoints rather than a collaborative caption editor.
Can AssemblyAI use audio to train its models?
It may use eligible data where its contract permits. Paid owners/admins can opt out for subsequent requests; EU-server and BAA conditions also affect eligibility. This policy is separate from LLM-provider training and production retention.
Is this your product? Claim this page to update your listing.