Real-time speech API platform

Deepgram

A managed speech platform with distinct speech-to-text and text-to-speech model families, streaming paths, language coverage, and usage meters.

Editorial verdict

Choose Deepgram when real-time speech is central and measured performance on the exact audio, language, model, and region justifies a specialized provider; evaluate speech recognition and generation separately.1234

Best for

  • Streaming transcription and turn-aware voice interactions
  • Teams that can benchmark domain audio and latency
  • Products needing both STT and TTS from one bounded speech provider
1234

Not ideal for

  • Teams assuming one model or price covers both modalities
  • Workloads lacking representative language and audio evaluation
  • Buyers unable to govern retention, model-improvement, and voice policy
1234
Main trade-off

Specialized streaming speech reduces integration latency work while increasing model-version, modality, add-on, data-policy, regional, and provider dependence.1234

Product boundary

Whether Deepgram's real-time speech-to-text or text-to-speech APIs fit the required latency, model, language, voice, data, and operating boundaries.

This page evaluates Deepgram's speech-to-text and text-to-speech APIs as two distinct modalities. Flux and Nova recognition paths, Aura and Flux speech-generation paths, add-ons, and Voice Agent APIs have separate capabilities and pricing; support in one modality does not imply support in the other.1234

For: Teams building live transcription or conversational voice systems that will benchmark the exact model, language, latency, and audio path

  • The production workload, language, model, provider, or deployment boundary changes
  • Current pricing, retention, data use, policy, or regional support changes
  • Measured quality, latency, reliability, or operating cost no longer fits

Why teams consider Deepgram

  • Modality separationOfficial documentation distinguishes recognition models from Aura and Flux speech generation.1234
  • Streaming focusDeepgram documents streaming recognition and low-latency speech-generation routes.1234
  • Usage visibilityOfficial pricing separates STT minutes, TTS characters, models, and add-ons.1234

Pricing

Speech-to-text is metered by audio duration and selected model or add-on; text-to-speech is metered by generated characters and selected voice model.1

Current official route

Pay as you go, Growth, or Enterprise

Verified 2026-07-27: Deepgram publishes model-specific STT minute rates and distinct Aura TTS character rates. Promotional, committed, add-on, and model-improvement-program terms must be evaluated separately.1

STT meter
Audio duration by model, mode, and optional intelligence add-ons1
TTS meter
Generated characters by Aura model1
Do not combine
STT and TTS are separate cost and capability decisions1
Pricing checked View official pricing

Deepgram vs alternatives

AssemblyAI

Choose when
Choose AssemblyAI when managed transcription quality and speech-intelligence features matter more than running the pipeline, after testing the exact model and governing session duration, retention, and training settings.
Avoid when
Teams requiring full control of model infrastructure
Compared with Deepgram
Managed transcription and intelligence reduce pipeline assembly while adding model, session, feature, retention, region, and provider dependence.56

Google Cloud Speech-to-Text

Choose when
Choose Google Cloud Speech-to-Text when Google Cloud placement, supported languages, batch economics, or service governance fit, after validating the exact V2 model, region, channel, and logging configuration.
Avoid when
Text-to-speech or general Gemini model use
Compared with Deepgram
Google Cloud integration and scale reduce platform assembly while adding recognizer, model, language, region, channel, IAM, logging, and cloud-specific cost boundaries.78

ElevenLabs

Choose when
Choose ElevenLabs when voice quality and catalog breadth are central, but only after the exact model, language, latency, character cost, output use, and voice-consent boundary have been verified.
Avoid when
Products unable to prove voice rights or consent
Compared with Deepgram
Voice quality and catalog breadth reduce voice-production work while adding character, model, catalog, consent, policy, provenance, and provider dependence.910

Cartesia

Choose when
Choose Cartesia when measured low-latency speech on the target path is decisive and the team can own API upgrades, concurrency, credit usage, voice rights, and the documented ZDR exclusions.
Avoid when
Teams unable to track model or API deprecations
Compared with Deepgram
Low-latency speech improves interaction speed while adding API-version, model, credit, concurrency, cloning, consent, retention, and provider lifecycle work.1112

Resources and sources

Official product, documentation, pricing, and policy sources

  • Deepgram pricing
    Open
  • Deepgram speech-to-text models and languages
    Open
  • Deepgram text-to-speech models and voices
    Open
  • Deepgram Flux TTS voices
    Open
  • AssemblyAI pricing
    Open
  • How AssemblyAI pricing works
    Open
  • Google Cloud Speech-to-Text
    Open
  • Google Cloud Speech-to-Text pricing
    Open
  • ElevenLabs API pricing
    Open
  • ElevenLabs text-to-speech documentation
    Open
  • Cartesia pricing
    Open
  • Cartesia voice cloning guide
    Open
  1. 1
    Deepgram pricing

    Deepgram · Accessed Official

  2. 2
    Deepgram speech-to-text models and languages

    Deepgram · Accessed Official

  3. 3
    Deepgram text-to-speech models and voices

    Deepgram · Accessed Official

  4. 4
    Deepgram Flux TTS voices

    Deepgram · Accessed Official

  5. 5
    AssemblyAI pricing

    AssemblyAI · Accessed Official

  6. 6
    How AssemblyAI pricing works

    AssemblyAI · Accessed Official

  7. 7
    Google Cloud Speech-to-Text

    Google Cloud · Accessed Official

  8. 8
    Google Cloud Speech-to-Text pricing

    Google Cloud · Accessed Official

  9. 9
    ElevenLabs API pricing

    ElevenLabs · Accessed Official

  10. 10
    ElevenLabs text-to-speech documentation

    ElevenLabs · Accessed Official

  11. 11
    Cartesia pricing

    Cartesia · Accessed Official

  12. 12
    Cartesia voice cloning guide

    Cartesia · Accessed Official