Since September 30, 2026, Deepgram's Flux text-to-speech (/v2/speak) reads two kinds of marker placed inside the text you send: one that inserts a silence, and one that tells it how to pronounce a word.
Prerequisites
A Deepgram API key in DEEPGRAM_API_KEY (paste it where the examples say YOUR_API_KEY)
The Deepgram Python SDK (pip install deepgram-sdk)
Steps
Pause. Add an escaped marker such as \{pause:800ms\} where the voice should wait. Pauses run from 500 ms to 3000 ms in 100 ms steps, up to 8 per request, on batch (REST) requests:
Python
from deepgram import DeepgramClientclient = DeepgramClient(api_key="YOUR_API_KEY")text = r"Your appointment is on Tuesday at 3 PM. \{pause:800ms\} Reply YES to confirm."response = client.speak.v2.audio.generate( text=text, model="flux-haley-en", encoding="mp3", speed=0.9,)audio_bytes = b"".join(response)with open("appointment_reminder.mp3", "wb") as f: f.write(audio_bytes)
Pronunciation (Early Access). Give the word and how it should sound in IPA. This works on batch and streaming:
Know the limits. Pronunciation cannot be combined with a pause, or with a speed other than 1.0. With a pause present, speed is capped at 1.15. Combinations outside these rules are rejected with CONTROL_COMBINATION_INVALID or PAUSE_SPEED_CAP_EXCEEDED.
Test it
Run the first script and play appointment_reminder.mp3: there should be a clear gap before "Reply YES to confirm." Deepgram notes that pronunciation results can vary while the feature is in Early Access, so generate each term a few times before relying on it.