Select a language model

The model dropdown, with the Suggested pill on the first recommended model highlighted.
Choose a voice

The voice control in the agent toolbar.

The Select Voice dialog. Each card carries a preview button and the voice ID.
Decide who speaks first
- User speaks first: the agent stays silent until the caller says something.
- AI speaks first: the agent opens the conversation. A second dropdown then chooses how:
- Dynamic message: the agent generates its own opener each call, based on your prompt.
- Custom message: the agent reads a fixed message you type in the field below.

The Welcome Message setting with the second dropdown open on Dynamic message.
More settings
The panels on the right of the agent editor cover everything else. None are required to make a first call.Write the prompt
Configure Knowledge Base
Configure Speech Settings
- Background sound: select a background sound that plays throughout the whole call to mimic an environment like a call center, making the conversation more humanlike and engaging.
- Responsiveness: how responsive the agent is. Set it lower if you want the agent to wait longer before responding, which can be useful when talking to folks like the elderly. The lower the value, the more wait time is added before the agent responds. You can also check “Dynamically adjust based on user input” to let the agent automatically tune its response timing during the call. When enabled, the agent observes how quickly the user speaks and adjusts accordingly — slower speakers get more patient response timing, while faster speakers get quicker responses.
- Interruption Sensitivity: how fast the agent gets interrupted by user interruptions. Set it lower if you want the agent to be more resilient to background speech.
- Backchanneling: Set up how often and what words the agent uses to acknowledge users.
- Boosted Keywords: Provides some biases towards certain words, making it easier to get recognized. Common ones are brand names, people’s names, etc.
- Speech Normalization: convert entities like date, currency, numbers into plain words, which can help prevent issues where audio generated was not pronouncing those right.
- Reminder frequency: how often the agent will remind the user when user is inactive.
- Pronunciation: set a pronunciation guide for specific words.
Configure Call Settings
- Voicemail related settings: set up voicemail detection and what to do when voicemail is detected. See more at Handle Voicemail.
- End call on silence: set up if user is inactive for a certain amount of time, the call will be ended.
- Call duration: set up maximum duration of the call.
- Pause before speaking: For the beginning of the call, if agent speaks first, it will wait for the configured duration before speaking, useful to handle scenarios when user is still picking up the phone.
Configure Realtime Transcription Settings
- Denoising Mode: how aggressively background noise is filtered out of the caller’s audio.
- Transcription Mode: which speech-to-text model transcribes the call.
- Endpointing: how long the agent waits before deciding the caller has finished a turn.
- Vocabulary Specialization: bias transcription toward a domain’s terms, so words like drug or product names come through correctly.
Configure Post-Call Data Extraction
Configure Security & Fallback Settings
- Data Storage and PII Settings: opt out of storing recordings, transcripts, or personally identifiable information.
- Secure URLs: serve recordings and logs from signed links that expire, anywhere from 1 minute to 7 days.
- Fallback Voice: the voice to switch to if your primary voice provider fails mid-call, so the call continues instead of dropping.
- Guardrails: limits on what the agent is allowed to say.
- Default Dynamic Variables: fallback values for prompt variables that aren’t supplied when the call starts.
Configure Webhook Settings
Configure MCPs

