Skip to main content
POST
JavaScript

Authorizations

Authorization
string
header
required

Authentication header containing API key (find it in dashboard). The format is "Bearer YOUR_API_KEY"

Body

application/json
response_engine
object
required

The Response Engine to attach to the agent. It is used to generate responses for the agent. You need to create a Response Engine first before attaching it to an agent.

Example:
voice_id
string
required

Unique voice id used for the agent. Find list of available voices and their preview in Dashboard.

Example:

"retell-Cimo"

agent_name
string | null

The name of the agent. Only used for your own reference.

Example:

"Jarvis"

version_description
string | null

Optional description of the agent version. Used for your own reference and documentation.

Example:

"Customer support agent for handling product inquiries"

version_title
string | null

Optional title of the agent version. Used for your own reference.

Example:

"Production hotfix"

voice_model
enum<string> | null

Select the voice model used for the selected voice. Each provider has a set of available voice models. Set to null to remove voice model selection, and default ones will apply. Check out dashboard for more details of each voice model.

Available options:
eleven_flash_v2,
eleven_flash_v2_5,
eleven_multilingual_v2,
eleven_v3,
eleven_v4_turbo,
sonic-3,
sonic-3-latest,
sonic-3.5,
sonic-3.6,
tts-1,
gpt-4o-mini-tts,
speech-02-turbo,
speech-2.8-turbo,
s1,
s2-pro,
s2.1-pro,
inworld-tts-2,
inworld-tts-2-flash,
null
fallback_voice_ids
string[] | null

When TTS provider for the selected voice is experiencing outages, we would use fallback voices listed here for the agent. Voice id and the fallback voice ids must be from different TTS providers. The system would go through the list in order, if the first one in the list is also having outage, it would use the next one. Set to null to remove voice fallback for the agent.

Example:
voice_temperature
number

Controls how stable the voice is. Value ranging from [0,2]. Lower value means more stable, and higher value means more variant speech generation. Check the dashboard to see what provider supports this feature. If unset, default value 1 will apply.

Example:

1

voice_speed
number

Controls speed of voice. Value ranging from [0.5,2]. Lower value means slower speech, while higher value means faster speech rate. If unset, default value 1 will apply.

Required range: 0.5 <= x <= 2
Example:

1

enable_dynamic_voice_speed
boolean

If set to true, will enable dynamic voice speed adjustment based on the user's speech rate and conversation context. If unset, default value false will apply.

Example:

true

enable_dynamic_responsiveness
boolean

If set to true, the agent will dynamically adjust how quickly it responds based on the user's speech rate and past turn-taking behavior in the call. If unset, default value false will apply.

Example:

true

volume
number

If set, will control the volume of the agent. Value ranging from [0,2]. Lower value means quieter agent speech, while higher value means louder agent speech. If unset, default value 1 will apply.

Example:

1

enable_expressive_mode
boolean

Master toggle for expressive mode. When true, the agent may add expressive voice tags to the audio it generates. Only applicable for platform voices. If unset, defaults to false.

Example:

true

expressive_emotion_tags
enum<string>[]

The expressive voice tags Retell pre-teaches the model to use when enable_expressive_mode is true. Custom tags defined in the system prompt are still allowed. If empty, the agent follows general expressive guidance without a fixed tag set.

Available options:
empathetic,
excited,
happy,
curious,
surprised,
sigh,
clear throat,
pause,
long pause,
emphasis
Example:
expressive_mode_prompt
string | null

Custom expressive voice guidance to use instead of the default Retell expressive prompt when enable_expressive_mode is true. If omitted or blank, the default expressive prompt will be used.

Example:

"Use [sigh] for thoughtful pauses and [excited] for good news."

responsiveness
number

Controls how responsive is the agent. Value ranging from [0,1]. Lower value means less responsive agent (wait more, respond slower), while higher value means faster exchanges (respond when it can). If unset, default value 1 will apply.

Required range: 0 <= x <= 1
Example:

1

interruption_sensitivity
number

Controls how sensitive the agent is to user interruptions. Value ranging from [0,1]. Lower value means it will take longer / more words for user to interrupt agent, while higher value means it's easier for user to interrupt agent. If unset, default value 1 will apply. When this is set to 0, agent would never be interrupted.

Required range: 0 <= x <= 1
Example:

1

enable_backchannel
boolean

Controls whether the agent would backchannel (agent interjects the speaker with phrases like "yeah", "uh-huh" to signify interest and engagement). Backchannel when enabled tends to show up more in longer user utterances. If not set, agent will not backchannel.

Example:

true

backchannel_frequency
number

Only applicable when enable_backchannel is true. Controls how often the agent would backchannel when a backchannel is possible. Value ranging from [0,1]. Lower value means less frequent backchannel, while higher value means more frequent backchannel. If unset, default value 0.8 will apply.

Example:

0.9

backchannel_words
string[] | null

Only applicable when enable_backchannel is true. A list of words that the agent would use as backchannel. If not set, default backchannel words will apply. Check out backchannel default words for more details. Note that certain voices do not work too well with certain words, so it's recommended to experiment before adding any words.

Example:
reminder_trigger_ms
number

If set (in milliseconds), will trigger a reminder to the agent to speak if the user has been silent for the specified duration after some agent speech. Must be a positive number. If unset, default value of 10000 ms (10 s) will apply.

Example:

10000

reminder_max_count
integer

If set, controls how many times agent would remind user when user is unresponsive. Must be a non negative integer. If unset, default value of 1 will apply (remind once). Set to 0 to disable agent from reminding.

Example:

2

ambient_sound
enum<string> | null

If set, will add ambient environment sound to the call to make experience more realistic. Currently supports the following options:

Available options:
coffee-shop,
convention-hall,
summer-outdoor,
mountain-outdoor,
static-noise,
call-center,
null
ambient_sound_volume
number

If set, will control the volume of the ambient sound. Value ranging from [0,2]. Lower value means quieter ambient sound, while higher value means louder ambient sound. If unset, default value 1 will apply.

Example:

1

language

Specifies what language(s) the agent will operate in. Accepts either a single locale (e.g. en-US) or an array of locales for multilingual agents (e.g. ["en-US","es-ES"]). The scalar value multi is deprecated but still accepted as a scalar, and is stored and returned as the ten locales it used to mean. It must not appear inside the array form. Send an explicit locale array instead. If unset, defaults to en-US.

Available options:
en-US,
en-IN,
en-GB,
en-AU,
en-NZ,
de-DE,
es-ES,
es-419,
hi-IN,
fr-FR,
fr-CA,
ja-JP,
pt-PT,
pt-BR,
zh-CN,
ru-RU,
it-IT,
ko-KR,
nl-NL,
nl-BE,
pl-PL,
tr-TR,
vi-VN,
ro-RO,
bg-BG,
ca-ES,
th-TH,
da-DK,
fi-FI,
el-GR,
hu-HU,
id-ID,
no-NO,
sk-SK,
sv-SE,
lt-LT,
lv-LV,
cs-CZ,
ms-MY,
af-ZA,
ar-SA,
az-AZ,
bs-BA,
cy-GB,
fa-IR,
fil-PH,
gl-ES,
he-IL,
hr-HR,
hy-AM,
is-IS,
kk-KZ,
kn-IN,
mk-MK,
mr-IN,
ne-NP,
sl-SI,
sr-RS,
sw-KE,
ta-IN,
ur-IN,
yue-CN,
uk-UA
Example:

"en-US"

webhook_url
string | null

The webhook for agent to listen to call events. See what events it would get at webhook doc. If set, will binds webhook events for this agent to the specified url, and will ignore the account level webhook for this agent. Set to null to remove webhook url from this agent.

Example:

"https://webhook-url-here"

webhook_events
enum<string>[] | null

Which webhook events this agent should receive. If not set, defaults to call_started, call_ended, call_analyzed.

Available options:
call_started,
call_ended,
call_analyzed,
transcript_updated,
transfer_started,
transfer_bridged,
transfer_cancelled,
transfer_ended
webhook_timeout_ms
integer

The timeout for the webhook in milliseconds. If not set, default value of 10000 will apply.

Example:

10000

boosted_keywords
string[] | null

Provide a customized list of keywords to bias the transcriber model, so that these words are more likely to get transcribed. Commonly used for names, brands, street, etc. Entries may reference dynamic variables with {{variable}} syntax.

Example:
contact_memory_config
object

Contact memory settings for phone calls and SMS chats. Creating an agent defaults enable_update to false and enable_read to true. Updates only change the supplied flags; omitted flags stay unchanged and an empty object has no effect. Set a flag to false to disable it. The configuration cannot be cleared. Existing agents without this configuration have both disabled.

data_storage_setting
enum<string>

Granular setting to manage how Retell stores sensitive data (transcripts, recordings, logs, etc.). This replaces the deprecated opt_out_sensitive_data_storage field.

  • everything: Store all data including transcripts, recordings, and logs.
  • everything_except_pii: Store data without PII when PII is detected.
  • basic_attributes_only: Store only basic attributes; no transcripts/recordings/logs. If not set, default value of "everything" will apply.
Available options:
everything,
everything_except_pii,
basic_attributes_only
Example:

"everything"

data_storage_retention_days
integer | null

Number of days to retain call/chat data before automatic deletion. Must be between 1 and 730 days. If not set, data is retained forever (no automatic deletion).

Required range: 1 <= x <= 730
Example:

30

opt_in_signed_url
boolean

Whether this agent opts in for signed URLs for public logs and recordings. When enabled, the generated URLs will include security signatures that restrict access and automatically expire after 24 hours.

Example:

true

signed_url_expiration_ms
integer | null

The expiration time for the signed url in milliseconds. Only applicable when opt_in_signed_url is true. If not set, default value of 86400000 (24 hours) will apply.

Example:

86400000

pronunciation_dictionary
object[] | null

A list of words / phrases and their pronunciation to be used to guide the audio synthesize for consistent pronunciation. Check the dashboard to see what provider supports this feature. Set to null to remove pronunciation dictionary from this agent.

enable_dnc_detection
boolean

If set to true, the agent recognizes requests to stop calling or contacting the user, confirms once, and on a clear yes ends the call with disconnection reason user_requested_dnc and sets do_not_call to true on the contact for the user's phone number. If unset, default value false will apply.

Example:

false

end_call_after_silence_ms
integer

If users stay silent for a period after agent speech, end the call. The minimum value allowed is 10,000 ms (10 s). By default, this is set to 600000 (10 min).

Example:

600000

max_call_duration_ms
integer

Maximum allowed length for the call, will force end the call if reached. The minimum value allowed is 60,000 ms (1 min), and maximum value allowed is 7,200,000 (2 hours). By default, this is set to 3,600,000 (1 hour).

Example:

3600000

voicemail_option
object | null

If this option is set, the call will try to detect voicemail in the first 3 minutes of the call. Actions defined (hangup, or leave a message) will be applied when the voicemail is detected. Set this to null to disable voicemail detection.

Example:
ivr_option
object | null

If this option is set, the call will try to detect IVR in the first 3 minutes of the call. Actions defined will be applied when the IVR is detected. Set this to null to disable IVR detection.

Example:
call_screening_option
object | null

If this option is set, the agent prompt will include call screen handling instructions for identity and call purpose questions. Set this to null to disable call screen prompt instructions.

post_call_analysis_data
object[] | null

Post call analysis data to extract from the call. This data will augment the pre-defined variables extracted in the call analysis. This will be available after the call ends.

Post-call analysis item (custom data or voice preset). Use for voice agent post_call_analysis_data; validates only call presets (call_summary, call_successful, user_sentiment).

post_call_analysis_model
enum<string> | null

The model to use for post call analysis. Default to gpt-5.6-terra.

Available options:
gpt-4.1,
gpt-4.1-mini,
gpt-4.1-nano,
gpt-5,
gpt-5-mini,
gpt-5-nano,
gpt-5.1,
gpt-5.2,
gpt-5.4,
gpt-5.4-mini,
gpt-5.4-nano,
gpt-5.5,
gpt-5.6-terra,
gpt-5.6-luna,
gpt-6-astra,
gpt-6-sol,
gpt-6.1-sol,
gpt-6-luna,
claude-4.5-sonnet,
claude-4.6-sonnet,
claude-5-opus,
claude-5.5-opus,
claude-5-sonnet,
claude-5.5-sonnet,
claude-4.5-haiku,
gemini-3.0-flash,
gemini-3.1-flash-lite,
gemini-3.5-flash,
gemini-3.5-flash-lite,
gemini-3.6-flash,
gemini-3.7-flash,
gemini-3.8-flash,
null
Example:

"gpt-4.1-mini"

begin_message_delay_ms
integer

If set, will delay the first message by the specified amount of milliseconds, so that it gives user more time to prepare to take the call. Valid range is [0, 5000]. If not set or set to 0, agent will speak immediately. Only applicable when agent speaks first.

Example:

1000

ring_duration_ms
integer

If set, the phone ringing will last for the specified amount of milliseconds. This applies for both outbound call ringtime, and call transfer ringtime. Default to 30000 (30 s). Valid range is [5000, 300000].

Required range: 5000 <= x <= 300000
Example:

30000

stt_mode
enum<string>

If set, determines whether speech to text should focus on latency or accuracy. Default to fast mode. When set to custom, custom_stt_config must be provided.

Available options:
fast,
accurate,
custom
Example:

"fast"

custom_stt_config
object | null

Custom STT configuration. Only used when stt_mode is set to custom.

vocab_specialization
enum<string>

If set, determines the vocabulary set to use for transcription. This setting only applies for English agents, for non English agent, this setting is a no-op. Default to general.

Available options:
general,
medical
Example:

"general"

allow_user_dtmf
boolean

If set to true, DTMF input will be accepted and processed. If false, any DTMF input will be ignored. Default to true.

Example:

true

allow_dtmf_interruption
boolean

If set to true, DTMF input will interrupt the agent even when interruption_sensitivity is 0. Can be overridden per conversation or subagent node. Default to false.

Example:

false

user_dtmf_options
object | null
denoising_mode
enum<string>

If set, determines what denoising mode to use. Use "no-denoise" to bypass all audio denoising. Default to noise-cancellation.

Available options:
no-denoise,
noise-cancellation,
noise-and-background-speech-cancellation
Example:

"noise-cancellation"

pii_config
object

Configuration for PII scrubbing from transcripts and recordings.

guardrail_config
object

Configuration for guardrail checks to detect and prevent prohibited topics in agent output and user input.

handbook_config
object

Toggle behavior presets on/off to influence agent response style and behaviors.

timezone
string | null

IANA timezone for the agent (e.g. America/New_York). Defaults to America/Los_Angeles if not set.

Example:

"America/New_York"

pre_session_tools
object[] | null

Integration (Agent Functions) tools run as a dependency graph during session setup, before the agent's first message. Outputs are injected as dynamic variables. On calls the graph gets one minute unless an outbound caller will be dialed after setup, in which case it gets five minutes. Past that session initialization continues and any remaining tools finish in the background, so their outputs no longer reach the agent's prompt. Set to null to clear.

Maximum array length: 15

A pre-session tool (app / custom / code, discriminated by type) plus its dependency edges. The tool's name is the depends_on handle and must be unique across the agent's session tools. Parameters may set a constant value or a description for the LLM to infer the value. Pre-session inference uses the agent prompt and dynamic variables.

post_session_tools
object[] | null

Integration (Agent Functions) tools run as a dependency graph during teardown, after post-call analysis. Each tool can be gated by a condition. On calls the graph is stopped after five minutes so teardown can finish. Set to null to clear.

Maximum array length: 15

A post-session tool with an optional condition. Parameter inference uses the agent prompt, dynamic variables, and conversation. Same tools as PreSessionTool plus send_sms, which is only supported on phone calls — the session is over, so the speak settings are ignored and an inferred sms_content is written by the LLM from the transcript at teardown.

Response

Successfully created a new agent.

agent_id
string
required

Unique id of agent.

Example:

"oBeDLoLOeuAbiuaMFXRtDOLriTJ5tSxD"

version
integer
required

Version of the agent.

Example:

0

response_engine
object
required

The Response Engine to attach to the agent. It is used to generate responses for the agent. You need to create a Response Engine first before attaching it to an agent.

Example:
voice_id
string
required

Unique voice id used for the agent. Find list of available voices and their preview in Dashboard.

Example:

"retell-Cimo"

last_modification_timestamp
integer
required

Last modification timestamp (milliseconds since epoch). Either the time of last update or creation if no updates available.

Example:

1703413636133

base_version
integer | null

Version that this draft was based on. Null for initial versions.

Example:

12

assigned_tags
string[]

Tags assigned to this agent version. Preferred tag is listed first.

is_published
boolean

Whether the agent is published.

Example:

false

agent_name
string | null

The name of the agent. Only used for your own reference.

Example:

"Jarvis"

version_description
string | null

Optional description of the agent version. Used for your own reference and documentation.

Example:

"Customer support agent for handling product inquiries"

version_title
string | null

Optional title of the agent version. Used for your own reference.

Example:

"Production hotfix"

voice_model
enum<string> | null

Select the voice model used for the selected voice. Each provider has a set of available voice models. Set to null to remove voice model selection, and default ones will apply. Check out dashboard for more details of each voice model.

Available options:
eleven_flash_v2,
eleven_flash_v2_5,
eleven_multilingual_v2,
eleven_v3,
eleven_v4_turbo,
sonic-3,
sonic-3-latest,
sonic-3.5,
sonic-3.6,
tts-1,
gpt-4o-mini-tts,
speech-02-turbo,
speech-2.8-turbo,
s1,
s2-pro,
s2.1-pro,
inworld-tts-2,
inworld-tts-2-flash,
null
fallback_voice_ids
string[] | null

When TTS provider for the selected voice is experiencing outages, we would use fallback voices listed here for the agent. Voice id and the fallback voice ids must be from different TTS providers. The system would go through the list in order, if the first one in the list is also having outage, it would use the next one. Set to null to remove voice fallback for the agent.

Example:
voice_temperature
number

Controls how stable the voice is. Value ranging from [0,2]. Lower value means more stable, and higher value means more variant speech generation. Check the dashboard to see what provider supports this feature. If unset, default value 1 will apply.

Example:

1

voice_speed
number

Controls speed of voice. Value ranging from [0.5,2]. Lower value means slower speech, while higher value means faster speech rate. If unset, default value 1 will apply.

Required range: 0.5 <= x <= 2
Example:

1

enable_dynamic_voice_speed
boolean

If set to true, will enable dynamic voice speed adjustment based on the user's speech rate and conversation context. If unset, default value false will apply.

Example:

true

enable_dynamic_responsiveness
boolean

If set to true, the agent will dynamically adjust how quickly it responds based on the user's speech rate and past turn-taking behavior in the call. If unset, default value false will apply.

Example:

true

volume
number

If set, will control the volume of the agent. Value ranging from [0,2]. Lower value means quieter agent speech, while higher value means louder agent speech. If unset, default value 1 will apply.

Example:

1

enable_expressive_mode
boolean

Master toggle for expressive mode. When true, the agent may add expressive voice tags to the audio it generates. Only applicable for platform voices. If unset, defaults to false.

Example:

true

expressive_emotion_tags
enum<string>[]

The expressive voice tags Retell pre-teaches the model to use when enable_expressive_mode is true. Custom tags defined in the system prompt are still allowed. If empty, the agent follows general expressive guidance without a fixed tag set.

Available options:
empathetic,
excited,
happy,
curious,
surprised,
sigh,
clear throat,
pause,
long pause,
emphasis
Example:
expressive_mode_prompt
string | null

Custom expressive voice guidance to use instead of the default Retell expressive prompt when enable_expressive_mode is true. If omitted or blank, the default expressive prompt will be used.

Example:

"Use [sigh] for thoughtful pauses and [excited] for good news."

responsiveness
number

Controls how responsive is the agent. Value ranging from [0,1]. Lower value means less responsive agent (wait more, respond slower), while higher value means faster exchanges (respond when it can). If unset, default value 1 will apply.

Required range: 0 <= x <= 1
Example:

1

interruption_sensitivity
number

Controls how sensitive the agent is to user interruptions. Value ranging from [0,1]. Lower value means it will take longer / more words for user to interrupt agent, while higher value means it's easier for user to interrupt agent. If unset, default value 1 will apply. When this is set to 0, agent would never be interrupted.

Required range: 0 <= x <= 1
Example:

1

enable_backchannel
boolean

Controls whether the agent would backchannel (agent interjects the speaker with phrases like "yeah", "uh-huh" to signify interest and engagement). Backchannel when enabled tends to show up more in longer user utterances. If not set, agent will not backchannel.

Example:

true

backchannel_frequency
number

Only applicable when enable_backchannel is true. Controls how often the agent would backchannel when a backchannel is possible. Value ranging from [0,1]. Lower value means less frequent backchannel, while higher value means more frequent backchannel. If unset, default value 0.8 will apply.

Example:

0.9

backchannel_words
string[] | null

Only applicable when enable_backchannel is true. A list of words that the agent would use as backchannel. If not set, default backchannel words will apply. Check out backchannel default words for more details. Note that certain voices do not work too well with certain words, so it's recommended to experiment before adding any words.

Example:
reminder_trigger_ms
number

If set (in milliseconds), will trigger a reminder to the agent to speak if the user has been silent for the specified duration after some agent speech. Must be a positive number. If unset, default value of 10000 ms (10 s) will apply.

Example:

10000

reminder_max_count
integer

If set, controls how many times agent would remind user when user is unresponsive. Must be a non negative integer. If unset, default value of 1 will apply (remind once). Set to 0 to disable agent from reminding.

Example:

2

ambient_sound
enum<string> | null

If set, will add ambient environment sound to the call to make experience more realistic. Currently supports the following options:

Available options:
coffee-shop,
convention-hall,
summer-outdoor,
mountain-outdoor,
static-noise,
call-center,
null
ambient_sound_volume
number

If set, will control the volume of the ambient sound. Value ranging from [0,2]. Lower value means quieter ambient sound, while higher value means louder ambient sound. If unset, default value 1 will apply.

Example:

1

language

Specifies what language(s) the agent will operate in. Accepts either a single locale (e.g. en-US) or an array of locales for multilingual agents (e.g. ["en-US","es-ES"]). The scalar value multi is deprecated but still accepted as a scalar, and is stored and returned as the ten locales it used to mean. It must not appear inside the array form. Send an explicit locale array instead. If unset, defaults to en-US.

Available options:
en-US,
en-IN,
en-GB,
en-AU,
en-NZ,
de-DE,
es-ES,
es-419,
hi-IN,
fr-FR,
fr-CA,
ja-JP,
pt-PT,
pt-BR,
zh-CN,
ru-RU,
it-IT,
ko-KR,
nl-NL,
nl-BE,
pl-PL,
tr-TR,
vi-VN,
ro-RO,
bg-BG,
ca-ES,
th-TH,
da-DK,
fi-FI,
el-GR,
hu-HU,
id-ID,
no-NO,
sk-SK,
sv-SE,
lt-LT,
lv-LV,
cs-CZ,
ms-MY,
af-ZA,
ar-SA,
az-AZ,
bs-BA,
cy-GB,
fa-IR,
fil-PH,
gl-ES,
he-IL,
hr-HR,
hy-AM,
is-IS,
kk-KZ,
kn-IN,
mk-MK,
mr-IN,
ne-NP,
sl-SI,
sr-RS,
sw-KE,
ta-IN,
ur-IN,
yue-CN,
uk-UA
Example:

"en-US"

webhook_url
string | null

The webhook for agent to listen to call events. See what events it would get at webhook doc. If set, will binds webhook events for this agent to the specified url, and will ignore the account level webhook for this agent. Set to null to remove webhook url from this agent.

Example:

"https://webhook-url-here"

webhook_events
enum<string>[] | null

Which webhook events this agent should receive. If not set, defaults to call_started, call_ended, call_analyzed.

Available options:
call_started,
call_ended,
call_analyzed,
transcript_updated,
transfer_started,
transfer_bridged,
transfer_cancelled,
transfer_ended
webhook_timeout_ms
integer

The timeout for the webhook in milliseconds. If not set, default value of 10000 will apply.

Example:

10000

boosted_keywords
string[] | null

Provide a customized list of keywords to bias the transcriber model, so that these words are more likely to get transcribed. Commonly used for names, brands, street, etc. Entries may reference dynamic variables with {{variable}} syntax.

Example:
contact_memory_config
object

Contact memory settings for phone calls and SMS chats. Creating an agent defaults enable_update to false and enable_read to true. Updates only change the supplied flags; omitted flags stay unchanged and an empty object has no effect. Set a flag to false to disable it. The configuration cannot be cleared. Existing agents without this configuration have both disabled.

data_storage_setting
enum<string>

Granular setting to manage how Retell stores sensitive data (transcripts, recordings, logs, etc.). This replaces the deprecated opt_out_sensitive_data_storage field.

  • everything: Store all data including transcripts, recordings, and logs.
  • everything_except_pii: Store data without PII when PII is detected.
  • basic_attributes_only: Store only basic attributes; no transcripts/recordings/logs. If not set, default value of "everything" will apply.
Available options:
everything,
everything_except_pii,
basic_attributes_only
Example:

"everything"

data_storage_retention_days
integer | null

Number of days to retain call/chat data before automatic deletion. Must be between 1 and 730 days. If not set, data is retained forever (no automatic deletion).

Required range: 1 <= x <= 730
Example:

30

opt_in_signed_url
boolean

Whether this agent opts in for signed URLs for public logs and recordings. When enabled, the generated URLs will include security signatures that restrict access and automatically expire after 24 hours.

Example:

true

signed_url_expiration_ms
integer | null

The expiration time for the signed url in milliseconds. Only applicable when opt_in_signed_url is true. If not set, default value of 86400000 (24 hours) will apply.

Example:

86400000

pronunciation_dictionary
object[] | null

A list of words / phrases and their pronunciation to be used to guide the audio synthesize for consistent pronunciation. Check the dashboard to see what provider supports this feature. Set to null to remove pronunciation dictionary from this agent.

enable_dnc_detection
boolean

If set to true, the agent recognizes requests to stop calling or contacting the user, confirms once, and on a clear yes ends the call with disconnection reason user_requested_dnc and sets do_not_call to true on the contact for the user's phone number. If unset, default value false will apply.

Example:

false

end_call_after_silence_ms
integer

If users stay silent for a period after agent speech, end the call. The minimum value allowed is 10,000 ms (10 s). By default, this is set to 600000 (10 min).

Example:

600000

max_call_duration_ms
integer

Maximum allowed length for the call, will force end the call if reached. The minimum value allowed is 60,000 ms (1 min), and maximum value allowed is 7,200,000 (2 hours). By default, this is set to 3,600,000 (1 hour).

Example:

3600000

voicemail_option
object | null

If this option is set, the call will try to detect voicemail in the first 3 minutes of the call. Actions defined (hangup, or leave a message) will be applied when the voicemail is detected. Set this to null to disable voicemail detection.

Example:
ivr_option
object | null

If this option is set, the call will try to detect IVR in the first 3 minutes of the call. Actions defined will be applied when the IVR is detected. Set this to null to disable IVR detection.

Example:
call_screening_option
object | null

If this option is set, the agent prompt will include call screen handling instructions for identity and call purpose questions. Set this to null to disable call screen prompt instructions.

post_call_analysis_data
object[] | null

Post call analysis data to extract from the call. This data will augment the pre-defined variables extracted in the call analysis. This will be available after the call ends.

Post-call analysis item (custom data or voice preset). Use for voice agent post_call_analysis_data; validates only call presets (call_summary, call_successful, user_sentiment).

post_call_analysis_model
enum<string> | null

The model to use for post call analysis. Default to gpt-5.6-terra.

Available options:
gpt-4.1,
gpt-4.1-mini,
gpt-4.1-nano,
gpt-5,
gpt-5-mini,
gpt-5-nano,
gpt-5.1,
gpt-5.2,
gpt-5.4,
gpt-5.4-mini,
gpt-5.4-nano,
gpt-5.5,
gpt-5.6-terra,
gpt-5.6-luna,
gpt-6-astra,
gpt-6-sol,
gpt-6.1-sol,
gpt-6-luna,
claude-4.5-sonnet,
claude-4.6-sonnet,
claude-5-opus,
claude-5.5-opus,
claude-5-sonnet,
claude-5.5-sonnet,
claude-4.5-haiku,
gemini-3.0-flash,
gemini-3.1-flash-lite,
gemini-3.5-flash,
gemini-3.5-flash-lite,
gemini-3.6-flash,
gemini-3.7-flash,
gemini-3.8-flash,
null
Example:

"gpt-4.1-mini"

begin_message_delay_ms
integer

If set, will delay the first message by the specified amount of milliseconds, so that it gives user more time to prepare to take the call. Valid range is [0, 5000]. If not set or set to 0, agent will speak immediately. Only applicable when agent speaks first.

Example:

1000

ring_duration_ms
integer

If set, the phone ringing will last for the specified amount of milliseconds. This applies for both outbound call ringtime, and call transfer ringtime. Default to 30000 (30 s). Valid range is [5000, 300000].

Required range: 5000 <= x <= 300000
Example:

30000

stt_mode
enum<string>

If set, determines whether speech to text should focus on latency or accuracy. Default to fast mode. When set to custom, custom_stt_config must be provided.

Available options:
fast,
accurate,
custom
Example:

"fast"

custom_stt_config
object | null

Custom STT configuration. Only used when stt_mode is set to custom.

vocab_specialization
enum<string>

If set, determines the vocabulary set to use for transcription. This setting only applies for English agents, for non English agent, this setting is a no-op. Default to general.

Available options:
general,
medical
Example:

"general"

allow_user_dtmf
boolean

If set to true, DTMF input will be accepted and processed. If false, any DTMF input will be ignored. Default to true.

Example:

true

allow_dtmf_interruption
boolean

If set to true, DTMF input will interrupt the agent even when interruption_sensitivity is 0. Can be overridden per conversation or subagent node. Default to false.

Example:

false

user_dtmf_options
object | null
denoising_mode
enum<string>

If set, determines what denoising mode to use. Use "no-denoise" to bypass all audio denoising. Default to noise-cancellation.

Available options:
no-denoise,
noise-cancellation,
noise-and-background-speech-cancellation
Example:

"noise-cancellation"

pii_config
object

Configuration for PII scrubbing from transcripts and recordings.

guardrail_config
object

Configuration for guardrail checks to detect and prevent prohibited topics in agent output and user input.

handbook_config
object

Toggle behavior presets on/off to influence agent response style and behaviors.

timezone
string | null

IANA timezone for the agent (e.g. America/New_York). Defaults to America/Los_Angeles if not set.

Example:

"America/New_York"

pre_session_tools
object[] | null

Integration (Agent Functions) tools run as a dependency graph during session setup, before the agent's first message. Outputs are injected as dynamic variables. On calls the graph gets one minute unless an outbound caller will be dialed after setup, in which case it gets five minutes. Past that session initialization continues and any remaining tools finish in the background, so their outputs no longer reach the agent's prompt. Set to null to clear.

Maximum array length: 15

A pre-session tool (app / custom / code, discriminated by type) plus its dependency edges. The tool's name is the depends_on handle and must be unique across the agent's session tools. Parameters may set a constant value or a description for the LLM to infer the value. Pre-session inference uses the agent prompt and dynamic variables.

post_session_tools
object[] | null

Integration (Agent Functions) tools run as a dependency graph during teardown, after post-call analysis. Each tool can be gated by a condition. On calls the graph is stopped after five minutes so teardown can finish. Set to null to clear.

Maximum array length: 15

A post-session tool with an optional condition. Parameter inference uses the agent prompt, dynamic variables, and conversation. Same tools as PreSessionTool plus send_sms, which is only supported on phone calls — the session is over, so the speak settings are ignored and an inferred sms_content is written by the LLM from the transcript at teardown.