> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tavus.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Text-to-Speech (TTS)

> Configure how your PAL sounds. Set a Tavus Voice by ID, or bring a voice from your own TTS provider account.

The TTS layer decides how your PAL sounds. Setting a [Tavus Voice](/sections/conversational-video-interface/voices) with `voice_id` is usually all it needs; the remaining fields are for bringing a voice from your own provider account.

Set **`layers.tts`** when you [Create PAL](/api-reference/pals/create-pal) or update a PAL. For how a PAL fits together, see [PAL overview](/sections/conversational-video-interface/pal/overview). For languages and locale-oriented setup, see [Language support](/sections/conversational-video-interface/language-support).

## Configuring the TTS Layer

Define the TTS layer under the `layers.tts` object. The snippets below show only the **`tts`** object for readability; in a full PAL payload it is nested under **`layers`** (see [Example configuration](#example-configuration)).

Below are the parameters available:

### 1. `voice_id`

A [Tavus Voice](/sections/conversational-video-interface/voices), referenced by ID. Tavus picks the provider and model that are the best fit for the language(s) of the conversation, so this is usually the entire TTS layer:

```json theme={null}
"tts": {
  "voice_id": "v0a1b2c3d4e5f"
}
```

Create a voice in [PAL Maker](https://maker.tavus.io/dev/voices/create) or with the [Voices API](/api-reference/voices/create-voice), then use the same `voice_id` on as many PALs and faces as you like. A `voice_id` here takes precedence over a face's `default_voice_id`.

### 2. `external_voice_id`

Use this when the voice lives in your own Cartesia, ElevenLabs, or Azure account. Tavus does not manage it: you pick the provider and model, and the same provider voice speaks every language in the conversation, carrying its own accent.

To find supported voice IDs, refer to the provider's documentation:

* <a href="https://docs.cartesia.ai/api-reference/voices/list" target="_blank">Cartesia</a>
* <a href="https://elevenlabs.io/docs/api-reference/voices/search" target="_blank">ElevenLabs</a>
* <a href="https://learn.microsoft.com/en-us/azure/ai-services/speech-service/language-support?tabs=tts" target="_blank">Azure</a> (fallback only; e.g. `en-US-JennyNeural`) - if using Azure, the voice determines the accent, not the language: any voice speaks the conversation's language, carrying that voice's own accent. For a natural result, choose a voice whose locale matches your target language, or use an Azure `*MultilingualNeural` voice.

```json theme={null}
"tts": {
  "external_voice_id": "external-voice-id"
}
```

<Note>
  You can use any publicly accessible custom voice from ElevenLabs or Cartesia without the provider's API key. If the custom voice is private, you still need to use the provider's API key.
</Note>

<Note>
  `external_voice_id` takes precedence over `voice_id` and over a face's `default_voice_id`. The two voice fields are mutually exclusive on one PAL, so send one. `voice_id` also cannot be combined with your own TTS `api_key`, because a Tavus Voice lives in Tavus's provider account rather than yours.
</Note>

### 3. `tts_engine`

Specifies the TTS engine.

* **Options**: `tavus-auto` **(default)**, `cartesia`, `elevenlabs`. Also `azure`, only as a fallback when your language is not otherwise supported.

<Tip>
  `tavus-auto` is the default. Use it unless you need a specific provider or voice. It selects the best TTS engine and model for each conversation; the underlying provider may change over time. For deterministic behavior, set an explicit provider and model.
</Tip>

<Note>
  If you use `tavus-auto`, you do not need to specify any other parameters in the `tts` layer.
</Note>

<Note>
  Use `azure` only if you need a language that is not supported by the default engines. Prefer `tavus-auto`, `cartesia`, or `elevenlabs` whenever your language is already covered. See [Additional language support via Azure](/sections/conversational-video-interface/language-support#additional-language-support-via-azure).
</Note>

```json theme={null}
"tts": {
  "tts_engine": "cartesia"
}
```

### 4. `api_key`

Authenticates requests to your selected third-party TTS provider. You can obtain an API key from one of the following:

<Warning>
  For Cartesia and ElevenLabs, only required when using private (non-public) voices. If you are using Azure as a fallback for an unsupported language, an API key is **required**.
</Warning>

* <a href="https://play.cartesia.ai/keys" target="_blank">Cartesia</a>
* <a href="https://elevenlabs.io/app/settings/api-keys" target="_blank">ElevenLabs</a> - if using pronunciation dictionaries, the key must have the `pronunciation_dictionaries_write` scope (or full account access). See <a href="https://elevenlabs.io/docs/api-reference/service-accounts/api-keys/create" target="_blank">ElevenLabs API key scopes</a>.
* <a href="https://portal.azure.com" target="_blank">Azure</a> (fallback only) - required when using Azure. Use your own Azure Speech resource key; the resource must be in the **East US** region (Tavus synthesizes via `eastus`; a key from another region returns an authentication error). Any standard neural voice available in East US works; only Custom Neural Voices need to be deployed in your resource.

```json theme={null}
"tts": {
  "api_key": "your-api-key"
}
```

### 5. `tts_model_name`

Model name used by the TTS engine. Refer to:

* <a href="https://docs.cartesia.ai/2025-04-16/build-with-cartesia/models" target="_blank">Cartesia</a>
* <a href="https://elevenlabs.io/docs/models" target="_blank">ElevenLabs</a>

<Warning>
  `tts_model_name` is not supported when `tts_engine` is `azure`. If you are using Azure as a language fallback, omit this field.
</Warning>

```json theme={null}
"tts": {
  "tts_model_name": "sonic-3"
}
```

### 6. `voice_settings`

Optional object for controlling speed, volume, and similar effects. **Which approach you use depends on your TTS engine and model:**

| Engine | Model | Approach |
| - | - | - |
| ElevenLabs | All models | `voice_settings` in PAL config |
| Cartesia | sonic-2 | `voice_settings` in PAL config |
| Cartesia | sonic-3 | **Either** `voice_settings` (global, set once per conversation) **or** prompt the LLM in `system_prompt` to output [Cartesia SSML tags](https://docs.cartesia.ai/build-with-cartesia/sonic-3/ssml-tags) for dynamic control. Not both. |

<Warning>
  **Cartesia sonic-3:** If you use `voice_settings` for speed/volume, those settings apply globally for the whole conversation and you cannot use SSML tags for dynamic, per-phrase control. If you want dynamic control, omit `voice_settings` and have the LLM output SSML tags instead. See [Cartesia volume, speed, and emotion](https://docs.cartesia.ai/build-with-cartesia/sonic-3/volume-speed-emotion).
</Warning>

**ElevenLabs (all models):** Set parameters in the `voice_settings` object:

| Parameter | ElevenLabs |
| - | - |
| `speed` | Range `0.7` to `1.2` (`0.7` = slowest, `1.2` = fastest) |
| `stability` | Range `0.0` to `1.0` (`0.0` = variable, `1.0` = stable) |
| `similarity_boost` | Range `0.0` to `1.0` (`0.0` = creative, `1.0` = original) |
| `style` | Range `0.0` to `1.0` (`0.0` = neutral, `1.0` = exaggerated) |
| `use_speaker_boost` | Boolean (enhances speaker similarity) |

<Note>
  See <a href="https://elevenlabs.io/docs/api-reference/voices/settings/get" target="_blank">ElevenLabs Voice Settings</a> for details.
</Note>

**Cartesia sonic-2:** Use the `voice_settings` object (e.g. `speed`, `emotion`). SSML tags are not used for sonic-2.

**Cartesia sonic-3:** You can use **either** of these, but not both:

* **`voice_settings`** - We accept speed/volume params for sonic-3. They apply **globally**, set once per conversation. Use this when you want a single default speed and volume for the entire conversation. Using `voice_settings` prevents dynamic SSML control.
* **SSML in LLM output** - Omit `voice_settings` for speed/volume and instead add instructions to your `system_prompt` so the LLM outputs [Cartesia SSML tags](https://docs.cartesia.ai/build-with-cartesia/sonic-3/ssml-tags) in its responses. This gives you dynamic, per-phrase control. See [Cartesia volume, speed, and emotion](https://docs.cartesia.ai/build-with-cartesia/sonic-3/volume-speed-emotion).

Emotion control is separate; see [Emotion Control](/sections/conversational-video-interface/quickstart/emotional-expression).

**Example: system prompt for Cartesia sonic-3 (dynamic speed and volume)**

If you are **not** using `voice_settings` for sonic-3, add instructions like this to your `system_prompt` so the LLM outputs Cartesia SSML tags:

```
When you want to emphasize a word or phrase, use Cartesia SSML tags for speed and volume:
- To slow down: <speed level="0.8">phrase</speed>
- To speed up: <speed level="1.2">phrase</speed>
- To speak louder: <volume level="1.2">phrase</volume>
- To speak more quietly: <volume level="0.8">phrase</volume>
You can combine tags, e.g. <speed level="0.9"><volume level="1.1">important point</volume></speed>.
Only use these tags when it improves clarity or emphasis; keep most of your response in plain text.
```

**Example: voice\_settings (ElevenLabs, Cartesia sonic-2, or Cartesia sonic-3 global)**

```json theme={null}
"tts": {
  "voice_settings": {
    "speed": 0.9
  }
}
```

For sonic-3, this sets global speed once per conversation; for sonic-2 and ElevenLabs, it applies as configured.

## Example Configuration

Below are example PALs, starting with the common case:

<CodeGroup>
  ```json Tavus Voice theme={null}
  {
    "pal_name": "AI Presenter",
    "system_prompt": "You are a friendly and informative video host.",
    "pipeline_mode": "full",
    "context": "You're delivering updates in a conversational tone.",
    "default_face_id": "rc9cff32ceba",
    "layers": {
      "tts": {
        "voice_id": "v0a1b2c3d4e5f"
      }
    }
  }
  ```

  ```json Cartesia theme={null}
  {
    "pal_name": "AI Presenter",
    "system_prompt": "You are a friendly and informative video host.",
    "pipeline_mode": "full",
    "context": "You're delivering updates in a conversational tone.",
    "default_face_id": "rc9cff32ceba",
    "layers": {
      "tts": {
        "tts_engine": "cartesia",
        "api_key": "your-api-key",
        "external_voice_id": "external-voice-id",
        "tts_model_name": "sonic-3"
      }
    }
  }
  ```

  ```json ElevenLabs theme={null}
  {
    "pal_name": "Narrator",
    "system_prompt": "You narrate long stories with clarity and consistency.",
    "pipeline_mode": "full",
    "context": "You're reading a fictional audiobook.",
    "default_face_id": "rc9cff32ceba",
    "layers": {
      "tts": {
        "tts_engine": "elevenlabs",
        "api_key": "your-api-key",
        "external_voice_id": "elevenlabs-voice-id",
        "voice_settings": {
          "speed": 0.9
        },
        "tts_model_name": "eleven_turbo_v2_5"
      }
    }
  }
  ```

  ```json Azure (fallback only) theme={null}
  {
    "pal_name": "Azure fallback PAL",
    "system_prompt": "You are a friendly host.",
    "pipeline_mode": "full",
    "default_face_id": "rc9cff32ceba",
    "layers": {
      "tts": {
        "tts_engine": "azure",
        "api_key": "your-azure-speech-key",
        "external_voice_id": "en-US-JennyNeural"
      }
    }
  }
  ```
</CodeGroup>

<Note>
  The Azure example above is only for cases where your target language is not supported by the default engines. Prefer Cartesia or ElevenLabs (or leave `tts_engine` unset for `tavus-auto`) whenever possible.
</Note>

<Note>
  Refer to [Create PAL](/api-reference/pals/create-pal) for a complete list of supported fields.
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.