Model Configurations

Voice Agents need AI Models to work, like LLM (Large Language Model), TTS (Voice) and STT (Transcriber). You can use any of your faviourite providers with Paladin Platform to run your Voice Agent.

How Model Configuration Works#

Paladin uses a two-level configuration system for AI models:

  1. Global configuration — A single set of model settings (LLM, TTS, STT) that applies to all agents by default.
  2. Agent-level overrides — Optional per-agent settings that override the global configuration for specific services.

If no overrides are set for an agent, it uses the global configuration as-is.

Global Configuration#

The global configuration is the default model setup shared throughout all your agents. Paladin ships with its own models by default. When your workspace is provisioned, you receive model credits to start with.

To configure the global models, go to Model Configurations in your dashboard: Open Model Configurations in your Paladin dashboard.

Model Configuration
Model Configuration

From here you can configure each service:

ServiceWhat it does
LLMThe language model that generates responses (e.g., OpenAI GPT-4.1, Anthropic Claude)
TTS (Voice)The text-to-speech model that converts responses to spoken audio (e.g., ElevenLabs, Cartesia)
STT (Transcriber)The speech-to-text model that transcribes user speech (e.g., Deepgram, AssemblyAI)
RealtimeA single speech-to-speech model that handles LLM, TTS, and STT in one (e.g., Gemini Live)

Select a provider from the dropdown and configure the API key, model, and any provider-specific settings. For Paladin's own models, see Service Keys for instructions on creating Service Keys.

Agent-Level Model Overrides#

You can override the global model configuration for any individual agent. This is useful when different agents have different requirements — for example, a customer support agent might use a faster, cheaper LLM while a sales agent uses a more capable one.

Configuring overrides#

  1. Open the agent you want to customize.
  2. Go to Settings in the agent detail page.
  3. Select the Model Overrides tab.
  4. You will see tabs for each service: LLM, Voice (TTS), and Transcriber (STT).
  5. Toggle Override on for the service you want to change.
  6. Configure the provider, model, and other settings as needed.
  7. Save your changes.

Selective overrides#

Each service can be toggled independently. When an override is off for a service, the agent inherits the global setting for that service. When an override is on, the agent uses the override setting instead.

LLM OverrideTTS OverrideSTT OverrideResult
OffOffOffAgent uses global config for all services
OnOffOffAgent uses custom LLM, global TTS and STT
OffOnOffAgent uses global LLM and STT, custom TTS
OnOnOnAgent uses custom config for all services

For example, if you only want to change the voice for a specific agent:

  1. Leave the LLM and Transcriber overrides off.
  2. Toggle the Voice override on.
  3. Select a different TTS provider or voice.
  4. The agent will use your custom voice while still using the global LLM and STT.

Realtime mode override#

You can also switch an individual agent to use a Realtime provider (such as Gemini Live) even if the global configuration uses standard LLM + TTS + STT. Toggle the Realtime switch in the Model Overrides tab, then configure the realtime provider, model, and voice.

Gemini 3.1 Live#

Gemini 3.1 Live is Google's realtime multimodal API that handles both LLM and voice in a single model. Instead of configuring separate LLM, TTS, and STT services, Gemini Live acts as an all-in-one realtime provider — it processes speech input, generates a response, and speaks it back, all over a single streaming connection.

Paladin supports Gemini 3.1 Live as a Realtime provider. The default model is gemini-3.1-flash-live-preview.

Available Voices#

You can choose from the following built-in voices:

VoiceDescription
PuckDefault voice
Charon
Kore
Fenrir
Aoede

Getting a Gemini API Key#

To use Gemini 3.1 Live with Paladin, you need a Google Gemini API key. Follow these steps:

  1. Go to Google AI Studio.
  2. Sign in with your Google account.
  3. Click on Get API Key in the left sidebar.
  4. Click Create API Key.
  5. Select an existing Google Cloud project or create a new one.
  6. Copy the generated API key and store it securely.

Configuring Gemini 3.1 Live in Paladin#

  1. Go to Model Configurations in your Paladin dashboard (https://app.paladin.northmanngrp.com/configurations/inference-providers https://app.paladin.northmanngrp.com/model-configurations ).
  2. Under the Realtime section, select google_realtime as the provider.
  3. Paste your Gemini API key.
  4. Select the model (gemini-3.1-flash-live-preview is available by default, or you can enter a model name manually).
  5. Choose a voice from the dropdown (default is Puck).
  6. Select the language (currently en is supported).

Gemini Live on Vertex AI#

If you want to run Gemini Live through your own Google Cloud project — for billing consolidation, VPC controls, regional residency, or enterprise IAM — Paladin also supports Gemini Live via Vertex AI as a separate provider (google_vertex_realtime). The default model is google/gemini-live-2.5-flash-native-audio.

Unlike Google AI Studio (which uses a single Gemini API key), Vertex AI authenticates with a service account belonging to your Google Cloud project.

Prerequisites#

  1. A Google Cloud project with billing enabled.
  2. The Vertex AI API enabled on that project:
bash
   gcloud services enable aiplatform.googleapis.com --project=YOUR_PROJECT_ID
  1. A service account with the Vertex AI User role (roles/aiplatform.user) on the project:
bash
   gcloud projects add-iam-policy-binding YOUR_PROJECT_ID \
     --member="serviceAccount:YOUR_SA@YOUR_PROJECT_ID.iam.gserviceaccount.com" \
     --role="roles/aiplatform.user"
  1. A JSON key for that service account (P12 keys are not supported).

Creating the service account key#

  1. In the GCP Console, go to IAM & Admin → Service Accounts.
  2. Pick an existing service account (or create a new one).
  3. Open the Keys tab → Add Key → Create new key.
  4. Choose JSON as the key type and click Create.
  5. The key file will download to your computer — store it securely and treat it as a secret.

Configuring Vertex AI Realtime in Paladin#

  1. Go to Model Configurations in your Paladin dashboard.
  2. Enable the Realtime toggle.
  3. Under the Realtime section, select google_vertex_realtime as the provider.
  4. Fill in the fields:
FieldWhat to put in
ModelVertex publisher/model id, e.g. google/gemini-live-2.5-flash-native-audio
VoiceOne of the built-in voices (Puck, Charon, Kore, Fenrir, Aoede)
LanguageBCP-47 code (e.g. en-US)
Project IdThe project_id value from your service-account JSON
LocationGCP region where the model is available (e.g. us-east4)
CredentialsPaste the entire contents of the service-account JSON file
API KeyLeave blank — Vertex AI does not use API keys
  1. Save the configuration.