Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), background reasoning (extended thinking), asynchronous function calling, session management, ephemeral tokens, live transcription, and live translation. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).
Permissions
Files
SKILL.md
gemini-live-api-dev
Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), background reasoning (extended thinking), asynchronous function calling, session management, ephemeral tokens, live transcription, and live translation. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).
Gemini Live API Development Skill
Overview
The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses and background reasoning.
[!NOTE]
The Live API connects directly via WebSockets. For WebRTC support or simplified integration, use a partner integration.
Models
Current Models (Use These)
gemini-3.8-live — Default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Supports interleaved reasoning, asynchronous function calling by default (behavior: NON_BLOCKING), and full-session client content updates.
gemini-3.8-live-extended-thinking — High-reasoning audio-to-audio model recommended when higher background reasoning is required during live interactions. Processes background reasoning and async tool calls (behavior: NON_BLOCKING required) while streaming continuous spoken conversational fillers; lifecycle managed via interaction_status (IN_PROGRESS vs IDLE).
gemini-3.5-transcribe-live — Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD.
gemini-3.5-live-translate-preview — Real-time speech-to-speech streaming translation across 70+ languages.
Output: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate.
[!IMPORTANT]
Use send_realtime_input / sendRealtimeInput for all real-time streaming user input (audio, video, and text). On Gemini 3.8 models, send_client_content / sendClientContent is supported across the full session lifecycle with explicit roles (user or model) to inject conversation context (turn_complete=true unconditionally interrupts active generation).
[!WARNING]
Do not use media in sendRealtimeInput. Use the specific keys: audio for audio data, video for images/video frames, and text for text input.
Quick Start
Authentication
Python
from google import genai
client = genai.Client(api_key="YOUR_API_KEY")
from google.genai import types
config = types.LiveConnectConfig(
response_modalities=[types.Modality.AUDIO],
system_instruction=types.Content(
parts=[types.Part(text="You are a helpful assistant.")]
)
)
asyncwith client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
pass# Session is active
[!IMPORTANT]
A single server event can contain multiple content parts simultaneously (e.g., audio chunks and transcript). Always process all parts in each event to avoid missing content.
Python
asyncfor response in session.receive():
content = response.server_content
if content:
# Audio — process ALL parts in each eventif content.model_turn:
for part in content.model_turn.parts:
if part.inline_data:
audio_data = part.inline_data.data
# Transcriptionif content.input_transcription:
print(f"User: {content.input_transcription.text}")
if content.output_transcription:
print(f"Gemini: {content.output_transcription.text}")
# Interruptionif content.interrupted isTrue:
pass# Stop playback, clear audio queue
JavaScript
// Inside the onmessage callbackconst content = response.serverContent;
if (content?.modelTurn?.parts) {
for (const part of content.modelTurn.parts) {
if (part.inlineData) {
const audioData = part.inlineData.data; // Base64 encoded
}
}
}
if (content?.inputTranscription) console.log('User:', content.inputTranscription.text);
if (content?.outputTranscription) console.log('Gemini:', content.outputTranscription.text);
if (content?.interrupted) { /* Stop playback, clear audio queue */ }
Background Reasoning (Extended Thinking)
Use gemini-3.8-live-extended-thinking when your voice agent must evaluate complex data, plan multiple steps, or handle long-running tools. The model speaks natural conversational fillers (e.g. "Checking flight options now...") while executing asynchronous tools in the background.
Non-blocking tools: All function declarations must set behavior="NON_BLOCKING". Synchronous blocking mode is not supported and returns an error.
Lifecycle tracking (interaction_status): Do not rely on turn_complete=True alone to detect turn completion. Monitor message.interaction_status (Python) / message.interactionStatus (JS):
"IN_PROGRESS": Server is reasoning, speaking conversational fillers, or waiting for async tool responses.
"IDLE": Server has completed all background reasoning and tool calls; session is ready for user input.
See references/migration.md and the Thinking in Live API Guide for complete Python and JavaScript implementation examples.
Live Translation (Gemini Live Translate)
The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide.
Model
gemini-3.5-live-translate-preview — The recommended translation model for all Live Translate use cases.
Configuration (TranslationConfig)
To enable translation, specify a TranslationConfig object inside your live session setup:
Python SDK: Configure the connection using translation_config on LiveConnectConfig:
Live Streaming Transcription (Gemini Live Transcribe)
The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook.
Model
gemini-3.5-transcribe-live
Modes
smart: cleans up filler words, resolves inline self-corrections, and structures formatting.
Response modality — Only TEXTorAUDIO per session, not both. Native audio models output audio (response_modalities=["AUDIO"]); enable output_audio_transcription if you need text transcripts.
Audio-only session — 15 min without compression
Audio+video session — 2 min without compression
Connection lifetime — ~10 min (use session resumption)
For step-by-step migration checklists and protocol deltas when upgrading from gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, or gemini-2.0-flash-live-001 to Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking, read references/migration.md.
Best Practices
Use headphones when testing mic audio to prevent echo/self-interruption
Enable context window compression for sessions longer than 15 minutes
Implement session resumption to handle connection resets gracefully
Use ephemeral tokens for client-side deployments — never expose API keys in browsers
Use send_realtime_input for real-time user input (audio, video, text). Use send_client_content with explicit user/model roles to inject context turns mid-stream
Send audioStreamEnd / audio_stream_end (Hybrid VAD) when the mic is paused or user finishes speaking
Clear audio playback queues on interruption signals (interrupted: true)
Process all parts in each server event — events can contain multiple content parts
Monitor interaction_status (IN_PROGRESS vs IDLE) when using gemini-3.8-live-extended-thinking rather than relying on turn_complete alone
Documentation Lookup
When MCP is Installed (Preferred)
If the search_docs tool (from the Google MCP server) is available, use it as your only documentation source:
Call search_docs with your query
Read the returned documentation
Trust MCP results as source of truth for API details — they are always up-to-date.
[!IMPORTANT]
When MCP tools are present, never fetch URLs manually. MCP provides up-to-date, indexed documentation that is more accurate and token-efficient than URL fetching.
When MCP is NOT Installed (Fallback Only)
If no MCP documentation tools are available, fetch from the official docs index:
Migration & Upgrading Guide — step-by-step checklists and code examples for Gemini 3.8 Live and Extended Thinking
Supported Languages
The Live API supports 70 languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, and many more. Native audio models automatically detect and switch languages.