Python library for Gemini Live TTS with voice characters, and safe math expression evaluation.
- uv — fast Python package manager
GEMINI_API_KEYenvironment variable
Install uv if you don't have it:
curl -LsSf https://astral.sh/uv/install.sh | shRun setup first — this creates a .venv, installs dependencies, and symlinks
gstts so you can use it from anywhere (/usr/local/bin if writable, otherwise ~/.local/bin):
./run.sh setuprun.sh Unified entry point (setup, test, gstts, etc.)
python/
gemini_live_tools/
gemini_live_api.py GeminiLiveAPI, character definitions, PCM/WAV helpers
math_eval.py Safe AST-based math expression evaluator
gstts.py Gemini Streaming TTS CLI
tests/ Python tests
js/
tts-audio-player.js Streaming WAV/PCM audio player (Web Audio API)
voice-character-selector.js Drop-in voice/character picker UI widget
docs/
streaming-tts-endpoint.md FastAPI streaming endpoint guide with cancellation
| Command | Description |
|---|---|
./run.sh setup |
Create .venv, install deps, symlink gstts to PATH |
./run.sh update |
Upgrade all dependencies in .venv |
./run.sh test |
Run Python tests (extra args passed to pytest, e.g. -v) |
./run.sh test-player |
Open TTS audio player test page in browser |
./run.sh gstts [args] |
Run Gemini Streaming TTS CLI |
./run.sh help |
Show available commands |
# First time setup (required before first use)
./run.sh setup
# After setup, use `gstts` from anywhere
gstts "Hello world" # read text aloud (uses default character)
gstts # interactive: pick a character, generate & play a greetinggstts -lc # list available characters
gstts -lv # list available Gemini voices
gstts "Hello" -c narrator # use a specific character
gstts "Hello" -c narrator -v Charon # character + voice override
gstts "Hello" -s "speak slowly and dramatically" # add a style instructionIn the interactive picker, Enter selects a character and proceeds to TTS.
Space sets the character as default in ~/gstts_config.json and exits
(quick-select mode — no audio generated).
Text can be piped from stdin — uses the default character from config:
cat README.md | gstts # read a file aloud
echo "Hello world" | gstts # pipe text
pbpaste | gstts # read clipboard
cat article.txt | gstts --summarize # summarize and read
cat notes.md | gstts -c narrator -p # pipe with character + prepareBy default, provided text is read as-is. Use -p to rewrite the text into
speech-friendly form via Gemini before synthesis (converts markdown, LaTeX,
abbreviations, etc. into natural speech):
gstts -p "The API returns 200 OK w/ a JSON payload incl. nested arrays"Use --summarize (or --summary) to condense text before reading it aloud.
This runs a two-step pipeline: first a separate LLM call shrinks the text to
~10% of its length, then prepare_text rewrites the summary for speech:
gstts --summarize "$(cat long-article.txt)" # summarize and read
cat research-paper.md | gstts --summarize # pipe + summarize
gstts --summarize -s "be funny" "long text..." # summarize with extra styleLow-latency streaming — audio starts playing in ~200-500ms. Bypasses sentence splitting and buffering; streams PCM directly from a single Live API session:
gstts -rt "Hello world" # realtime with default character
gstts -rt -c narrator "Hello world" # realtime with specific character
gstts -rt -p "text with markdown" # prepare first, then realtime
cat article.txt | gstts -rt # pipe with realtime playbackRun gstts --help for the full list:
gstts "text" -c narrator # use a specific character
gstts "text" -c narrator -v Charon # character + voice override
gstts "text" -s "speak slowly" # add a style instruction
gstts "text" -p # prepare text for speech (rewrite markdown, etc.)
gstts "text" --summarize # summarize before reading
gstts "text" -rt # realtime low-latency mode
gstts "text" --no-live # disable Live API (use generate_content)
gstts "text" --parallelism 4 # parallel TTS (4 concurrent chunks)
gstts "text" --output greeting.wav # save audio to file
gstts "text" --debug # verbose output
Character preference is stored in ~/gstts_config.json:
{ "character": "crisp" }When text is passed without --character, the config character is used automatically.
After picking from the menu, you're prompted to save the selection as the new default.
./run.sh setup # create .venv, install deps, symlink gstts
./run.sh test # run Python tests
./run.sh test -v # run tests with verbose output
./run.sh test-player # open TTS audio player test page in browser
./run.sh update # upgrade dependenciesPin to a tagged release:
pip install "gemini-live-tools @ git+https://github.com/ibenian/gemini-live-tools.git@v0.1.16#subdirectory=python"To find the latest tag:
git ls-remote --tags https://github.com/ibenian/gemini-live-tools.git | grep -v '\^{}' | awk -F/ '{print $3}' | sort -V | tail -1from google import genai
from gemini_live_tools import GeminiLiveAPI, ParallelTTSStatus
from gemini_live_tools import safe_eval_math, eval_math_sweep, MATH_NAMES
client = genai.Client(api_key="...")
api = GeminiLiveAPI(api_key="...", client=client)
# Single-shot TTS
prepared = api.prepare_text("Hello world", character_name="crisp")
wav = api.synthesize_wav(prepared, character_name="crisp")
# Single-shot TTS via Live API (falls back to generate_content on failure)
wav = api.synthesize_wav(prepared, character_name="crisp", use_live=True)
# Parallel streaming TTS (sync — yields one WAV chunk per sentence in order)
for chunk in api.stream_parallel_wav(prepared, parallelism=4, character_name="crisp"):
play(chunk) # play each sentence as it arrives
# Parallel streaming TTS with Live API
for chunk in api.stream_parallel_wav(prepared, parallelism=4, character_name="crisp", use_live=True):
play(chunk)
# Parallel streaming TTS (async — for FastAPI / aiohttp)
async for chunk in api.astream_parallel_wav(prepared, parallelism=4, character_name="crisp"):
yield chunk
# Realtime streaming TTS — lowest latency (~200-500ms to first audio)
# Single Live API session, yields raw PCM s16le 24kHz as it arrives
for pcm_chunk in api.stream_realtime_pcm("Hello world", character_name="crisp"):
stream.write(np.frombuffer(pcm_chunk, dtype=np.int16))
# Realtime streaming TTS (async — for FastAPI)
async for pcm_chunk in api.astream_realtime_pcm("Hello world", character_name="crisp"):
yield pcm_chunk
# Use ParallelTTSStatus standalone for your own streaming loops
status = ParallelTTSStatus(n=total_chunks)
status.start(parallelism=4)
status.mark_received(idx=0, delivery_mode="live") # L icon — received via Live API
status.mark_received(idx=1, delivery_mode="fallback") # * icon — received via generate_content
status.mark_playing(idx=0) # shows ▶ on status line
status.mark_played()
status.finish() # prints final Played N/N line
# Math eval
result, err = safe_eval_math("norm([3, 4])") # → 5.0
result, err = safe_eval_math("sin(pi/2)") # → 1.0See docs/streaming-tts-endpoint.md for a full FastAPI streaming endpoint example with client-side cancellation.
js/voice-character-selector.js is a self-contained browser widget that exposes window.GeminiVoiceCharacterSelector. It provides two components:
CharacterPicker — a searchable, grouped character palette (like a command palette). Features:
- Opens via a trigger button or
Cmd+K/Ctrl+K - Live search across character name, label, and group
- Characters organized into groups: Core, Academic, Accents, Dramatic, Musical, Fiction, etc.
- Tracks recently used characters
- Persists selection to
localStorage
setupVoiceSelect — populates a <select> element with all available Gemini voices, with optional localStorage persistence.
Copy the file into your project and include it:
<script src="/static/voice-character-selector.js"></script><button id="characterBtn">Character</button>
<select id="voiceSelect"></select>
<!-- Palette and backdrop should be direct children of body for correct positioning -->
<div id="characterPalette" class="style-palette" hidden>
<input id="characterSearch" class="style-search" type="text" placeholder="Search characters..." />
<div id="characterList" class="style-list"></div>
</div>
<div id="characterBackdrop" class="style-backdrop" hidden></div>const lib = window.GeminiVoiceCharacterSelector;
// Move palette and backdrop to body so they're never clipped by overflow/stacking contexts
document.body.appendChild(document.getElementById('characterPalette'));
document.body.appendChild(document.getElementById('characterBackdrop'));
// Populate voice <select> — returns the currently selected voice
let selectedVoice = lib.setupVoiceSelect(document.getElementById('voiceSelect'), {
storageKey: 'myAppVoice',
defaultValue: 'Charon',
});
// Wire up the character picker
let selectedCharacter = 'crisp';
const picker = new lib.CharacterPicker({
buttonEl: document.getElementById('characterBtn'),
paletteEl: document.getElementById('characterPalette'),
searchEl: document.getElementById('characterSearch'),
listEl: document.getElementById('characterList'),
backdropEl: document.getElementById('characterBackdrop'),
options: lib.CHARACTER_OPTIONS,
groupMap: lib.CHARACTER_GROUPS,
groupOrder: lib.CHARACTER_GROUP_ORDER,
storageKey: 'myAppCharacter',
recentsKey: 'myAppCharacterRecents',
defaultId: 'crisp',
hotkey: 'k', // opens palette on Cmd/Ctrl+K
onChange: (characterId) => {
selectedCharacter = characterId;
// auto-switch voice to the character's recommended default
const opt = lib.CHARACTER_OPTIONS.find(o => o.id === characterId);
if (opt?.defaultVoice) {
document.getElementById('voiceSelect').value = opt.defaultVoice;
selectedVoice = opt.defaultVoice;
}
},
});
selectedCharacter = picker.init();See CONTRIBUTING.md for how to add voice characters and more.
- Python 3.10+
GEMINI_API_KEYenvironment variable- uv — only needed for
./run.sh(dev/CLI workflow). Library consumers can use standard pip.
This software is provided for educational and informational purposes only. The authors and contributors make no representations or warranties regarding the accuracy, completeness, or suitability of this software for any particular purpose. Use is entirely at your own risk. The authors shall not be held liable for any direct, indirect, incidental, special, or consequential damages arising from the use of or inability to use this software.