Text to speech for CircuitPython using the SVOX Pico
engine by SVOX AG, released under the Apache License 2.0. It speaks English (en-US) at 16 kHz,
handles numbers, abbreviations and dates on the board, and plays through audiomixer. The
engine and voice ship in the library, so it runs on stock CircuitPython with no core module.
say() waits until the text has been spoken. say(text, wait=False) returns right away;
call update() from your main loop to keep speaking.
This library depends on:
- Adafruit CircuitPython 11 or later
The engine is a precompiled native module, picotts_native.armv7emsp.mpy, for the RP2350.
The engine needs about 1.1 MB of RAM and the en-US voice (en_US_ta.bin and
en_US_lh0_sg.bin, in the library) another 1.43 MB, about 2.5 MB in all, so it needs a
board with PSRAM:
| Board | Status |
|---|---|
| Fruit Jam (RP2350) | Working |
| Feather RP2350 with 8 MB PSRAM | Runs, audio check pending |
This library does not run on Blinka.
Please ensure all dependencies are available on the CircuitPython filesystem. This is easily achieved by downloading the Adafruit library and driver bundle or individual libraries can be installed using circup.
Make sure that you have circup installed in your Python environment.
Install it with the following command if necessary:
pip3 install circupWith circup installed and your CircuitPython device connected use the
following command to install:
circup install adafruit_picottsOr the following command to update an existing version:
circup updateimport audiobusio
import board
import adafruit_picotts as speech
audio = audiobusio.I2SOut(board.A0, board.A1, board.A2)
tts = speech.TTS(audio)
tts.say("Hello from Circuit Python.")See examples/picotts_fruitjam.py for the Fruit Jam's TLV320 DAC.
Rendering takes about 0.6 times real time on the RP2350, but the engine analyzes each sentence
before any of it is spoken, which takes from under a second to several seconds for a long
sentence. TTS fills a buffer before it starts playing, so a sentence of up to 15 s plays
without a break.
TTS() loads the voice and allocates its buffers once, taking the largest block the heap
allows for the speech buffer (up to 30 s of audio, 960 KB). Create it early, before other large
allocations.
src/svox is the SVOX Pico engine source. src/Makefile builds it with
py/dynruntime.mk into adafruit_picotts/picotts_native.<arch>.mpy:
cd src
make MPY_DIR=path/to/circuitpython ARCH=armv7emspUse the CircuitPython tree of the version the file will run on.
The speech engine and en-US voice are SVOX Pico, Copyright (C) 2008-2009 SVOX AG, licensed under the Apache License 2.0. This library is MIT licensed.
API documentation for this library can be found on Read the Docs.
For information on building library documentation, please check out this guide.
Contributions are welcome! Please read our Code of Conduct before contributing to help this project stay welcoming.