Local Text-to-Speech adds self-hosted text-to-speech to any Drupal site. Visitors click "Listen to this page" and hear the content read aloud by one of 54 voices across 9 languages. No cloud APIs, no subscriptions, no data leaving your server.
Features
- 54 voices across 9 languages: American English, British English, Japanese, Mandarin Chinese, French, Hindi, Spanish, Italian, Portuguese
- Language-aware voice filtering: automatically shows only voices matching the content language
- Adjustable speech speed: 0.5x to 2.0x
- Audio caching: generated files cached on disk to reduce server load
- Block and field integration: place a player site-wide or per content type
- Drush commands: test, batch generate, list voices, clear cache
- Accessibility-optimised: ARIA labels, keyboard shortcuts, semantic HTML
- Privacy-first: all processing local, only publicly viewable content converted
- Data sovereignty: no cloud, no third-party SaaS, all data stays on your server
How it works
Local TTS uses the Kokoro TTS engine (an 82M-parameter model) compiled to a native binary via Kokoros. Text is converted to phonemes by eSpeak NG, then the Kokoro ONNX model generates natural-sounding speech. The entire pipeline runs locally on your server.
You need Local TTS if
- You want to add text-to-speech to your Drupal site without a cloud service
- You want multi-language speech from a single module with no per-request costs
- You need GDPR-compliant audio where no visitor data leaves your infrastructure
- You want a privacy-first accessibility feature that works without third-party JavaScript
Installation
System dependency
eSpeak NG must be installed on the server for phoneme processing:
# macOS brew install espeak-ng # Ubuntu/Debian sudo apt-get install espeak-ng
Module installation
composer require drupal/local_tts cd web/modules/contrib/local_tts composer run build-binary drush en local_tts
The composer install step downloads model and voice data files (~337MB). The binary must be built from source (requires Rust/Cargo). After enabling, configure the eSpeak NG data path at /admin/config/media/local-tts.
Verify installation
drush local-tts:test "Hello World"Drush command reference
| Command | Description |
|---|---|
local-tts:test |
Test TTS generation with instant playback |
local-tts:read |
Read entity content aloud (with caching) |
local-tts:batch |
Batch generate audio for multiple entities |
local-tts:voices |
List available voices, optionally filtered by language |
local-tts:cache-clear |
Clear all cached audio files |
FAQ
- Does my content leave the server?
-
No. All speech generation happens locally using the bundled Kokoro binary and ONNX model files. No text or audio is sent to any external service.
- How large are the model files?
-
The model is approximately 310MB and the voice data is approximately 27MB. Files are downloaded automatically during
composer install. - Does Local TTS work with page caching?
-
Yes. Audio is loaded via AJAX when the visitor clicks the play button, so the page itself can be fully cached (including behind Varnish or a CDN). Cached audio files are served directly without regeneration.
- Does Local TTS set cookies or track users?
-
No. The module only processes the content text server-side when requested. It does not set cookies, store user IDs, or create visitor profiles.
- Can I add custom voices?
-
Not at present. Local TTS uses the voices bundled with the Kokoro model. New voices would require training a custom ONNX model.
- Module documentation: installation, configuration, voice reference, and Drush commands
- Installation guide: system dependencies, binary build, and Nginx configuration
- Available voices: all 54 voices across 9 languages with voice IDs
- FAQ: caching, privacy, load management, and troubleshooting
Project information
- Project categories: Accessibility, Artificial Intelligence (AI), Content display
- Created by jurriaanroelofs on , updated
Stable releases for this project are covered by the security advisory policy.
There are currently no supported stable releases.
