Local Text-to-Speech adds self-hosted text-to-speech to any Drupal site. Visitors click "Listen to this page" and hear the content read aloud by one of 54 voices across 9 languages. No cloud APIs, no subscriptions, no data leaving your server.

Features

  • 54 voices across 9 languages: American English, British English, Japanese, Mandarin Chinese, French, Hindi, Spanish, Italian, Portuguese
  • Language-aware voice filtering: automatically shows only voices matching the content language
  • Adjustable speech speed: 0.5x to 2.0x
  • Audio caching: generated files cached on disk to reduce server load
  • Block and field integration: place a player site-wide or per content type
  • Drush commands: test, batch generate, list voices, clear cache
  • Accessibility-optimised: ARIA labels, keyboard shortcuts, semantic HTML
  • Privacy-first: all processing local, only publicly viewable content converted
  • Data sovereignty: no cloud, no third-party SaaS, all data stays on your server

How it works

Local TTS uses the Kokoro TTS engine (an 82M-parameter model) compiled to a native binary via Kokoros. Text is converted to phonemes by eSpeak NG, then the Kokoro ONNX model generates natural-sounding speech. The entire pipeline runs locally on your server.

You need Local TTS if

  • You want to add text-to-speech to your Drupal site without a cloud service
  • You want multi-language speech from a single module with no per-request costs
  • You need GDPR-compliant audio where no visitor data leaves your infrastructure
  • You want a privacy-first accessibility feature that works without third-party JavaScript

Installation

System dependency

eSpeak NG must be installed on the server for phoneme processing:

# macOS
brew install espeak-ng

# Ubuntu/Debian
sudo apt-get install espeak-ng

Module installation

composer require drupal/local_tts
cd web/modules/contrib/local_tts
composer run build-binary
drush en local_tts

The composer install step downloads model and voice data files (~337MB). The binary must be built from source (requires Rust/Cargo). After enabling, configure the eSpeak NG data path at /admin/config/media/local-tts.

Verify installation

drush local-tts:test "Hello World"

Drush command reference

Command Description
local-tts:test Test TTS generation with instant playback
local-tts:read Read entity content aloud (with caching)
local-tts:batch Batch generate audio for multiple entities
local-tts:voices List available voices, optionally filtered by language
local-tts:cache-clear Clear all cached audio files

FAQ

Does my content leave the server?

No. All speech generation happens locally using the bundled Kokoro binary and ONNX model files. No text or audio is sent to any external service.

How large are the model files?

The model is approximately 310MB and the voice data is approximately 27MB. Files are downloaded automatically during composer install.

Does Local TTS work with page caching?

Yes. Audio is loaded via AJAX when the visitor clicks the play button, so the page itself can be fully cached (including behind Varnish or a CDN). Cached audio files are served directly without regeneration.

Does Local TTS set cookies or track users?

No. The module only processes the content text server-side when requested. It does not set cookies, store user IDs, or create visitor profiles.

Can I add custom voices?

Not at present. Local TTS uses the voices bundled with the Kokoro model. New voices would require training a custom ONNX model.

  • Module documentation: installation, configuration, voice reference, and Drush commands
  • Installation guide: system dependencies, binary build, and Nginx configuration
  • Available voices: all 54 voices across 9 languages with voice IDs
  • FAQ: caching, privacy, load management, and troubleshooting

Project information

Releases