Configuration
Audify is configured through environment variables or a .keys file.
Priority: environment variables > .keys file > defaults.
The .keys File
Create a .keys file in the project root (or your working directory):
cp .keys.example .keys
# TTS provider (kokoro, openai, aws, google)
TTS_PROVIDER=kokoro
# Kokoro TTS (local)
KOKORO_API_URL=http://localhost:8887/v1
# OpenAI TTS
OPENAI_API_KEY=sk-your-key
OPENAI_TTS_MODEL=gpt-4o-mini-tts
OPENAI_TTS_VOICE=coral
# AWS Polly
AWS_ACCESS_KEY_ID=your-key
AWS_SECRET_ACCESS_KEY=your-secret
AWS_REGION=us-east-1
AWS_POLLY_VOICE=Joanna
AWS_POLLY_ENGINE=neural
# Google Cloud TTS
GOOGLE_APPLICATION_CREDENTIALS=/path/to/credentials.json
GOOGLE_TTS_VOICE=en-US-Neural2-F
GOOGLE_TTS_LANGUAGE_CODE=en-US
# Ollama (local LLM)
OLLAMA_API_URL=http://localhost:11434
OLLAMA_MODEL=magistral:24b
OLLAMA_TRANSLATION_MODEL=qwen3:30b
# Commercial LLM API keys
DEEPSEEK=sk-your-deepseek-key
ANTHROPIC=sk-ant-your-anthropic-key
OPENAI=sk-your-openai-key
GEMINI=your-google-api-key
Warning
Never commit .keys to version control. It is already in .gitignore.
TTS Provider Comparison
Provider |
Local? |
Free? |
GPU needed? |
Key Features |
|---|---|---|---|---|
Kokoro |
Yes |
Yes |
Recommended |
Fast, low-latency synthesis |
OpenAI |
No |
No |
No |
High quality, easy setup |
AWS Polly |
No |
No |
No |
Enterprise, multiple engines |
Google Cloud |
No |
No |
No |
Multilingual, neural voices |
Google Cloud TTS Configuration
When using Google Cloud TTS, the default voice is set to en-US-Neural2-F (English US, Female).
To use different languages or voices, set the environment variables:
GOOGLE_TTS_VOICE=<voice-id> # e.g., es-ES-Neural2-A for Spanish
GOOGLE_TTS_LANGUAGE_CODE=<lang-code> # e.g., es-ES for Spanish
Warning
Important: The voice locale must match the language. If you specify a voice like en-US-Chirp-HD-F, you should set GOOGLE_TTS_LANGUAGE_CODE=en-US. Mismatched values will cause errors.
Docker Services
The docker-compose.yml provides local services via profiles:
Service |
Profile |
Port |
Description |
|---|---|---|---|
Kokoro TTS |
|
8887 |
GPU-accelerated speech synthesis |
Ollama |
|
11434 |
Local LLM for script/translation |
Audify API |
|
8000 |
REST API (depends on Kokoro & Ollama) |
docker compose --profile kokoro --profile ollama up -d # Start TTS + LLM
docker compose --profile kokoro --profile ollama --profile api up -d # Include API
docker compose ps # Check status
docker compose logs -f # Follow logs
docker compose down # Stop all services
Supported Languages
English, Spanish, French, German, Italian, Portuguese, Polish, Turkish, Russian, Dutch, Czech, Arabic, Chinese, Hungarian, Korean, Japanese, Hindi.
Translation supports any language pair available in your configured LLM.
Pipeline Environment Variables
These variables control pipeline behaviour during audiobook creation:
Variable |
Default |
Description |
|---|---|---|
|
|
Set to |
|
|
Set to |