Usage
Basic TTS Conversion
Convert ebooks directly to audio without LLM processing:
# EPUB to audiobook
audify run book.epub
# PDF to audio
audify run document.pdf
# Specify language (affects voice selection and TTS)
audify run book.epub --language pt
# Translate content (text is translated before TTS)
audify run book.epub --translate es
# Combined: source language with translation
audify run book.epub --language en --translate es
# Choose TTS provider
audify run book.epub --tts-provider openai
Understanding –language and –translate
--language(or-l): Sets the language for TTS voice selection and audio output. Default:en(English).--translate(or-t): Translates the extracted text to a target language before TTS synthesis.Can be used alone when source language is autodetected; otherwise provide
--languagefor explicit source languageExample:
--language en --translate esmeans “extract English text, translate to Spanish, then synthesize Spanish speech”Uses your configured LLM (local Ollama or commercial API) to perform translation
Translation Examples
# Translate an English book to Spanish audio
audify run english-book.epub --language en --translate es
# Portuguese book to French audio
audify run book.epub --language pt --translate fr
# Without source language specified, translation uses the autodetected language
audify run mixed-language-book.epub --translate es
LLM-Powered Audiobook Generation
Use an LLM to transform text into engaging audiobook scripts before TTS:
# Default audiobook style
audify audiobook book.epub
# Limit chapters
audify audiobook book.epub --max-chapters 5
# Custom voice and language
audify audiobook book.epub --voice af_bella --language en
# With translation (LLM processes source text, then script is translated for synthesis)
audify audiobook book.epub --translate pt
# Full workflow: extract as English, LLM processes English text, translate script to Spanish, synthesize
audify audiobook book.epub --language en --translate es -m "api:deepseek/deepseek-chat"
Note
When using --translate with audiobook generation, the LLM processes the extracted text in the source language to generate the script, then that script is translated to the target language before synthesis. This ensures the LLM has access to the original text for best results.
Using commercial LLM APIs
# DeepSeek (cost-effective)
audify audiobook book.epub -m "api:deepseek/deepseek-chat"
# Claude (high quality)
audify audiobook book.epub -m "api:anthropic/claude-3-5-sonnet-20240620"
# GPT-4
audify audiobook book.epub -m "api:openai/gpt-4-turbo-preview"
# Gemini
audify audiobook book.epub -m "api:gemini/gemini-1.5-pro"
See Commercial APIs for API key setup.
Task System
Use the --task flag to control how the LLM transforms your text:
# Audiobook style (default)
audify audiobook book.epub --task audiobook
# Podcast/lecture style
audify audiobook book.epub --task podcast
# Summary
audify audiobook book.epub --task summary
# Guided meditation
audify audiobook book.epub --task meditation
# Classroom lecture
audify audiobook book.epub --task lecture
# Custom prompt file
audify audiobook book.epub --prompt-file my-prompt.txt
See Tasks for details on creating custom tasks.
Directory Processing
Process multiple files into a single audiobook:
# All supported files in a directory
audify audiobook path/to/articles/
# With translation
audify audiobook path/to/articles/ --translate es
Supported file types: EPUB, PDF, TXT, MD
Each file becomes a separate episode with a synthesized title, and all episodes are combined into a single M4B audiobook with chapter markers.
Unified Convert Command
The convert command unifies run and audiobook functionality:
# Direct TTS (equivalent to audify run)
audify convert book.epub --task direct
# Audiobook (equivalent to audify audiobook)
audify convert book.epub --task audiobook
# Podcast style
audify convert book.epub --task podcast
Listing Available Options
# List available tasks
audify list-tasks
# Validate a custom prompt file
audify validate-prompt my-prompt.txt
# List languages
audify --list-languages
# List TTS voices
audify --list-voices
# List TTS models
audify --list-models
# List TTS providers
audify --list-tts-providers
Output Structure
Basic TTS (audify run)
data/output/[book_name]/
chapters.txt # Metadata
cover.jpg # Cover image (EPUB)
chapters_001.mp3 # Chapter audio files
chapters_002.mp3
book_name.m4b # Final audiobook
LLM Audiobook (audify audiobook)
audiobooks/[book_name]/
episodes/
episode_001.mp3
episode_002.mp3
scripts/
episode_001_script.txt
original_text_001.txt
chapters.txt
book_name.m4b