Fully offline document-to-speech converter using Tortoise TTS. High-quality autoregressive synthesis with voice cloning. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Fully offline document-to-speech converter using Sesame CSM-1B. Conversational speech model with Llama backbone and voice cloning. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Fully offline document-to-speech converter using OpenVoice V2. Tone color conversion and voice cloning via MeloTTS. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
C0m3b4ck/DocToSpeech-Chatterboxv7.0updated 1mo ago
Fully offline document-to-speech converter using Resemble Chatterbox TTS. 350M parameter model with paralinguistic tags and voice cloning. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Finrandojin/alexandria-audiobookv5.0updated 1mo ago
A multi-voice AI audiobook generator built on Qwen3-TTS — annotate scripts with an LLM, assign unique voices to each character, per-line style instructions for delivery, clone voices from reference audio, design new voices from text descriptions, train custom voices with LoRA fine-tuning, and export to MP3 or Audacity multi-track projects
[LINUX + NVIDIA ONLY] Real-time interactive world model. Drive an infinite, action-conditioned world rollout at 720p/16fps on a single desktop GPU (~19GB VRAM). https://github.com/amap-cvlab/ABot-World
6Morpheus6/stable-diffusion-webui-forgev2.0updated 1mo ago
[NVIDIA ONLY] The most efficient way to run FLUX (Optimized to run even on low memory machines, as low as 3GB VRAM with 512x512 resolution) https://github.com/lllyasviel/stable-diffusion-webui-forge
Blizaine/Qwen3-TTS-MLX-WebUI-Enhancedv5.0updated 1mo ago
High-quality text-to-speech with Beautiful Web UI & API, optimized for Apple Silicon using MLX. Features include Custom Voice (preset speakers), Voice Design (natural language), and Voice Cloning. With enhanced features for saving custom voices and long-form / endless TTS streaming.