Fully offline document-to-speech converter using Coqui TTS models (XTTS v2, Bark, VITS, YourTTS) and StyleTTS 2. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio with voice cloning support. No API keys or cloud services required.
Fully offline document-to-speech converter using OpenVoice V2. Tone color conversion and voice cloning via MeloTTS. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Reconstruct 3D Gaussian Splatting worlds from video using non-rigid alignment. Supports fast and extensive modes with 2DGS/3DGS rendering. https://github.com/lukasHoel/video_to_world
BazedFrog/SongGeneration-Studiov3.7updated 1mo ago
AI Song Generation with Full Style Control - Generate complete songs with lyrics, vocals, and instrumental tracks using Tencent AI Lab's SongGeneration (LeVo) model. [NVIDIA ONLY]
Local LoRA trainer for MiniMax H3, Flux2 Klein 9B & Krea 2 — trains against the exact models you deploy. Desktop GUI + headless CLI. NVIDIA only (RTX 30/40/50, driver 555+).
A god roleplaying sandbox — AI characters talk to each other with automated conversations, character development, and god interference. Works with any OpenAI-compatible local LLM (e.g. LM Studio).
Describe an image, get a 100% schema-valid Ideogram 4 JSON prompt — generated fully locally with an embedded llama.cpp (no Ollama or LM Studio required).
An all-in-one, 100% local AI video, image & music studio. Its Director mode turns a single prompt into a full music video or short film — LLM-planned, shot by shot. Built on the WanGP pipeline (Wan 2.1/2.2, LTX-2.3, Qwen, Hunyuan Video, Flux). Requires an NVIDIA GPU (6GB+ VRAM).
pinokiofactory/Hunyuan3d-2-lowvramv3.7updated 1mo ago
Text/Image to 3D (Cross Platform: Mac + Windows + Linux): High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models. https://github.com/deepbeepmeep/Hunyuan3D-2GP