Fast Lipsync application for smaller GPU's.
Creations by @morpheus
104 totalSuper Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP
[NVIDIA ONLY] The most efficient way to run FLUX (Optimized to run even on low memory machines, as low as 3GB VRAM with 512x512 resolution) https://github.com/lllyasviel/stable-diffusion-webui-forge
AI Song Generation with Full Style Control - Generate complete songs with lyrics, vocals, and instrumental tracks using Tencent AI Lab's SongGeneration (LeVo) model. [NVIDIA ONLY]
[NVIDIA Only] Dead simple web UI for training FLUX LoRA with LOW VRAM support (From 12GB)
[AMD ONLY] Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video, Flux and more. (On Windows supported by all dedicated AMD GPUs from RDNA 2 - RDNA 4)
Video to 3D: 4D Face Reconstruction from any Video or Image Sequence. Normal Map, Depth Map and 3D Mesh Generation.
[NVIDIA ONLY] AllTalk-TTS is a unified UI for E5-TTS, XTTS, Vite TTS, Piper TTS, Parler TTS and RVC, based on CoquiTTS, including a finetune mode.
[NVIDIA ONLY] Stable Diffusion WebUI Forge supporting Flux, Qwen, wan, nunchaku and more in a lightweight WebUI. https://github.com/Haoming02/sd-webui-forge-classic/tree/neo
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Generating Consistent Long Depth Sequences for Open-world Videos
Super fast Multilingual TTS supporting 54 voices across 8 languages.
Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers. https://github.com/Zyphra/Zonos
A Web UI for easy subtitle using whisper model.
Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech application
Create 3D Meshes of Body Poses from Images.
X-Voice is a multilingual text-to-speech system that enables one speaker to speak 27 languages.
A FastAPI wrapper for KokoroTTS. Integrates with Open-WebUI and other API-driven AI applications.
X-Voice
One-click launcher for the original HiDream-O1-Image web UI using lazy-downloaded drbaph Dev or Full FP8 checkpoints through a root FP8 runner. Requires an NVIDIA CUDA GPU.
