Video to Openpose & DWPose (All OS supported) https://github.com/sdbds/vid2pose
Creations by @morpheus
104 totalContribute to 6Morpheus6/vid_2_pose development by creating an account on GitHub.
Optimized Training script for Ace-Step with low VRAM support for local GPUs.
Bring portraits to life! https://github.com/KwaiVGI/LivePortrait
The open-source ElevenLabs alternative. Local voice cloning, design, creation and cinematic video dubbing with real-time dictation.
Local GPU-accelerated music video generator: Gradio UI, analysis, SDXL backgrounds, NVENC output.
Minimal Flux Web UI powered by Gradio & Diffusers (Flux Schnell + Flux Merged)
AudioCraft Plus is an all-in-one WebUI for the original AudioCraft, adding many quality features on top https://github.com/GrandaddyShmax/audiocraft_plus
All in one Gradio interface for chatterbox. Voice cloning from uploaded audio samples, automatic text processing for long content and real-time speech generation with configurable parameters. (Minimum Requirements 4GB VRAM / Recommended Requirements 8GB VRAM)
An open-source, modern-design ChatGPT/LLMs UI/Framework. Supports speech-synthesis, multi-modal, and extensible (function call) plugin system. https://github.com/lobehub/lobe-chat
High-Quality Text-to-Speech for Indian Languages
Automatically remove watermarks from videos generated by Sora AI.
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. https://github.com/comfyanonymous/ComfyUI
Based on BFS - Best Face Swap, VisoMaster, and SwapAnyHead.
Generate music in different genres using text and audio prompts.
Separate Anything You Describe (https://huggingface.co/spaces/Audio-AGI/AudioSep)
Kimodo generates high-quality 3D human and robot motions and is controlled through text prompts
Minimal Stable Diffusion UI
[NVIDIA ONLY] Make virtual avatars talk whatever you want with an image and an audio clip https://github.com/antgroup/echomimic_v2
Xkaliber Agent is a powerful, autonomous AI interface built on Electron. It goes beyond standard chat by allowing local Ollama models to autonomously manage your file system, interact with APIs, and control physical robotics via GPIO. It features offline text-to-speech, multimodal vision, WhatsApp integration,
