Kimodo generates high-quality 3D human and robot motions and is controlled through text prompts
Creations by @morpheus
105 totalMinimal Stable Diffusion UI
[NVIDIA ONLY] Make virtual avatars talk whatever you want with an image and an audio clip https://github.com/antgroup/echomimic_v2
Xkaliber Agent is a powerful, autonomous AI interface built on Electron. It goes beyond standard chat by allowing local Ollama models to autonomously manage your file system, interact with APIs, and control physical robotics via GPIO. It features offline text-to-speech, multimodal vision, WhatsApp integration,
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
A professional, Suno-like music generation studio for HeartLib. https://github.com/fspecii/HeartMuLa-Studio
Suno-like music generation studio for HeartMuLa/heartlib - AI-powered music creation with reference audio style transfer
[NVIDIA, ROCM] One app to train them all. LORA training and Model finetuning for Z-Image, Qwen Image, FLUX.1, Flux.2 Dev and Klein, Chroma, SD 1.5 - 3.5, SDXL, Würstchen-v2, Stable Cascade, PixArt-Alpha, PixArt-Sigma, Sana, Hunyuan Video and inpainting models.
🎙️ Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning. High-quality text-to-speech synthesis supporting zero-shot voice cloning and streaming inference with natural emotional expression.
Official implementation of Kimodo, a kinematic motion diffusion model for high-quality human(oid) motion generation.
When Expressive Talking Head Generation Meets Diffusion Probabilistic Models (https://github.com/ali-vilab/dreamtalk)
Local-first voice synthesis studio powered by Qwen3-TTS.
Fork of the original SongGeneration project from Tencent.
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. https://github.com/comfyanonymous/ComfyUI
Sandboxed LM Studio-only OpenClaw fork packaged for Pinokio.
Distill and quantize models using TorchAO with intelligent GPU/CPU management
[NVIDIA ONLY] A minimal Gradio interface for Automatic Speech Recognition. Transcribe Audio in Malayalam language.
Prebuilt DeepSpeed wheels for Windows with NVIDIA GPU support. Supports GTX 10 - RTX 50 series. Compiled with pytorch 2.7, 2.8 and cuda 12.8
Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP
clone voices into different languages by using just a quick 3-second audio clip. (a local version of https://huggingface.co/spaces/coqui/xtts)
