Unified Image Understanding and Generation. Text-to-Image Generation, In-context Generation, Instruction-guided Image Editing, Visual Understanding (Minimum Requirements 12GBV RAM / 48GB RAM, Recommended Requirements 24GB VRAM / 32GB RAM)
Projects by @morpheus
72 total[NVIDIA ONLY] Make virtual avatars talk whatever you want with an image and an audio clip https://github.com/antgroup/echomimic_v2
High-Quality Text-to-Speech for Indian Languages
[NVIDIA ONLY] End-to-end multimodal SVG generator capable of generating complex and detailed SVGs, from simple icons to intricate anime characters. (Minimum Requirements 12GB VRAM / 32GB RAM, Recommended Requirements 24GB VRAM / 24GB RAM)
Video to Openpose & DWPose (All OS supported) https://github.com/sdbds/vid2pose
Separate Anything You Describe (https://huggingface.co/spaces/Audio-AGI/AudioSep)
[NVIDIA ONLY] Advanced Web UI for CogVideo (text to video, image to video, video to video, extend video, etc) -- Generate videos with less than 10GB VRAM
[NVIDIA ONLY] Remove Objects in videos with inpainting. Recommended requirements 16 - 24 GB VRAM / 48 GB RAM, Minimal requirements 12GB VRAM / 32 GBRAM
Pinokio wrapper: installs HeartMuLa heartlib + downloads checkpoints + launches a Gradio UI for music generation.
Minimal Flux Web UI powered by Gradio & Diffusers (Flux Schnell + Flux Merged)
🦙 Let 2 models debate about a topic you pick. Create custom Ollama models with your own system prompts and parameters and use them to debate ot publish on ollama.com Easy-to-use Gradio interface for building personalized AI models with temperature control and custom instructions.
Unify Efficient Fine-Tuning of 100+ LLMs https://github.com/hiyouga/LLaMA-Factory
Customizing Realistic Human Photos via Stacked ID Embedding https://huggingface.co/spaces/TencentARC/PhotoMaker-V2
create a story by generating consistent images https://github.com/HVision-NKU/StoryDiffusion
A simple, high-quality image generation tool to create stunning illusions.
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
🎙️ Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning. High-quality text-to-speech synthesis supporting zero-shot voice cloning and streaming inference with natural emotional expression.
The Gen AI Platform for Pro Studios https://github.com/invoke-ai/InvokeAI
AudioCraft Plus is an all-in-one WebUI for the original AudioCraft, adding many quality features on top https://github.com/GrandaddyShmax/audiocraft_plus
[NVIDIA ONLY] Generate videos with less than 10GB VRAM https://github.com/THUDM/CogVideo
