Run image generation, GGUF language models, Whisper speech recognition, and Kokoro speech synthesis locally from one offline studio.
Global radar
Self-improving CLI agent by Nous Research with local memory, skills, and messaging workflows.
open source chat UI for Ollama https://github.com/ivanfioravanti/chatbot-ollama
An all-in-one, 100% local AI video, image & music studio. Its Director mode turns a single prompt into a full music video or short film — LLM-planned, shot by shot. Built on the WanGP pipeline (Wan 2.1/2.2, LTX-2.3, Qwen, Hunyuan Video, Flux). Requires an NVIDIA GPU (6GB+ VRAM).
Pyramd Flow Video Generation AI (text-to-video & image-to-video) https://github.com/jy0205/Pyramid-Flow
An enhanced local port of finegrain-image-enhancer powered by Refiners (https://huggingface.co/spaces/finegrain/finegrain-image-enhancer), which was adapted from philz1337x's Clarity Upscaler (https://github.com/philz1337x/clarity-upscaler)
1-click WanGP Launcher. Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP
Minimal Flux Web UI powered by Gradio & Diffusers (Flux Schnell + Flux Merged)
MiniMax H3 omni-modal video generation in ComfyUI. Text/image/video/audio in, video with native 32kHz stereo audio out (768p default, 1080p+ supported). Disk-optimized: pruned INT8 + NVFP4 weights (~63GB instead of ~290GB). NVIDIA only.
Estúdio IA para Criação de Letras e Prompts Estruturados para Suno.ai e Mureka.ai — gere músicas muito melhores com engenharia de prompt avançada.
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching https://huggingface.co/spaces/mrfakename/E2-F5-TTS
Expressive TTS with voice cloning, prompt-driven speech synthesis built on LTX-2.3 by Resemble AI
Describe any song in plain English, compose a structured caption and lyrics locally, and generate it with MiniMax Music 3.
[NVIDIA ONLY] The most efficient way to run FLUX (Optimized to run even on low memory machines, as low as 3GB VRAM with 512x512 resolution) https://github.com/lllyasviel/stable-diffusion-webui-forge
AI-Powered Text-to-Speech with Voice Cloning using Chatterbox TTS and Gradio interface. Includes Turbo, Multilingual (23+ languages), and Original models.
Style Aligned Image Generation via Shared Attention https://style-aligned-gen.github.io/
Run and train AI models with a unified local interface. https://github.com/unslothai/unsloth
Customizing Realistic Human Photos via Stacked ID Embedding https://huggingface.co/spaces/TencentARC/PhotoMaker-V2
