A comprehensive AI-powered video production studio. Features local batch processing for automated dubbing (XTTS), smart audio censorship (Whisper), and visual NSFW blurring (NudeNet) wrapped in a modern dashboard UI.
Global radar
This tool creates videos where a person's lips move in sync with spoken audio, using just a single image and an audio file. Users upload a photo and voice recording, then customize the video length...
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
IndicF5: High-Quality Text-to-Speech for Indian Languages , including voice cloning
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest AI-driven technologies. The solution offers an industry leading WebUI, and serves as the foundation for multiple commercial products.
This is an implementation of iperov's DeepFaceLab and DeepFaceLive in Stable Diffusion Web UI 1111 by AUTOMATIC1111.
A Family of Open Sourced Music Foundation Models
Amica is an open source interface for interactive communication with 3D characters with voice synthesis and speech recognition.
This codebase is for a React and Electron-based app that executes the FreedomGPT LLM locally (offline and private) on Mac and Windows using a chat-based interface
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the model.
