A macOS-optimized distribution of the Retrieval-based Voice Conversion WebUI, featuring enhanced compatibility and performance for Apple Silicon (M1/M2/M3/M4/M5) and Intel Macs.
JoyCaption is an image captioning Visual Language Model (VLM) being built from the ground up as a free, open, and uncensored model for the community to use in training Diffusion models.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.