A CLI text-to-speech tool using the Kokoro model, supporting multiple languages, voices (with blending), and various input formats including EPUB books and PDF documents.
This app allows you to convert and separate audio files into vocals and instruments using various models. You can also generate speech from text in different languages. Upload or download models, a...
state-of-the-art tuning-free method to achieve ID-Preserving generation with only single image, supporting various downstream tasks. https://instantid.github.io/
Qwen3-aligner is an all-in-one audio processing tool powered by **Qwen3-ASR** and **Qwen3-ForcedAligner** from Alibaba's Tongyi Qianwen series. It supports single-file transcription & alignment, batch processing, SRT timestamp repair, and comes with a **bilingual (Chinese/English) Gradio UI**. No limit on audio duration.
[NVIDIA ONLY] Qwen-Image-Edit-2511 with 19 lazy-loaded LoRAs for single & multi-image editing (anime, pose transfer, relighting, upscaling, style transfer, and more). 4-step fast inference. https://github.com/PRITHIVSAKTHIUR/Qwen-Image-Edit-2511-LoRAs-Fast-Lazy-Load