A CLI text-to-speech tool using the Kokoro model, supporting multiple languages, voices (with blending), and various input formats including EPUB books and PDF documents.
This app allows you to convert and separate audio files into vocals and instruments using various models. You can also generate speech from text in different languages. Upload or download models, a...
state-of-the-art tuning-free method to achieve ID-Preserving generation with only single image, supporting various downstream tasks. https://instantid.github.io/
Qwen3-aligner is an all-in-one audio processing tool powered by **Qwen3-ASR** and **Qwen3-ForcedAligner** from Alibaba's Tongyi Qianwen series. It supports single-file transcription & alignment, batch processing, SRT timestamp repair, and comes with a **bilingual (Chinese/English) Gradio UI**. No limit on audio duration.