Lightning-Fast, On-Device, Multilingual TTS โ Gradio, ONNX, 44.1kHz
Creations by @pierrunoyt
57 totalZero-shot voice cloning across 24 languages and 21 Chinese dialects, plus instruction-driven voice design and semantic and acoustic speech editing.
Run AuK Base and AuK-Flash locally for speech generation, editing, enhancement, and separation.
AI-Powered Text-to-Speech with Voice Cloning using Chatterbox TTS and a Gradio interface. Includes Turbo, Multilingual (23+ languages), and Original models. Runs locally; CUDA GPU recommended, CPU supported. Windows, Mac, and Linux.
High-quality rapid TTS voice cloning model (150x+ realtime) โ 48kHz speech, voice cloning
Hy-MT2 multilingual translation โ Gradio UI with 38 language and variant choices for Hy-MT2-1.8B, Hy-MT2-7B, and Hy-MT2-30B-A3B.
Fast Image Generation with Sana Diffusion Model
State-of-the-art open-source speech recognition model supporting 14 languages. 2B parameter ASR model from Cohere Labs.
Expressive TTS with voice cloning, prompt-driven speech synthesis built on LTX-2.3 by Resemble AI
3.8B foundational text-to-image model by Microsoft โ Lens and Lens-Turbo variants
A local AI chatbot powered by SmolLM3-3B with a Gradio web interface.
NVIDIA PiD โ Pixel Diffusion Decoder for high-resolution latent decoding. Gradio UI for Z-Image + 4ร PiD upscale, plus CLI demos for Flux.
All-in-one Gradio UI for the MOSS-TTS Family: voice cloning, dialogue generation, voice design from text, and sound effects.
Pinokio launcher for Higgs Audio v3 TTS with Gradio UI, SGLang-Omni backend, and automatic model download.
Bulk transcribe many YouTube videos, whole playlists, or your own uploaded audio/video files at once with faster-whisper. Outputs txt, srt, vtt, or json.
Zero-shot multilingual TTS (600+ languages) with voice cloning and voice design โ Gradio UI (app/app.py)
Pixel-space PRX text-to-image pipeline (~7B params, Qwen3-VL text encoder, no VAE)
NVIDIA's Audio Flamingo 3 - Large Audio-Language Model for speech, sound, and music understanding with Gradio web interface
Ideogram 4 (nf4 / fp8) open-weights text-to-image model (9.3B params, Qwen3-VL-8B text encoder, structured JSON prompting, native 2k resolution)
2B-parameter fully continuous, end-to-end autoregressive text-to-speech with zero-shot voice cloning. https://huggingface.co/rednote-hilab/dots.tts-base
