High-quality rapid TTS voice cloning model (150x+ realtime) โ 48kHz speech, voice cloning
Projects by @pierrunoyt
48 total๐ฃ๏ธ PersonaPlex - NVIDIA's real-time speech-to-speech conversational AI model. Natural full-duplex conversations with customizable personas and voices.
BreezeBlue Breeze TTS 2 โ open-weight bilingual (EN/ZH) text-to-speech with voice clone, voice design, voice direction, and inline vocal events. Requires an NVIDIA GPU with 12GB+ VRAM.
State-of-the-art open-source speech recognition model supporting 14 languages. 2B parameter ASR model from Cohere Labs.
Zero-shot voice cloning across 24 languages and 21 Chinese dialects, plus instruction-driven voice design and semantic and acoustic speech editing.
๐ต YouTube to MP3 downloader with a simple Gradio UI and bundled FFmpeg. Paste a YouTube link to download MP3.
Pixel-space PRX text-to-image pipeline (~7B params, Qwen3-VL text encoder, no VAE)
๐๏ธ Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning. High-quality text-to-speech synthesis supporting zero-shot voice cloning and streaming inference with natural emotional expression.
YouTube to MP3, Cohere transcription, TranslateGemma translation.
Bulk transcribe many YouTube videos, whole playlists, or your own uploaded audio/video files at once with faster-whisper. Outputs txt, srt, vtt, or json.
Instant, Ultra-Realistic Text-to-Speech
NVIDIA's Audio Flamingo 3 - Large Audio-Language Model for speech, sound, and music understanding with Gradio web interface
Ultra-lightweight text-to-speech (15M-80M params) โ CPU optimized, 8 voices, ONNX-powered
Hy-MT2 multilingual translation โ Gradio UI with 38 language and variant choices for Hy-MT2-1.8B, Hy-MT2-7B, and Hy-MT2-30B-A3B.
Advanced 3B parameter language model with Gradio web interface, GPU acceleration, and complete privacy
Liquid Audio - LFM2.5-Audio-1.5B: speech-to-speech, ASR, and TTS powered by Liquid AI.
Tokenizer-free TTS for context-aware speech, voice cloning, and voice design. 2B params, 48kHz, 30 languages (Gradio UI).
Gradio web interface for Photoroom's PRX-1024-t2i-beta text-to-image model
A web interface for the Moondream3 vision-language model featuring image captioning, visual question answering, object detection, and object pointing.
High-quality Text-to-Speech powered by VyvoTTS LFM2 model with easy-to-use web interface
