The open-source ElevenLabs alternative. Local voice cloning, video dubbing, and real-time dictation — 646 languages, no API keys.
Creations by @pierrunoyt
57 totalAdvanced 3B parameter language model with Gradio web interface, GPU acceleration, and complete privacy
🌍 TranslateGemma - Google's open-source multilingual translation AI. Translate text across 55+ languages and extract/translate text from images. Powered by Gemma 3 architecture.
🎵 YouTube to MP3 downloader with a simple Gradio UI and bundled FFmpeg. Paste a YouTube link to download MP3.
Gradio web interface for Photoroom's PRX-1024-t2i-beta text-to-image model
High-quality Text-to-Speech powered by VyvoTTS LFM2 model with easy-to-use web interface
A web interface for the Moondream3 vision-language model featuring image captioning, visual question answering, object detection, and object pointing.
Instant, Ultra-Realistic Text-to-Speech
Ultra-lightweight text-to-speech (15M-80M params) — CPU optimized, 8 voices, ONNX-powered
Liquid Audio - LFM2.5-Audio-1.5B: speech-to-speech, ASR, and TTS powered by Liquid AI.
Tokenizer-free TTS for context-aware speech, voice cloning, and voice design. 2B params, 48kHz, 30 languages (Gradio UI).
Standalone Text-to-Speech using Orpheus TTS with a Gradio UI
LFM2.5-VL-450M (Liquid AI): compact vision–language model for image understanding. Gradio UI with upload/URL, prompt, and generation sliders.
Lightweight CPU text-to-speech with preset voices and optional Hugging Face-authenticated voice cloning.
🎙️ Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning. High-quality text-to-speech synthesis supporting zero-shot voice cloning and streaming inference with natural emotional expression.
Generalizable real-world image restoration (diffusers + Gradio). CUDA recommended; first run downloads HF weights.
YouTube to MP3, Cohere transcription, TranslateGemma translation, OmniVoice TTS. https://github.com/PierrunoYT/VidLingo-Pinokio
Advanced text-to-speech with voice cloning, multi-speaker support, and background music generation using Higgs Audio V2
Zero-shot multilingual voice cloning with cross-lingual synthesis, disentangled emotion control, pronunciation guidance, and speaking-speed control.
⚡️ Efficient 6B parameter image generation model with sub-second inference. Generate high-quality, photorealistic images with only 8 inference steps. Features bilingual text rendering (Chinese & English) and Single-Stream Diffusion Transformer architecture.
