mikecastrodemaria/crispz.pinokiov7.0updated 3mo ago
Standalone local hi-res fix: enlarge images with Real-ESRGAN and reinject clean detail with Z-Image Turbo img2img. 100% local — no ComfyUI, no SwarmUI, no cloud.
AI-Powered Text-to-Speech with Voice Cloning using Chatterbox TTS and Gradio interface. Includes Turbo, Multilingual (23+ languages), and Original models.
Robust automatic speech recognition for challenging real-world audio. Handles noise, far-field, echo, reverberation, and more using a foundation model trained on 2.6M samples across 54 acoustic scenarios.