[NVIDIA ONLY] Generate an image from multiple images https://github.com/bytedance/UNO
Projects on Pinokio
[Mac Only] Super Fast MLX Powered Video Transcription https://github.com/RayFernando1337/MLX-Auto-Subtitled-Video-Generator/ by https://x.com/RayFernando1337
Advanced Gradio UI for Stable Audio https://github.com/RoyalCities/RC-stable-audio-tools
Pinokio launcher for the MLX-only SongGeneration Studio.
An advanced vision foundation model from MicroSoft https://huggingface.co/spaces/gokaygokay/Florence-2
the simplest self-building coding agent https://github.com/yoheinakajima/ditto
Accelerating any conditional diffusion model for few steps image generation https://gojasper.github.io/flash-diffusion-project/
A Step Towards Music Generation Foundation Model
Roblox Foundation Model for 3D Intelligence --- Cross Platform (Mac, Windows, Linux): Requires 16GB+ VRAM PC or 18GB+ Memory Macs https://github.com/Roblox/cube
Integrates Florence2 and SAM2 models for detailed image captioning and object detection. Florence2 generates detailed captions that are then used to perform phrase grounding. The Segment Anything Model 2 (SAM2) converts these phrase-grounded boxes into masks. https://huggingface.co/spaces/SkalskiP/florence-sam
Unified Image Understanding and Image Generation with Data and Model Scaling https://github.com/peanutcocktail/Janus
Describe UI and see it rendered live. Ask for changes and convert HTML to React, Svelte, Web Components, etc. Like vercel v0, but open source https://github.com/wandb/openui
AudioX Diffusion Transformer for Anything-to-Audio Generation
Nari Dia is a powerful text-to-speech (TTS) application based on the Dia-1.6B model from Nari Labs. This application allows you to convert text into natural-sounding speech with various customization options.
[Mac Onlyl] An all-in-one LLMs Chat UI for Apple Silicon Mac using MLX Framework. https://github.com/qnguyen3/chat-with-mlx
remove or change any video background https://huggingface.co/spaces/innova-ai/video-background-removal
Model Context Protocol https://modelcontextprotocol.io/introduction
[Mac only] a speech-text foundation model for real time dialogue https://github.com/kyutai-labs/moshi
Turn any raw text into a high-quality dataset for AI finetuning https://github.com/e-p-armstrong/augmentoolkit
