All Models
HeartMuLa
Most open-source music models give you one capability. HeartMuLa gives you four: a lyrics-conditioned song generator, a high-fidelity music codec, a lyrics transcription model, and an audio-text alignment model — all open-sourced together as a coherent foundation. The 3B generator handles multilingual lyrics across English, Chinese, Japanese, Korean, and Spanish, with style controlled through simple comma-separated tags. An internal 7B version already reaches Suno-level quality, with the open 7B release planned.
by @AIOZAI
UI-Venus 1.5
Give UI-Venus 1.5 a natural language instruction and a screenshot — it will find the right button, navigate the interface, and complete the task, just like a human would. No accessibility APIs, no DOM parsing, no special permissions needed. The unified 2B/8B/30B-A3B model family achieves state-of-the-art results on major GUI benchmarks including AndroidWorld (77.6%) and ScreenSpot-Pro (69.6%), with a full Android automation framework supporting 40+ mainstream apps out of the box.
by @AIOZAI
Text Generation with SmolLM-135M
Text Generation with SmolLM-135M involves utilizing a compact language model with 135 million parameters to automatically generate text. This model, although smaller in size, is proficient at producing coherent and structured textual content.
by @AIOZAI
Melanoma Skin Cancer Classification Challenge
Model for Melanoma Skin Cancer Classification Challenge
by @AIOZAI
DeepSeek-OCR
DeepSeek-OCR reimagines optical character recognition as a context compression problem — treating visual documents not as images to scan, but as information to compress and decode through an LLM-centric vision encoder. It converts documents, PDFs, and images to clean markdown, extracts text with layout awareness, parses figures, and localizes specific elements by reference — all at ~2500 tokens per second on a single A100 with vLLM. Multiple resolution modes from 64 to 400+ vision tokens let you tune the quality-speed tradeoff for your use case.
by @AIOZAI
LightRAG
LightRAG is a simple, fast, and powerful RAG system that goes beyond chunk retrieval by automatically building a knowledge graph from your documents — then querying both the graph and vector store simultaneously for richer, more contextually aware answers. Published at EMNLP 2025 and trusted by 29k+ developers, it works with any LLM, supports production-grade storage backends, and ships with a Web UI featuring live knowledge graph visualization.
by @AIOZAI
OmniLottie
OmniLottie is the first end-to-end model capable of generating Lottie animations directly from text descriptions, images, or video clips — producing structured, editable JSON output rather than raster video. Built on a 4B vision-language model and trained on MMLottie-2M, a dataset of 2 million annotated animations, it introduces a custom Lottie tokenizer that makes complex vector animation learnable by a language model. Accepted to CVPR 2026.
by @AIOZAI
Iris Flower Classification Challenge
Model for Iris Flower Classification Challenge
by @AIOZAI
Supertonic 2
Most TTS systems make you choose between speed, quality, and privacy. Supertonic 2 refuses that tradeoff — a featherweight 66M parameter model that runs 167× faster than real-time, entirely on your device, with zero network dependency. Powered by ONNX Runtime, it deploys across 11 platforms from iOS to Rust to the browser, supports 5 languages, and correctly reads complex real-world expressions that trip up every major cloud TTS API.
by @AIOZAI
Qwen3-TTS
Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.
by @AIOZAI
FireRed-Image-Edit-1.1
Edit anything in an image with a simple text instruction. FireRed-Image-Edit covers the full spectrum of image editing — from swapping backgrounds and retouching portraits to restoring old photos and performing virtual try-on across multiple images. With leading benchmark scores among all open-source models and bilingual Chinese–English instruction support, it's built for both researchers and real-world applications.
by @AIOZAI
Helios-Distilled
Generate up to 60 seconds of high-quality video from a text prompt, a single image, or an existing video clip — all running at real-time speeds on a single H100. Helios-Distilled is the most efficient variant of the 14B Helios family, distilled to 3 inference steps while maintaining strong visual coherence and temporal consistency. No KV-cache, no quantization, no anti-drifting hacks — just fast, clean video generation out of the box.
by @AIOZAI
Email Spam Classification Challenge
Model for Email Spam Classification Challenge
by @AIOZAI
Pneumonia Chest X-Ray Classification Challenge
Model for Pneumonia Chest X-Ray Classification Challenge
by @AIOZAI
Face Anti-Spoofing Challenge
Model for Face Anti-Spoofing Challenge
by @AIOZAI
Background Removal
Background Removal is an image processing technique, used to separate the main object from the background of a photo. Removing the background helps highlight the product, subject, or character, bringing a professional and aesthetically pleasing look to the image.
by @AIOZAI
Image Super-Resolution with SMFANet
Image Super-Resolution with SMFANet involves utilizing the SMFANet model architecture to enhance the resolution and quality of images. SMFANet is a deep learning network designed for super-resolution tasks, aiming to generate high-quality, detailed images from low-resolution inputs.
by @AIOZAI
Low-light Image Enhancement
Low light Image Enhancement is a task focused on improving the quality and visibility of images captured in low-light conditions. This task involves applying image processing techniques and algorithms to enhance details, reduce noise, and increase brightness in photos taken in dimly lit environments.
by @AIOZAI
MediaPipe Face Detection
Face detection is a computer vision technique that involves identifying and locating human faces within an image or video. The goal of face detection is to detect the presence of faces, and draw bounding boxes around them, without necessarily identifying specific facial features or landmarks.
by @AIOZAI
MediaPipe Face Mesh Plotting
Face mesh detection, also known as facial landmark detection or face pose estimation, is the task of identifying and localizing specific keypoints or landmarks on a human face. It involves detecting the positions of facial features, such as eyes, eyebrows, nose, mouth, and jawline, in an image or video.
by @AIOZAI