All Models
Qwen2.5-72B-Instruct
Qwen2.5-72B-Instruct is a flagship 72.7-billion parameter, open-weight model engineered for high-stakes AI production, advanced reasoning, and complex conversational tasks. Optimised with an 80-layer architecture using Grouped Query Attention (GQA), it delivers exceptional performance in coding, mathematics, and structured data tasks like JSON and table processing.
by @AIOZAI
Qwen3-0.6B
Qwen3-0.6B is a compact 0.6B-parameter causal language model from Alibaba Cloud’s Qwen team. Featuring a hybrid Thinking/Non-Thinking mode, it balances deep reasoning with fast, efficient responses. With support for a 32K token context window, over 100 languages, and native tool-calling via the Qwen-Agent framework, it is well-suited for lightweight chatbots, math/code reasoning, and agentic applications.
by @AIOZAI
Handwritten Digit Recognition Challenge
Baseline source code for Handwritten Digit Recognition Challenge
by @AIOZAI
Ministral-3 3B-Base 2512
Ministral 3 3B Base 2512 is a compact multimodal foundation model from Mistral AI. It combines a 3.4B-parameter language decoder with a 0.4B frozen vision encoder for native image understanding. Distilled from the 24B Mistral Small 3.1 through an iterative “Cascade Distillation” process, it preserves much of its teacher model’s capability while supporting a 256K token context window via YaRN RoPE scaling.
by @AIOZAI
MiMo-7B-Base
A 7B-parameter decoder-only language model developed by Xiaomi. Trained from scratch for mathematical and coding reasoning, it serves as the foundational checkpoint of the MiMo-7B series, providing a strong base for fine-tuning and reinforcement learning.
by @AIOZAI
Granite-4.0-Micro
A 3B-parameter long-context instruct model from IBM, finetuned for enhanced instruction following and tool-calling. Supports 12 languages including English, Chinese, Arabic, and Japanese. Built on a dense Transformer with GQA, RoPE, SwiGLU, and 128K context length. Trained using SFT, RL alignment, and model merging techniques for enterprise applications.
by @AIOZAI
Qwen-Image 2512
Qwen-Image-2512 is a 20B-parameter text-to-image generation model built on the Multimodal Diffusion Transformer (MMDiT) architecture with a frozen Qwen2.5-VL semantic encoder. It delivers photorealistic image generation, with improved human rendering, natural textures, and accurate text rendering in both English and Chinese for posters, infographics, and slides.
by @AIOZAI
Wine Quality Classification Challenge
Baseline source code for Wine Quality Classification Challenge
by @AIOZAI
Wan2.2-Animate-14B
A character animation and replacement video model: give it one character image and a driving video, and it transfers full-body motion and facial expression onto your character in either animation or replacement mode.
by @AIOZAI
MiniMax-M2.5
A large agentic model built for coding, tool use, search, and office work, reporting strong results on SWE-Bench Verified and other agentic benchmarks.
by @AIOZAI
Olmo-3.1-32B-Instruct
A 32B instruction-tuned model from Ai2, built for reasoning, coding, and instruction-following, and released as a fully open model with public code, checkpoints, and training data.
by @AIOZAI
SmolLM3
A compact 3B model that punches above its size: dual-mode reasoning, six languages, and up to 128k-token context, released fully open with weights, data mixture, and training details.
by @AIOZAI
Smoker Classification Challenge
Baseline source code for Smoker Classification Challenge
by @AIOZAI
Devstral Small 2 24B Instruct 2512
An agentic coding model that explores codebases, edits across multiple files, and drives software-engineering agents — light enough to run on a single GPU, with a 256k context window and vision support.
by @AIOZAI