All Models

Search
all
AIOZ AI
qwen2-5-72b-instruct

Qwen2.5-72B-Instruct

Qwen2.5-72B-Instruct is a flagship 72.7-billion parameter, open-weight model engineered for high-stakes AI production, advanced reasoning, and complex conversational tasks. Optimised with an 80-layer architecture using Grouped Query Attention (GQA), it delivers exceptional performance in coding, mathematics, and structured data tasks like JSON and table processing.

user-avatar
9
0
qwen3-0-6b

Qwen3-0.6B

Qwen3-0.6B is a compact 0.6B-parameter causal language model from Alibaba Cloud’s Qwen team. Featuring a hybrid Thinking/Non-Thinking mode, it balances deep reasoning with fast, efficient responses. With support for a 32K token context window, over 100 languages, and native tool-calling via the Qwen-Agent framework, it is well-suited for lightweight chatbots, math/code reasoning, and agentic applications.

user-avatar
1
8
handwritten_digit_baseline

Handwritten Digit Recognition Challenge

Baseline source code for Handwritten Digit Recognition Challenge

user-avatar
32
5
ministral-3-3b-base-2512

Ministral-3 3B-Base 2512

Ministral 3 3B Base 2512 is a compact multimodal foundation model from Mistral AI. It combines a 3.4B-parameter language decoder with a 0.4B frozen vision encoder for native image understanding. Distilled from the 24B Mistral Small 3.1 through an iterative “Cascade Distillation” process, it preserves much of its teacher model’s capability while supporting a 256K token context window via YaRN RoPE scaling.

user-avatar
21
5
wine_model

wine model

Model for Wine Quality Classification Challenge

mimo-7b-base

MiMo-7B-Base

A 7B-parameter decoder-only language model developed by Xiaomi. Trained from scratch for mathematical and coding reasoning, it serves as the foundational checkpoint of the MiMo-7B series, providing a strong base for fine-tuning and reinforcement learning.

user-avatar
69
13
wine_model

wine model

Model for Wine Quality Classification Challenge

granite-4-0-micro

Granite-4.0-Micro

A 3B-parameter long-context instruct model from IBM, finetuned for enhanced instruction following and tool-calling. Supports 12 languages including English, Chinese, Arabic, and Japanese. Built on a dense Transformer with GQA, RoPE, SwiGLU, and 128K context length. Trained using SFT, RL alignment, and model merging techniques for enterprise applications.

user-avatar
68
11
extremelypowerfulfromdistrict2

extremelypowerfulfromdistrict2

extremelypowerfulfromdistrict2

qwen-image-2512

Qwen-Image 2512

Qwen-Image-2512 is a 20B-parameter text-to-image generation model built on the Multimodal Diffusion Transformer (MMDiT) architecture with a frozen Qwen2.5-VL semantic encoder. It delivers photorealistic image generation, with improved human rendering, natural textures, and accurate text rendering in both English and Chinese for posters, infographics, and slides.

user-avatar
106
18
wine_quality_baseline

Wine Quality Classification Challenge

Baseline source code for Wine Quality Classification Challenge

user-avatar
92
14
wan2-2-animate-14b

Wan2.2-Animate-14B

A character animation and replacement video model: give it one character image and a driving video, and it transfers full-body motion and facial expression onto your character in either animation or replacement mode.

user-avatar
73
17
spaceship-model-2

spaceship model 2

spaceship model 2

minimax-m2-5

MiniMax-M2.5

A large agentic model built for coding, tool use, search, and office work, reporting strong results on SWE-Bench Verified and other agentic benchmarks.

user-avatar
120
24
olmo-3-1-32b-instruct

Olmo-3.1-32B-Instruct

A 32B instruction-tuned model from Ai2, built for reasoning, coding, and instruction-following, and released as a fully open model with public code, checkpoints, and training data.

user-avatar
149
22
smollm3-3b

SmolLM3

A compact 3B model that punches above its size: dual-mode reasoning, six languages, and up to 128k-token context, released fully open with weights, data mixture, and training details.

user-avatar
132
18
smoker_classification_baseline

Smoker Classification Challenge

Baseline source code for Smoker Classification Challenge

user-avatar
122
31
devstral-small-2-24b-instruct

Devstral Small 2 24B Instruct 2512

An agentic coding model that explores codebases, edits across multiple files, and drives software-engineering agents — light enough to run on a single GPU, with a 256k context window and vision support.

user-avatar
140
39