The fastest way to find the right AI model for your task, hardware, and budget.
This site covers text & language models— GPT, Claude, Llama, DeepSeek, Gemini, Qwen and more. Image, video & audio generation are a different category and not included.
All 24+ models ranked by overall score
Head-to-head across 14 task categories
MMLU, GPQA, SWE-bench, AIME and more
Best models per industry vertical
Whitepaper-derived strengths & weaknesses
Hardware-tier deployment guide
Price per 1M tokens vs overall score
Your task, your hardware, your priority
Deep model analysis: VibeThinker 3B breakthrough, Ornith-1.0 agents, SubQ 12M context, Qwythos 9B distillation, LongCat 2.0 reveal
We track every major release. Here's what's new in 2026.
DeepReinforce open-sources coding models — 397B MoE beats Claude Opus 4.7 on Terminal-Bench 2.1. MIT license.
First closed-weight Qwen ever. 1M context, reasoning-native agent model, Arena Elo 1,475. Qwen3.7-Plus adds vision.
SWE-bench leader. Now with sub-agent orchestration, adaptive thinking, and 2.5× fast mode at 1/3 the cost.
Built a full SysY compiler in Rust via 672 tool calls. MIT license, 1M context, native multimodal.
Best open-source math and code model. MIT license. FP8 optimized for domestic hardware.
744B MoE with 40B active. Async RL for agentic coding. First major model trained without NVIDIA.
We've done the research so you don't have to. 24+ models. 12+ benchmarks. Real weaknesses verified against whitepapers. And a smart recommender that finds your best model in seconds.
Enter the Dashboard→