Models

Models are listed by how often they appear, not by how well they did. There is no score here: a measured run and a number someone posted are not the same evidence, and averaging them would produce a ranking that looks rigorous and is not.

ModelRunsMeasuredReported
Opus 5claude-opus-51459
Kimi K3 (high)kimi-k31376
Fable 5fable-5954
Qwen 3.8 Maxqwen3-8-max651
GPT-5.6gpt-5-6202
Grok 4.5 (high)grok-4-5202
1-bit Kimi K3kimi-k3-1bit101
Claude Opus 4.7claude-opus-4-7101
Claude Opus 4.8claude-opus-4-8101
DeepSeek V4 Flash-0731deepseek-v4-flash-0731101
GLM 5.2 (high)glm-5-2101
GPT-5.6 Solgpt-5-6-sol101
Kimi K2.6kimi-k2-6101
Luna 5.6 (max reasoning)luna-5-6101