← All matchups

Fable 5vsQwen 3.8 Max

1 prompt where both models ran the exact same instructions. Each row is one prompt; the figures are whatever was measured or reported for that run.

Fable 5 and Qwen 3.8 Max ran the same prompt on 1 task, side by side. Cost, duration and outcome for each — 3 measured locally, no aggregate score.

Shared prompts
1
Compared
models
Measured runs
3

No winner is declared. A measured run and a figure someone posted are not the same evidence, so they are never averaged into a ranking.

Dofus Retro — Coin des Bouftous
GamesReplayable pack
Fable 5
$18.7not measuredCompletedmeasured
$11.8not measuredCompletedmeasured
Qwen 3.8 Max
$0.51not measuredCompletedmeasured

What this page shows, and how to read it

Fable 5 is a large language model from Anthropic. On the prompts below, it ran through Claude Code. Qwen 3.8 Max is a large language model from Alibaba. On the prompts below, it ran through Qwen Code.

The two sides share 1 prompt on this site: "Dofus Retro — Coin des Bouftous". It is a games prompt, rated hard, run locally in Bench Arena on Jul 23, 2026, tagged canvas2d, single-file and retro. Every run has a video capture on the prompt page. Across it, Fable 5 completed all 2 of its runs, and Qwen 3.8 Max completed its run. Fable 5 has measured costs from $11.8 to $18.7. Qwen 3.8 Max has a measured cost of $0.51.

How to read the outcomes. Completed means the prompt alone produced a working artefact. Failed means this run did not — the build broke, the result would not run, or the agent stopped short of a working state. That is a fact about one run on one prompt, not a verdict on the tool: the same programs complete other prompts elsewhere on this site. Timed out means the run was cut at its time limit with the artefact unfinished, so whatever it cost bought a partial result. Stopped and Interrupted mark runs ended from the outside before they concluded. A low cost attached to a run that did not finish is not a saving — it is the price of an attempt, which is why every figure on this page travels with its status.

Every figure above carries a trust label. Measured means SamePrompt ran it locally in Bench Arena, reconciled the tokens on the harness log and recomputed the cost from them. Reported means the author of the source announced the figure; it is shown as stated and cannot be verified here. The two are never summed, averaged or ranked, and this page declares no winner: a cheap run that failed and an expensive run that completed are two facts, not a score.