Qwen3.8 Flash Beats Max Prime in Forge App Benchmark
The cheaper Qwen3.8 Flash outperformed Qwen3.8 Max Prime on Atlassian Forge with significantly lower running costs.

Stock photo for illustration only, not from the actual event
- Qwen3.8 Flash scored 0.7961, defeating Max Prime's 0.6953 on Atlassian Forge 1.0
- Flash's Forge run cost 47 cents compared to Max Prime's $16.46
- Max Prime lost to Flash on Gauntlet 7.2 with 0.6846 against 0.8364
- A single manifest line regarding storage indexes proved to be the scoring turning point
Recent benchmark results within the Qwen3.8 model family revealed a surprising outcome where the much cheaper Flash model outperformed the flagship Qwen3.8 Max Prime on the Atlassian Forge 1.0 application benchmark. Flash achieved a score of 0.7961, edging out Max Prime's 0.6953. However, on the Gauntlet 7.2 payments web application benchmark, Max Prime scored 0.8364, outperforming Flash's 0.6846. The execution costs showed a massive gap, with Flash costing 47 cents while Max Prime cost $16.46.
Max Prime is marketed as a faster version of Qwen3.8 Max, a 2.4-trillion-parameter flagship model. Flash features a 125-billion-parameter main model and consumes 6 billion parameters per token. Both models were driven by the goose agent on builds 3.0.100 for Flash and 3.0.107 for Max Prime, using a 150-call budget hosted on Alibaba via OpenRouter, executed a day apart. During the Forge build, both models performed similarly until hitting a strict platform ceiling limit.
The ceiling limit originated from a single line in Max Prime's manifest file, specifically a storage index containing two range attributes where Atlassian Forge strictly permits only one. When both models were subsequently queried about this rule across 100 calls costing roughly 60 cents, they mostly stated a range takes one attribute, before incorrectly stating "one" for the partition as well, whereas Atlassian's documentation specifies several. Given only the raw contract, both models wrote two ordering attributes in 14 out of 15 answers, but once Atlassian's sentence regarding range was included in the prompt, both adhered to a single range attribute in 8 out of 8 attempts.
"Optimizes your index for the use of query conditions. This parameter can only have one attribute."
Atlassian's custom entities reference

Stock photo for illustration only, not from the actual event
This scenario highlights a common challenge in AI evaluation where models fail to inherently recall specific platform rules without explicit prompt guidance. Even massive language models can stumble on strict platform linting requirements if constraints are not explicitly defined. Furthermore, the stark cost difference demonstrates the high efficiency of smaller specialized models for scoped tasks.
The Atlassian Forge 1.0 benchmark tasks models with building a Jira sprint scope-creep ledger inside an internet-free sandbox, evaluated by running the application against three seeded Jira sites. Both models achieved nearly identical core scores of 0.9537 for Flash and 0.9519 for Max Prime with zero critical defects. The score divergence stemmed entirely from the platform's band-capping mechanism.
Flash's worst failure sat in the current platform band capped at 0.799 due to dark mode issues, while Max Prime failed a check in the lower working ledger band capped at 0.699. The scorer then applies a slight downward adjustment based on missed points, resulting in Max Prime's final score of 0.6953. Without this index error, Max Prime's earned score of 0.9269 would have placed it fifth on the Forge leaderboard.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment