BottleCap Launches ThinkingCap-Qwen3.8-27B Model
BottleCap AI releases ThinkingCap-Qwen3.8-27B, cutting thinking tokens by an average of 37.2% with only a 0.86pp accuracy cost on vLLM and SGLang.

Stock photo for illustration only, not from the actual event
- Reduces thinking tokens by an average of 37.2% across 12 benchmarks
- Incurs an accuracy cost of just 0.86pp
- Drop-in replacement for vLLM and SGLang with 5 quantized builds
- Licensed under PolyForm Small Business 1.0.0
BottleCap AI has officially released ThinkingCap-Qwen3.8-27B, a new artificial intelligence model building upon the foundation of Qwen3.6-27B. The research team focused heavily on reducing the generation of excess thinking tokens that frequently exceed what simple questions actually require, all while preserving core reasoning abilities, instruction-following capabilities, and safety behaviors.
This latest release deliberately targets math, reasoning, long-context, and agentic benchmarks using the default chat template setting of reasoning_effort=xhigh for evaluation. Across all tested benchmarks, the model achieved thinking token reductions ranging from 10.7% to 65.5%, bringing the pooled mean thinking tokens down from 15,735 to 12,144.
Knowledge and multilingual tasks experienced the most significant contractions. MMMLU dropped by 65.5% from 1,656 down to 571 tokens, MMLU-Pro fell by 57.3%, and GPQA-Diamond decreased by 43.1% to 7,267 tokens. Meanwhile, IFBench required 46.4% fewer thinking tokens while maintaining a nearly flat accuracy of 79.71%, compared to the base model's 79.75%.

Stock photo for illustration only, not from the actual event
Long-context retrieval also saw improvements, with AA-LCR accuracy rising by 2.25pp from 81.75% to 84.00% alongside a 38.6% reduction in thinking tokens. Agentic results remained closely aligned with the base model, as τ²-bench traded 1.01pp of accuracy for a 30.9% cut, and Terminal-Bench 2.1 lost 0.56pp for a 10.7% reduction. However, the most expensive trade occurred on AIME 2026, where accuracy dropped by 3.85pp from 98.13% to 94.27% for 30.2% less thinking.
"BottleCap team recommends xhigh for the best accuracy-to-token balance. It says individual thinking modes will get attention in a future release."
BottleCap Research Team
Reasoning models often suffer from over-thinking on tasks that do not demand extensive computational traces. BottleCap's methodology successfully curbs these redundant tokens without sacrificing significant model performance, marking an important step toward lowering enterprise inference costs and accelerating response times.
For deployment, the model serves as a direct drop-in replacement for Qwen3.8-27B on vLLM or SGLang, featuring 5 published quantized builds including FP8, NVFP4, GGUF, and MLX formats. The repository is gated, and commercial usage beyond the small-business license requires a specific BottleCap agreement.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment