BottleCap AI has launched ThinkingCap-Qwen3.8-27B, the second mannequin in its ThinkingCap sequence. It’s a fine-tune of the Qwen group’s Qwen3.8-27B with one slim purpose: shorter reasoning traces. Throughout 12 benchmarks, it spends 37.2% fewer pondering tokens on common. Macro-average accuracy strikes from 86.65% to 85.79%, a 0.86pp drop.
Deployable? Sure. It drops in for Qwen3.8-27B on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds. The repo is gated, and industrial use past the small-business license wants a BottleCap settlement.
What Downside Does ThinkingCap Goal?
Reasoning fashions typically spend extra pondering tokens than a query wants. BottleCap’s place is that lots of these additional tokens don’t change the ultimate reply. The first launch within the sequence utilized this concept to Qwen3.6-27B.
The target this time was intentionally conservative. BottleCap didn’t attempt to add information or change reply model. Reasoning capacity, instruction following and security behaviour have been meant to go via untouched. The analysis group additionally centered more durable on math, reasoning, long-context and agentic benchmarks.
Benchmark Outcomes at xhigh Effort
All foremost numbers use reasoning_effort=xhigh, the chat template default. Each benchmark will get shorter, with cuts starting from 10.7% to 65.5%.
Data and multilingual duties shrink probably the most. MMMLU drops 65.5% (1,656 to 571 tokens) and MMLU-Professional drops 57.3%. GPQA-Diamond falls from 12,772 to 7,267 tokens, a 43.1% lower. IFBench thinks 46.4% much less with accuracy practically flat (79.75% to 79.71%).
Lengthy-context retrieval improves. AA-LCR accuracy rises 2.25pp, from 81.75% to 84.00%, with 38.6% fewer pondering tokens. LiveCodeBench v6 edges up 0.07pp whereas pondering 20.3% much less.
Agentic outcomes maintain near the bottom. τ²-bench provides up 1.01pp for a 30.9% lower. Terminal-Bench 2.1 loses 0.56pp, properly inside its ±4.26 interval, for a ten.7% lower.
The most costly commerce is AIME 2026. Accuracy falls 3.85pp, from 98.13% to 94.27%, for 30.2% much less pondering.
Please be aware that the 37.2% determine is the imply of the 12 per-benchmark reductions. Pooled imply pondering tokens fall from 15,735 to 12,144.
BottleCap additionally reviews a funds curve. Below a 16K-token cap per response, ThinkingCap scores greater than the bottom mannequin. Truncated traces fall from 0.51% to 0.34%, and looping from 0.06% to 0.05%.









