# BottleCap AI ships a Qwen3.8-27B fine-tune that thinks 37% shorter

> BottleCap AI released ThinkingCap-Qwen3.8-27B, a Qwen3.8-27B fine-tune that cuts thinking tokens 37.2% on average across 12 benchmarks for a 0.86pp accuracy drop, per MarkTechPost.

- **Topic**: Models
- **Published**: 2026-09-25T02:15:52.022Z
- **Canonical URL**: https://highsignal.sh/stories/bottlecap-ai-ships-a-qwen3-8-27b-fine-tune-that-thinks-37-shorter-50080132

## Why It Matters

Reasoning models often burn tokens on steps that don't change the answer. BottleCap's pitch is that you can cut that spend without losing much accuracy — a tradeoff worth watching for anyone paying per reasoning token.

## Key Findings & Analysis

### What happened

- BottleCap AI released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap series — a fine-tune of the Qwen team's Qwen3.8-27B aimed only at shorter reasoning traces, MarkTechPost reports. The first release applied the same idea to Qwen3.6-27B. BottleCap says it did not try to add knowledge or change answer style, and intended reasoning, instruction following and safety behavior to pass through unchanged. The repo is gated, and commercial use beyond a small-business license requires a BottleCap agreement.

### The numbers, as BottleCap reports them

- All main results use reasoning_effort=xhigh, the chat template default, and run through one harness on a single NVIDIA H200 with vLLM 0.29.0. Every one of 12 benchmarks gets shorter, with cuts from 10.7% to 65.5%, per MarkTechPost. Macro-average accuracy slips from 86.65% to 85.79%, a 0.86pp drop. The 37.2% figure is the mean of the 12 per-benchmark reductions; pooled mean thinking tokens fall from 15,735 to 12,144; per MarkTechPost, multi-seed accuracy is reported as a mean with a 95% interval, with seeds ranging from 32 on AIME 2026 down to 1 on MMLU-Pro and MMMLU.

### Where it gains and where it costs

- Knowledge and multilingual tasks shrink most: MMMLU thinks 65.5% less (1,656 to 571 tokens) and MMLU-Pro 57.3% less, MarkTechPost reports. GPQA-Diamond falls from 12,772 to 7,267 tokens (43.1% cut). IFBench thinks 46.4% less with accuracy nearly flat (79.75% to 79.71%). Long-context retrieval improves: AA-LCR gains 2.25pp (81.75% to 84.00%) on 38.6% fewer tokens. LiveCodeBench v6 edges up 0.07pp while thinking 20.3% less; τ²-bench gives up 1.01pp for a 30.9% cut; Terminal-Bench 2.1 loses 0.56pp, inside its ±4.26 interval, for a 10.7% cut. The costliest trade is AIME 2026: 3.85pp lower accuracy (98.13% to 94.27%) for 30.2% less thinking. Under a 16K-token cap, BottleCap says ThinkingCap scores higher than the base model, with truncated traces down from 0.51% to 0.34% and looping from 0.06% to 0.05%.

### Caveats and deployment

- These are the model maker's own evaluations, run under a shared harness rather than by an independent party, and some correctness comes from the 95% intervals and seed counts rather than from large margins. Compression also stacks with Qwen3.8-27B's reasoning-effort dial: averaged over 11 benchmarks and measured against the base model at xhigh, BottleCap reports the base cuts 52.1% of thinking at medium for -9.16pp versus ThinkingCap's 60.2% for -9.90pp; at low, -55.4%/-9.71pp versus -62.3%/-10.79pp. With thinking off, ThinkingCap trails the base by 5.7pp, and BottleCap recommends xhigh for the best accuracy-to-token balance. The 28B bf16 checkpoint accepts image and text input, and BottleCap publishes five quantized builds (FP8 31 GB, NVFP4 weight-only 21 GB, NVFP4 W4A4 23 GB, GGUF 16–55 GB, MLX 4-bit DWQ 21 GB) that drop in for Qwen3.8-27B on vLLM or SGLang. MarkTechPost reports MTP speculative decoding (3 draft tokens) as accuracy-neutral on AIME 2026, accepting 53% of drafted tokens at about 2.6 tokens per step.

## Primary Sources & Citations

- [BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost](https://www.marktechpost.com/2026/09/24/bottlecap-ai-releases-thinkingcap-qwen3-8-27b-37-2-fewer-thinking-tokens-at-a-0-86pp-accuracy-cost/) — *MarkTechPost* (Reporting)

---

[← Back to front page](https://highsignal.sh/) | [Daily Brief](https://highsignal.sh/brief) | [All stories](https://highsignal.sh/latest)
