← Back to feed News · August 26, 2026 · 1 min
News

Local AI Benchmarks: Why High Token Speed Delays Actual Answers

High token generation speeds on laptops do not guarantee quick answers. Benchmark tests show compact local models finish tasks faster and outscore massive cloud systems simply by eliminating unnecessary internal reasoning.

Photo: Tom Tunguz

Stop looking at raw generation speed when picking a local artificial intelligence model for your laptop. A blistering token-per-second rate looks great on a spec sheet, but in real daily use, it frequently leaves you staring at an empty screen far longer than an apparently slower setup.

As venture capitalist Tomasz Tunguz pointed out, Qwen3.6-35B-A3B churns out text at 113.4 tokens per second—more than 2.2 times faster than Qwen3.8-27B at 51.9 tokens per second. Yet the seemingly fast model finishes the actual job slower, taking an average of 10.0 seconds versus 7.2 seconds for the 27B model across 25 benchmarked tasks. The reason comes down to background bloat: the 35B model burns through 1,143 tokens on internal deliberation—3.1 times more preliminary "thinking" than the 369 tokens required by the 27B model—before finally delivering a result. In one extreme case, it spun through 993 tokens of hidden reasoning just to output a six-word classification.

Judging local models by raw typing speed is practically useless; total time-to-answer is the only metric that matters at your desk. Better yet, you do not need expensive cloud subscriptions to get top-tier results. Benchmark data from Artificial Analysis ranks the compact Qwen3.8-27B at number one out of 135 models on its Intelligence Index with a score of 52, outperforming Z.ai's monolithic 753-billion-parameter GLM-5.2 cloud giant at 51. Modern laptop silicon running concise, well-reasoned local models can match frontier cloud intelligence without wasting your battery and time on pointless internal monologues.

Source Tom Tunguz → © 2026 «Gadgety». Full or partial copying — with a link to this page.