Skip to content
AI ConnectPowered by VELENTIS
AI-generated2 min

According to Simon Willison's Benchmark: Qwen3.8-27B Struggles with Word-Based Addition

A benchmark analysis by Simon Willison reveals that Qwen3.8-27B achieves only 23.57 percent accuracy on word-based addition when explicit reasoning tokens are omitted.

This article was AI-generated and published automatically. Context, labelling and all sources at the end of the article.

(KI-generiertes Symbolbild: Gemini / AI Connect)

In a detailed benchmark analysis, programmer and AI researcher Simon Willison exposed the mathematical limits of the open-weight model Qwen3.8-27B. The experiment focused on evaluating how reliably modern language models handle fundamental arithmetic when the required output deviates from standard numerical digits. Instead of producing digits, the test required the model to perform integer additions and output the results exclusively in spelled-out English words. The evaluation was conducted using a GGUF-quantized version of the model to measure inference performance under standardized conditions. The findings highlighted a significant divide between linguistic fluency and actual arithmetic reasoning in transformer architectures.

The evaluation relied on a comprehensive dataset of 5,070 individual test cases categorized by operand length. Throughout the test run, explicit reasoning tokens, representing deliberate thinking phases prior to the final response, were deliberately disabled. This setup aimed to isolate whether pure autoregressive token prediction could internally maintain multi-step mathematical calculations without scratchpad tokens. The task forced the model to simultaneously compute the exact sum and decode the resulting number into natural language words. Even across moderate operand sizes, the model began exhibiting severe breakdowns in its computation pipeline.

Across all 5,070 test cases, Qwen3.8-27B attained an overall calculation accuracy of just 23.57 percent in the absence of thinking tokens. Performance degraded precipitously as the number of digits in the operands increased. While the model achieved a respectable accuracy rate of 97.04 percent on three-digit additions, its precision dropped sharply with larger values. On test cases involving operands with ten to thirteen digits, accuracy collapsed to a mere 6.44 percent. This steep decline demonstrates that autoregressive token generation fails to preserve carryover logic as operand complexity expands.

A particularly striking observation was the stark divergence between syntactic competence and factual precision. In over 96 percent of the failed test cases, the grammatical structure of the generated English words remained entirely correct. The model consistently formed natural, syntactically coherent number phrases while delivering mathematically incorrect answers. This phenomenon illustrates a structural vulnerability of autoregressive transformers trained to optimize statistical word continuation. The system presents a convincing illusion of competence that conceals a complete breakdown in the underlying mathematical reasoning.

Willison's benchmark underscores fundamental constraints in how transformer networks process arithmetic without dedicated scaffolding. Autoregressive architectures struggle to reliably translate mathematical logic into word tokens when deprived of explicit reasoning phases. The experiment proves that fluent text generation must not be conflated with genuine logical understanding or computational capability. For practical deployments, the findings demonstrate that complex arithmetic in language models must rely on explicit reasoning tokens or specialized external tools. Without these safeguards, even models with 27 billion parameters remain prone to basic calculation failures behind polished grammar.

What this means for you

For developers and enterprise users, the benchmark confirms that linguistic fluency cannot guarantee computational reliability without explicit reasoning phases. Teams deploying models for quantitative, financial, or data-driven workloads should never rely on raw text generation for arithmetic tasks. Instead, incorporating explicit chain-of-thought tokens or external code-execution tools remains essential to prevent silent calculation errors hidden behind fluent prose.

Evidence

Solidly sourced
46/100
  • Simon Willison's benchmark experiment on the open-weight model Qwen3.8-27B evaluated 5,070 test cases of integer addition with spelled-out word output.

    single source
  • Without reasoning tokens enabled, Qwen3.8-27B achieved an overall arithmetic accuracy of only 23.57 percent on word additions.

    single source
  • Accuracy dropped from 97.04 percent for three-digit numbers down to 6.44 percent for ten- to thirteen-digit operands.

    single source
  • Despite the high arithmetic error rate, the grammatical structure of the outputs remained correct in over 96 percent of cases.

    single source

The evidence score is computed, not hand-set: from confidence, the number of sources and the share of verified statements.

Source & transparency

As of: October 06, 2026

AI-generatedAI-generated: produced automatically from vetted sources with technical quality checks (source, quote and figure verification); no human sign-off of each item before publication

Sources
1
Verified statements
0 / 4
Evidence score
46Solidly sourced

Want to put this into practice?

We connect you with suitable AI providers from the DACH region, free of charge and without obligation.

What's next?