In a detailed benchmark analysis, programmer and AI researcher Simon Willison exposed the mathematical limits of the open-weight model Qwen3.8-27B. The experiment focused on evaluating how reliably modern language models handle fundamental arithmetic when the required output deviates from standard numerical digits. Instead of producing digits, the test required the model to perform integer additions and output the results exclusively in spelled-out English words. The evaluation was conducted using a GGUF-quantized version of the model to measure inference performance under standardized conditions. The findings highlighted a significant divide between linguistic fluency and actual arithmetic reasoning in transformer architectures.
The evaluation relied on a comprehensive dataset of 5,070 individual test cases categorized by operand length. Throughout the test run, explicit reasoning tokens, representing deliberate thinking phases prior to the final response, were deliberately disabled. This setup aimed to isolate whether pure autoregressive token prediction could internally maintain multi-step mathematical calculations without scratchpad tokens. The task forced the model to simultaneously compute the exact sum and decode the resulting number into natural language words. Even across moderate operand sizes, the model began exhibiting severe breakdowns in its computation pipeline.
Across all 5,070 test cases, Qwen3.8-27B attained an overall calculation accuracy of just 23.57 percent in the absence of thinking tokens. Performance degraded precipitously as the number of digits in the operands increased. While the model achieved a respectable accuracy rate of 97.04 percent on three-digit additions, its precision dropped sharply with larger values. On test cases involving operands with ten to thirteen digits, accuracy collapsed to a mere 6.44 percent. This steep decline demonstrates that autoregressive token generation fails to preserve carryover logic as operand complexity expands.
A particularly striking observation was the stark divergence between syntactic competence and factual precision. In over 96 percent of the failed test cases, the grammatical structure of the generated English words remained entirely correct. The model consistently formed natural, syntactically coherent number phrases while delivering mathematically incorrect answers. This phenomenon illustrates a structural vulnerability of autoregressive transformers trained to optimize statistical word continuation. The system presents a convincing illusion of competence that conceals a complete breakdown in the underlying mathematical reasoning.
Willison's benchmark underscores fundamental constraints in how transformer networks process arithmetic without dedicated scaffolding. Autoregressive architectures struggle to reliably translate mathematical logic into word tokens when deprived of explicit reasoning phases. The experiment proves that fluent text generation must not be conflated with genuine logical understanding or computational capability. For practical deployments, the findings demonstrate that complex arithmetic in language models must rely on explicit reasoning tokens or specialized external tools. Without these safeguards, even models with 27 billion parameters remain prone to basic calculation failures behind polished grammar.

