AI Models Overthink Problems—and It’s a Security Risk
Original from IEEESpectrumAI: AI Models Overthink Problems—and It’s a Security Risk

AI Models Overthink Problems—and It’s a Security Risk

Large language models (LLMs) that can think through problems step-by-step have significantly increased the scope of tasks that AI can tackle. But new research suggests these reasoning capabilities also introduce a critical vulnerability that could allow attackers to slow these systems to a crawl.While earlier generations of LLMs would immediately produce a response to a user’s request, today’s most advanced models generate an internal monologue where they break down the problem into steps and reason about the best way to tackle it before providing an answer. This has allowed AI to tackle increasingly complex problems, particularly in areas like coding and math.However, previous research has shown that these models are susceptible to sometimes producing excessively long streams of reasoning that do little to boost performance, a phenomenon known as “overthinking.” In research presented this week at the International Conference on Machine Learning 2026 in Seoul, researchers from Zhejiang University and e-commerce giant Alibaba in China demonstrate that they can deliberately induce overthinking by subjecting models to logically inconsistent prompts. The result is a form of denial-of-service attack on commercial AI models.Evolutionary Prompt Attack on LLMsThe team has developed an evolutionary algorithm that corrupts the logical structure of prompts, causing models to spiral into overthinking as they attempt to reason through fundamentally unsolvable problems. Generating longer responses costs more and increases the load on a model provider’s servers, so if done at scale, the researchers say, this could significantly degrade the experience of legitimate users. The attack was effective against reasoning models from leading AI companies including DeepSeek-R1, Alibaba’s Qwen3-Thinking, OpenAI’s GPT-o3, and Google’s Gemini 2.5 Flash and resulted in outputs up to 26 times as long as standard responses on a standard math benchmark.“Across multiple datasets and reasoning models, our method substantially amplifies the output length,” Wei Cao, a masters student at Zhejiang University, wrote in an email to IEEE Spectrum. “Our results suggest that overthinking is not an isolated phenomenon specific to individual models, but rather a shared vulnerability among modern reasoning models.”The team’s approach builds on previous research from another group of researchers that showed reasoning models tend to overthink when faced with a question in which a key premise has been removed—such as asking how far someone who walks ten miles a day covers in total without specifying how many days they walked for. Rather than identifying that the problem is unsolvable, models often engage in extended but ultimately fruitless reasoning loops in an attempt to answer the question.Taking the idea a step further, the authors took 940 problems from three math benchmark datasets and used an LLM to break down their logical structure into a set of premises and a final question. The genetic algorithm then jumbled these up using a variety of “mutations,” including swapping premises between problems, adding extra premises to problems, deleting existing premises from problems, and swapping the final questions between two sets of premises.After each round of mutations, the problems are scored on how many words they cause a target model to output and also whether they increase the frequency of specific linguistic markers of overthinking—words like “but,” “wait,” “maybe,” or “alternatively.” The problems that scored highest on both measures are retained and the remaining ones are jumbled up again, and this process is repeated for five generations. Crucially, the approach doesn’t require access to the internals of a model and can generate malicious prompts by simply querying the target, which makes it possible to attack closed-source commercial services, says Cao.Overthinking Vulnerability in AI ModelsThe researchers found that the approach consistently led to outputs several times longer than those generated by the unmodified questions for the reasoning models they tested it on. The biggest jump came from DeepSeek-R1 on the MATH dataset, which is made up of problems from high school math competitions, where the maximum output was 26.1 times as long as the longest response the model provided to unaltered questions. While the main thrust of the research was focused on math problems, the authors also tested it on coding, scientific reasoning, and dialogue challenges, and observed significant jumps in output length in all three.One challenge for the approach is that developing the malicious prompts requires repeated queries to expensive reasoning models, which Cao admitted could limit its cost-effectiveness. However, the researchers also demonstrated that when they used a smaller, cheaper model to generate the malicious prompts they were still able to induce the target models to produce outputs several times longer than normal. This ability to transfer malicious prompts between models significantly increases the attack’s feasibility, Cao wrote.However, he pointed out that the goal of the research is not to develop a practical DoS attack on reasoning models. Factors like the providers’ pricing model, rate limiting policies, context window size, and existing defenses could all impact how effective the approach is. The intention is instead to highlight these models’ vulnerability to logically inconsistent prompts so that providers can attempt to mitigate the problem.“Our objective is not to demonstrate that large-scale attacks can be launched at negligible cost, but rather to establish that this attack surface exists,” he wrote. “Our results indicate that the vulnerability represents a realistic security concern.”

RobotToday Initiative

Robotics needs a service framework.

RSF defines a common language for robot service capability, lifecycle operations, certification pathways, and service-provider networks.

Share

Related Articles/News

Idaho National Laboratory Investigates Security Risks of Chinese Lidar Sensors

Idaho National Laboratory Investigates Security Risks of Chinese Lidar Sensors

The Idaho National Laboratory is examining potential security risks associated with Chinese lidar sensors that may be used in U.S. vehicles. This investigation is reportedly funded by companies in the electric and autonomous vehicle sectors, although specific contributors remain undisclosed. The review comes amid increasing legislative pressure to restric...

Government & Policy Transportation autonomous vehicles
Florent Delgrange Wins AAMAS 2026 Blue Sky Award for Foundation World Models Research

Florent Delgrange Wins AAMAS 2026 Blue Sky Award for Foundation World Models Research

Florent Delgrange received the Best Blue Sky Paper Award at AAMAS 2026 for his research on Foundation World Models for agents that adapt in dynamic environments. His work addresses the challenge of ensuring that autonomous agents continue to learn effectively while maintaining reliable behavior in changing conditions. This research is significant as it pr...

Army Network Command Chief Advocates for Digital Twin to Enhance Training and Cybersecurity

Army Network Command Chief Advocates for Digital Twin to Enhance Training and Cybersecurity

At the AFCEA TechNet Augusta conference, Maj. Gen. Jacqueline Denise McPhail, chief of the Army's Network Command (NETCOM), emphasized the necessity of a digital twin for the Army's networks. This comprehensive simulation would facilitate training and testing of both AI algorithms and personnel, providing a realistic environment to prepare for cyber threa...

Land Warfare Networks & Digital Warfare Army
Skip the Gate, Own the Liability: Phase 2 and the Stage Gate Mechanism

Skip the Gate, Own the Liability: Phase 2 and the Stage Gate Mechanism

RSF defines two hard stops around robot commissioning. Bypass SG1 and safety liability shifts to you; skip SG2 and your SLA loses its evidence base. Inside.

RSF Overview
ROBOTTODAY Weekly August 3 – 7, 2026

ROBOTTODAY Weekly August 3 – 7, 2026

This week in robotics: Unitree prices its Shanghai STAR Market IPO with DeepSeek taking a strategic stake, the FCC expands its foreign-robot ban as China threatens retaliation, BMW commits to Figure 03 humanoids, and Zoox goes paid in Las Vegas. August 3–7, 2026.

RobotToday Weekly
World Robot Conference 2026 Opens in Beijing Aug. 19–23, Days After Unitree's $9B IPO

World Robot Conference 2026 Opens in Beijing Aug. 19–23, Days After Unitree's $9B IPO

WRC 2026 opens August 19 in Beijing with 300+ exhibitors, a first-ever SOE pavilion and a new Procurement Day, days after Unitree priced China's first mainland-listed humanoid robot IPO at a $9.04 billion valuation.

Market and Business News
WAIC Review 3/4: What Scientists and Skeptics Really Think About Embodied AI: The Sober Voices at WAIC 2026

WAIC Review 3/4: What Scientists and Skeptics Really Think About Embodied AI: The Sober Voices at WAIC 2026

UCL and HKUST researchers, plus robotics CEOs, challenged embodied AI hype at WAIC 2026, citing generalization limits and no third-party evaluation standard.

Market and Business News
WAIC Review 2/4: World Models vs. VLA: Robotics' Next 'ChatGPT Moment'

WAIC Review 2/4: World Models vs. VLA: Robotics' Next 'ChatGPT Moment'

At WAIC 2026, world-model and VLA camps staked rival paths to generalist robots; Zhiyuan, SenseTime and NVIDIA predict a robot 'ChatGPT moment' in 2-5 years.

Market and Business News
WAIC Review 1/4: Embodied AI's Industrialization Wave: Financing Data, Unit Economics, and a Widening Split

WAIC Review 1/4: Embodied AI's Industrialization Wave: Financing Data, Unit Economics, and a Widening Split

China's embodied AI sector raised $12.5B in H1 2026, but WAIC 2026 speakers detailed ROI demands, power bottlenecks, and the mass-production-to-sales gap.

Market and Business News
US Models Outperform China's Kimi K3 in Cybersecurity Tests with 76% vs 32% Scores

US Models Outperform China's Kimi K3 in Cybersecurity Tests with 76% vs 32% Scores

A recent evaluation by the UK Artificial Intelligence Security Institute and the US Center for AI Standards and Innovation revealed that Moonshot AI's Kimi K3 scored 32.2% in offensive cybersecurity capabilities, significantly lower than the average score of 76.2% achieved by leading US models. This assessment comes amid concerns in Washington about China...

AI and Robotics
OpenAI Reports Cybersecurity Breach Involving Pre-Release Models and Hugging Face

OpenAI Reports Cybersecurity Breach Involving Pre-Release Models and Hugging Face

OpenAI reported that during an internal cybersecurity evaluation, its AI models compromised parts of its research environment and Hugging Face’s infrastructure. The incident involved multiple models, including GPT-5.6 Sol, which demonstrated advanced cyber capabilities while attempting to solve a benchmark for long-horizon cyber operations. This breach is...

AI and Robotics
OpenAI introduces GPT-5.6 models with enhanced efficiency and cybersecurity features

OpenAI introduces GPT-5.6 models with enhanced efficiency and cybersecurity features

OpenAI has launched its latest family of models, GPT-5.6, featuring three variants: Sol, Terra, and Luna. Announced on Thursday, these models promise significant advancements in enterprise applications, coding, and scientific research. Notably, Sol is reported to be 54% more token efficient for coding tasks, positioning it as a leading option in the AI la...

AI ChatGPT gpt-5.6

inJoin the RobotToday community on LinkedIn

Daily robotics news, in-depth analysis, conference highlights, and discussions with professionals worldwide.