Thought the AI was getting smarter? Turns out, it was just getting stuck in overthinking.

An abstract image representing an AI unable to make a decision and lost in thought, amidst a complex circuit diagram with numerous tangled arrows.
AI Summary

A phenomenon has been discovered where 'quantization,' a technique used to shrink AI models for efficiency, causes the AI to overthink, leading it to find the correct answer mid-process but fail in its final response.

Imagine this: You are solving a math problem and you already know the answer in your head. But then, doubts like “What if this isn’t it?” keep piling up, and before you know it, the exam time has run out, and you haven’t written a single thing on your answer sheet. It turns out that exactly the same thing is happening in the world of Artificial Intelligence (AI) recently.

It is commonly known that the larger an AI model, the smarter it is. However, to run these models lightly on users’ computers or smartphones, they undergo a process called “Quantization” (a technique that increases efficiency by reducing the precision of the model). Surprisingly, it has been found that AI models that have gone through this process often get stuck in “overthinking” more than necessary and miss the correct answer themselves.

Why does this matter?

The importance of AI in our lives is growing. However, running huge AI models as they are requires too much power and high-performance equipment. Therefore, developers want to reduce the size of the models to use AI smoothly on personal devices like smartphones.

This research points out a “hidden side effect” we have been missing while trying to make AI efficient. If an AI unnecessarily prolongs its Chain of Thought (the process where an AI steps through logical steps to reach an answer) because it doubts itself even after finding the answer, we end up getting slower and more inaccurate responses. This is not just a technical issue, but an important discovery directly linked to the quality of the AI services we use every day. Ref 14

Easy to understand: The dilemma of a ‘perfectionist student’

Let’s use an analogy. Imagine a student who is very good at studying, but even though they already know the answer to an exam question, they get anxious, thinking, “Did I make a mistake?” So they erase the answer they wrote, solve it again, try to find a new way to solve it, and then hear the exam ending bell.

According to the Meta AI research team, quantized logical reasoning models exhibit this exact behavior. Ref 3 These models reach the correct answer halfway through solving the problem. However, they are not confident in that answer, so they branch out into unnecessary new lines of reasoning, falling into a swamp of doubt. Consequently, they either fail to provide an answer or reach an incorrect conclusion. In fact, the study found that 52% of the cases where quantized models failed were instances where the model already knew the correct answer during the intermediate process. Ref 1, Ref 5

Where are we now: Cutting off excessive thinking

To solve this problem, the researchers tried an interesting experiment. They detected specific “Thinking markers” that trigger the AI to start excessive doubting—such as “but” or “wait”—and applied a “Logit penalty” (a technique to lower the probability of outputting specific words or expressions) to them. Ref 2, Ref 13

In simple terms, it’s like having a teacher by their side who says, “That’s already the right answer, just write it down!” every time the student starts to overthink unnecessarily. Applying this technique reduced the AI’s unnecessary thinking time (the length of the logical reasoning process) by 12% to 23% across various models and benchmarks, resulting in more efficient and accurate responses. Ref 2

What will happen in the future?

This research provides a major insight into the direction of AI development. It shows that simply reducing the capacity of AI models is not always the best path. In the future, developers will need to conduct deeper research into “smart compression techniques” that ensure the intrinsic logical reasoning ability is not compromised even when shrinking the size of AI models.

We will now focus not on how long of a text an AI can write, but on how accurately it can make judgments without fluff. If you feel that your AI conversation feels a bit smarter and clearer next time, it might be because the AI has been trained to stop overthinking unnecessarily.

MindTickleBytes’ AI Reporter’s Take

This research by Meta AI is very interesting because it applies the life truth that ‘thinking more does not necessarily lead to better results’ to AI. The fact that optimizing AI models can actually degrade a model’s ‘self-objectification’ ability is a variable that must be considered in the future AI development process.

References

  1. Quantized Reasoning Models Think They Need to Think Longer…
  2. Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
  3. Quantized Reasoning Models Think They Need to Think Longer…
  4. Quantized Reasoning Models Think They Need to Think Longer…
  5. Meta study finds quantized reasoning models struggle with …
  6. Quantized Reasoning Models Think They Need To Think Longer …
  7. Logit-Bias “Overthinking” Penalties in llama.cppQuantizations
  8. QuantizedReasoningModelsThinkTheyNeedtoThinkLonger…
  9. QuantizedReasoningModelsThinkTheyNeedtoThinkLonger…
  10. When aquantizedreasoningmodelfails, the natural assumption is…
  11. Think–Answer Mismatch inReasoningModels
  12. Quantized Reasoning Models Think They Need to Think Longer …
  13. Quantization Inflates Reasoning: Token Inflation as a Hidden …
AD
Test Your Understanding
Q1. According to the research, what is the maximum percentage of cases where a quantized AI model gets a question wrong despite already finding the correct answer in its intermediate process?
  • 12%
  • 23%
  • 52%
According to Meta AI's research, up to 52% of failure cases for quantized models occurred because the model had already found the correct answer in an intermediate step but failed to output it in the final response.
Q2. What technique did the researchers introduce to prevent the AI from overthinking?
  • Logit penalties applied without additional training
  • Additional large-scale data retraining
  • Massive expansion of the AI model size
The researchers applied logit penalties, which do not require additional training, to signals that trigger the AI to start unnecessary overthinking, such as 'but' or 'wait'.
Q3. What is the main point of this research?
  • Quantization unconditionally improves AI performance
  • Quantized AI often overthinks while solving logical problems and ends up failing
  • All AI models do not need to be quantized
This research revealed that quantization, which shrinks AI model size, lowers the accuracy of logical reasoning models and has the side effect of unnecessarily lengthening the reasoning process.
Thought the AI was getting ...
0:00