Skip to content
Alcuse.com
Menu
  • Home
  • Arts Entertainments
  • Auto
  • Business
  • Cryptocurrency
  • Digital Marketing
  • Education
  • Finance
  • Gaming
  • Health Fitness
  • Home Kitchen
  • Legal Law
  • Lifestyle Fashion
  • Medicine
  • Pets
  • Real Estate
  • Relationship
  • Shopping Product Reviews
  • Sports
  • Technology
  • Tours Travel
  • Privacy Policy
  • Contact US
  • Sitemap
Menu

New Apple study challenges whether AI models truly “reason” through problems

Posted on June 12, 2025

In early June, Apple researchers released a study suggesting that simulated reasoning (SR) models, such as OpenAI’s o1 and o3, DeepSeek-R1, and Claude 3.7 Sonnet Thinking, produce outputs consistent with pattern-matching from training data when faced with novel problems requiring systematic thinking. The researchers found similar results to a recent study by the United States of America Mathematical Olympiad (USAMO) in April, showing that these same models achieved low scores on novel mathematical proofs.

The new study, titled “The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity,” comes from a team at Apple led by Parshin Shojaee and Iman Mirzadeh, and it includes contributions from Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar.

The researchers examined what they call “large reasoning models” (LRMs), which attempt to simulate a logical reasoning process by producing a deliberative text output sometimes called “chain-of-thought reasoning” that ostensibly assists with solving problems in a step-by-step fashion.

To do that, they pitted the AI models against four classic puzzles—Tower of Hanoi (moving disks between pegs), checkers jumping (eliminating pieces), river crossing (transporting items with constraints), and blocks world (stacking blocks)—scaling them from trivially easy (like one-disk Hanoi) to extremely complex (20-disk Hanoi requiring over a million moves).

Figure 1 from Apple’s “The Illusion of Thinking” research paper.


Credit:

Apple

“Current evaluations primarily focus on established mathematical and coding benchmarks, emphasizing final answer accuracy,” the researchers write. In other words, today’s tests only care if the model gets the right answer to math or coding problems that may already be in its training data—they don’t examine whether the model actually reasoned its way to that answer or simply pattern-matched from examples it had seen before.

Ultimately, the researchers found results consistent with the aforementioned USAMO research, showing that these same models achieved mostly under 5 percent on novel mathematical proofs, with only one model reaching 25 percent, and not a single perfect proof among nearly 200 attempts. Both research teams documented severe performance degradation on problems requiring extended systematic reasoning.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Recent Posts

  • What should I expect from a Sexual harassment lawyer consultation?
  • What are the benefits of software composition analysis tools?
  • What is natural justice in unjust dismissal Canada?
  • Can users find factories on global sources website?
  • 구글 검색 누락 카페24도 생기나요?
  • 음식은 강남달토에서 빨리 나오는 편이야?
  • 강남 가라오케 중심가예요?
  • 강남호빠는 처음 가는 사람에게 추천되나요?
  • What questions should I ask an employment lawyer Toronto?
  • Can the TikTok API improve content strategy?
  • Is commercial energy management worth the investment in Bermuda Dunes CA?
  • Decen Masters: Crowdsourcing Alpha in an Automated World
  • Can a Locksmith near me install window security locks?
  • What is the cost of gutters installation?
  • Does termination pay apply to executive-level employees?
  • 해외스포츠중계는 스마트 TV에서 볼 수 있나요?
  • 무료 스포츠중계 사이트 오류 잦나요?
  • 오피 계약 해지 시 중개수수료는 어떻게 되나요?
  • Is transportation included with a Curacao jetski tour?
  • How is the EV charging industry supporting fleet electrification?

Archives

  • July 2026
  • May 2026
  • April 2026
  • March 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • November 2023
  • October 2023

Categories

  • Arts Entertainments
  • Auto
  • Business
  • Cryptocurrency
  • Digital Marketing
  • Education
  • Finance
  • Gaming
  • Health Fitness
  • Home Kitchen
  • Legal Law
  • Lifestyle Fashion
  • Medicine
  • Pets
  • Real Estate
  • Relationship
  • Shopping Product Reviews
  • Sports
  • Technology
  • Tours Travel
Slot gacor hari ini
Slot online
TOTO SLOT
Slot
©2026 Alcuse.com | Design: Newspaperly WordPress Theme