Close Menu
MyAppsPlus

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Anthropic ‘formalizes’ Fermat’s Last Theorem like never before using Claude — but it still took 11 days to write out

    September 13, 2026

    Never use harsh cleaning products to clean your smartphone — do this instead

    September 13, 2026

    MAKERphone 2.0 Lets You Build a 4G Phone and Vibe-Code Its Apps

    September 13, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    MyAppsPlusMyAppsPlus
    Sunday, September 13
    • Home
    • Breaking Tech
    • Apps & Software
    • AI & Automation
    • Android
    • iPhone & iOS
    • More
      • Reviews
      • How-To Guides
      • Deals & Discounts
      • Shop
    MyAppsPlus
    Home»AI & Automation»Are We on the Verge of an Intelligence Explosion? Maybe Not.
    AI & Automation

    Are We on the Verge of an Intelligence Explosion? Maybe Not.

    myappsplusBy myappsplusAugust 28, 2026004 Mins Read
    Share Facebook Twitter Pinterest Copy Link LinkedIn Tumblr Email Telegram WhatsApp
    Follow Us
    Google News Flipboard
    Are We on the Verge of an Intelligence Explosion? Maybe Not.
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    There’s growing excitement in the AI industry about the idea that today’s leading models could build the next generation of the technology. But a new study recently found top AI agents struggle on the kind of genuinely open-ended research problems required to push the field forward.

    Large language models have made rapid progress in many of the day-to-day jobs involved in machine learning research, such as writing code, generating and curating data, and running experiments. Last year, startup Sakana AI’s AI Scientist-v2 even managed to write a paper that cleared peer review for the prestigious International Conference on Learning Representations.

    These advances have led to speculation that models are close to being able to build better versions of themselves with little human oversight—a process called recursive self-improvement. The idea underpins predictions that we may be on the verge of an intelligence explosion that could quickly lead to AI superintelligence.

    In a recent paper, researchers put the idea to the test using a new approach they call shadow evaluations. This involves taking the research question from a high-quality, unpublished machine learning paper and asking AI agents to solve the problem. The original paper’s authors then grade the results. When the team tested Claude Opus 4.8 on two papers submitted to the prestigious machine-learning conference NeurIPS 2026, the authors rejected both.

    “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” Sayash Kapoor from Princton University, who co-led the study, told MIT Technology Review.

    Previous efforts to get AI agents to do machine learning research have often targeted problems focused on engineering, such as reproducing previous research or training smaller models against a benchmark.

    In the new experiments, the researchers challenged models with more open-ended tasks that required them to devise hypotheses, decide what evidence is needed to validate them, judge when a research direction was fruitless, and go back to the drawing board.

    One research question was whether the personality traits a language model displays can be measured and adjusted by observing and editing its weights; the other attempted to detect when a model that works with tabular data has quietly stopped being reliable.

    Be Part of the Future

    100% Free.No Spam.Unsubscribe any time.

    In each case, the AI researchers were given $3,000 of API credits, a budget for time on GPUs to run machine learning experiments, a dedicated Linux virtual machine, and unrestricted internet access. They were then given six days to produce a paper that could pass NeurIPS’ stringent peer-review criteria.

    In both cases, the models got a good start. The agents surveyed the literature effectively, came up with opening hypotheses that mirrored those of the authors, and successfully ran hundreds of experiments.

    But they quickly went off the rails. Although they could monitor their own use of time and their API and GPU budgets, they rushed through the process. One left 110 hours of unused time on the clock, and both failed to spend even 50 percent of their API budget.

    Both agents also settled on a research direction within just 10 hours and failed to change approaches despite repeated negative feedback from another AI designed to review drafts of their papers. The reviewer identified problems the human authors would also flag in the final paper, but the models simply added caveats to their findings and ploughed on. Ultimately the papers received a “strong reject” and a “reject” decision from the human reviewers based on NeurIPS grading protocol.

    The authors admit their approach has limitations. The reviewers knew AI had written the submissions, and some of the team are on record as doubting an imminent intelligence explosion. The original human-authored papers also took far longer than six days to produce and used many more GPU hours to reach their conclusions (though, as the researchers note, the models did not use their allocated budget in any case).

    Nonetheless, the results suggest that today’s models still have some way to go before they can tackle the most challenging problems in machine learning research. Until that happens, the dream of recursive self-improvement is likely to remain a distant prospect.

    Edd Gent
    Edd Gent
    Edd is a freelance science and technology writer based in Bangalore, India. His main areas of interest are engineering, computing, and biology, with a particular focus on the intersections between the three.

    A wooden judge's gavel sits on a marble table

    An ‘AI Legal Team’ Has Won Its First Case. It’s a Rare Victory for Access to Justice.

    Colorful onscreen static

    Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy.

    DeepMind’s Weather AI Predicts Hurricanes a Day Earlier Than Traditional Forecasting

    Shelly Fan
    Aug 17, 2026
    Future

    Genevieve Grant
    Aug 27, 2026
    Future

    Liming Zhu
    Aug 20, 2026
    Artificial Intelligence

    Explosion intelligence Maybe Verge
    Follow on Google News Follow on Flipboard
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    myappsplus
    • Website

    Related Posts

    Machine learning algorithm sets XRP price for October 1, 2026

    September 13, 2026

    Obama Sounds The Alarm On ‘Dangerous’ AI, Urges Dems To Take Action: NYT

    September 13, 2026

    This Artificial Intelligence (AI) Chip Stock Will Soar After Sept. 30 (Hint: It’s Not Micron)

    September 13, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    The 6 AI-free Linux distros I recommend most

    August 19, 20264 Views

    AI, automation, robot dogs ensure on-site nuclear safety

    September 7, 20262 Views

    This tiny AI box could save me from upgrading my perfectly good laptop

    September 6, 20262 Views
    Latest Reviews

    Why Fluper is the No.1 Mobile App Development Company in the UAE, Saudi Arabia, and the Middle East.

    myappsplusAugust 18, 2026

    How New Kuwait And Indonesia Tech Deals At Baker Hughes (BKR) Have Changed Its Investment Story

    myappsplusAugust 18, 2026

    Google is reportedly planning to move all Pixel production out of China

    myappsplusAugust 18, 2026
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Why Fluper is the No.1 Mobile App Development Company in the UAE, Saudi Arabia, and the Middle East.

    August 18, 20260 Views

    How New Kuwait And Indonesia Tech Deals At Baker Hughes (BKR) Have Changed Its Investment Story

    August 18, 20260 Views

    Google is reportedly planning to move all Pixel production out of China

    August 18, 20260 Views
    Our Picks

    Anthropic ‘formalizes’ Fermat’s Last Theorem like never before using Claude — but it still took 11 days to write out

    September 13, 2026

    Never use harsh cleaning products to clean your smartphone — do this instead

    September 13, 2026

    MAKERphone 2.0 Lets You Build a 4G Phone and Vibe-Code Its Apps

    September 13, 2026

    Subscribe to Updates

    Subscribe to our newsletter and get the latest tech news, app updates, AI trends, smartphone reviews, and exclusive deals delivered straight to your inbox.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms & Conditions
    © 2026 MyAppsPlus. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.