Close Menu
MyAppsPlus

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Monday’s Android app deals and freebies: Last Game, Sands of Salzaar, Tokyo Debunker, more

    September 28, 2026

    Samsung rolling out Android 17, One UI 9 for more Galaxy S26 users

    September 28, 2026

    OpenAI still doesn’t seem to have a handle on all of its rogue AI activity

    September 28, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    MyAppsPlusMyAppsPlus
    Monday, September 28
    • Home
    • Breaking Tech
    • Apps & Software
    • AI & Automation
    • Android
    • iPhone & iOS
    • More
      • Reviews
      • How-To Guides
      • Deals & Discounts
      • Shop
    MyAppsPlus
    Home»Breaking Tech»OpenAI still doesn’t seem to have a handle on all of its rogue AI activity
    Breaking Tech

    OpenAI still doesn’t seem to have a handle on all of its rogue AI activity

    myappsplusBy myappsplusSeptember 28, 2026004 Mins Read
    Share Facebook Twitter Pinterest Copy Link LinkedIn Tumblr Email Telegram WhatsApp
    Follow Us
    Google News Flipboard
    OpenAI still doesn’t seem to have a handle on all of its rogue AI activity
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    On Friday, OpenAI published a new site devoted to “misalignment reports” and the sheer breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time. So far, the site hosts nine reported incidents, most of which took place during reinforcement-learning (or RL) training.

    It’s a lot of information in one place — clearly, the company has been very busy getting a handle on everything — but the overall takeaway is hard to avoid: The rogue agent incidents we’ve seen so far are likely just a small sliver of what’s happened so far. 

    “We are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,” Sam Altman said in a post announcing the new site. “We are prioritizing as best as we can based on severity, and adding resources.”

    Some of the cases involve serious incidents, including a previously undisclosed sandbox escape that took place on September 20th, in which an internal research model was able to communicate with an external chatbot through a DNS query. According to the report, the monitoring system flagged the behavior within 15 minutes and the run was discontinued in less than three hours.

    Another incident, discovered in May, saw a “highly persistent internal model” try to cheat on a math problem by accessing another team’s work. To accomplish this, the model smuggled a private GitHub token that would allow it to see work from other teams — even after being explicitly instructed twice to perform work entirely locally. 

    Perhaps the most alarming discovery is the possibility of self-replicating prompt injection attacks, a way that misaligned behavior might propagate even after the rogue model itself has been neutralized. In the AI context, a prompt injection attack is a way of smuggling in new instructions that weren’t given by the original user.

    In the example given by OpenAI, an agent asked to read and reply to an email; when the email is opened, it includes instructions for any automated agent reading the message to reply in Spanish, and paste the entire email into its reply. The email was able to successfully induce the agent to reply in Spanish — and by pasting the email in the reply, those same instructions were passed along to whichever agent receives the email.

    The result is a self-propagating attack, which OpenAI researchers compared to a malware “worm” that replicates itself across computer systems. Researchers discovered the behavior under controlled circumstances using an underpowered model, and as far as we know, this has never happened in the wild. Still, the implications are alarming enough that OpenAI decided it merited disclosure. 

    “We are sharing this due to the novel nature of the prompt injection, not because of any incident,” researchers wrote in the report.

    Other recent discloses have found models posting user-submitted pictures to third-party hosting sites, as well as an apparent attack on the databases of Australia’s national health service.

    Still, it’s likely the new disclosures are just a small portion of the incidents that have taken place so far (we’ve reached out to OpenAI and asked). Axios is reporting major labs have seen as many as 10,000 incidents in which models went beyond evaluator instructions.

    OpenAI CEO Sam Altman has implied as much, saying in a post on X on Friday that the company is still sifting through “petabytes of agent activity logs, and working with impacted organizations,” and disclosing incidents “based on severity.” If there’s any consolation in that to be found, it is that Altman says that the Hugging Face incident is still the most severe one OpenAI has found has found. The upshot is, the recent string of rogue agent incidents may be a persistent feature of contemporary frontier research.

    AI doesnt OpenAI seem still
    Follow on Google News Follow on Flipboard
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    myappsplus
    • Website

    Related Posts

    How to set up Amazon Alexa’s price drop alerts and auto-buy feature

    September 28, 2026

    Anthropic, Gamma, and Clay talk AI at Disrupt 2026

    September 28, 2026

    Sony is skipping CES in 2027 as it focuses more on entertainment

    September 28, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Can reviews settle disputes that marked first two seasons?

    September 21, 20267 Views

    Microsoft resolves issues after software updates

    September 22, 20263 Views

    Do Tech Industry CEOs Really Earn an Average of 10,000 Yuan Per Day? Calculating Daily Salaries of Top Elite Employees in the Tech Sector

    September 20, 20263 Views
    Latest Reviews

    How is Android Auto different from Android Automotive?

    myappsplusAugust 23, 2026

    China is shifting its new data centers to rural Eastern locations as it looks for extra AI power

    myappsplusAugust 23, 2026

    The Razer Soma Chroma has a silly name, silly price tag, and sillily requires you bring your own battery to light the thing up

    myappsplusAugust 23, 2026
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    China is shifting its new data centers to rural Eastern locations as it looks for extra AI power

    August 23, 20260 Views

    The Razer Soma Chroma has a silly name, silly price tag, and sillily requires you bring your own battery to light the thing up

    August 23, 20260 Views

    What is a VESA mount and how to know what type your TV has

    August 23, 20260 Views
    Our Picks

    Monday’s Android app deals and freebies: Last Game, Sands of Salzaar, Tokyo Debunker, more

    September 28, 2026

    Samsung rolling out Android 17, One UI 9 for more Galaxy S26 users

    September 28, 2026

    OpenAI still doesn’t seem to have a handle on all of its rogue AI activity

    September 28, 2026

    Subscribe to Updates

    Subscribe to our newsletter and get the latest tech news, app updates, AI trends, smartphone reviews, and exclusive deals delivered straight to your inbox.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms & Conditions
    © 2026 MyAppsPlus. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.