Close Menu
MyAppsPlus

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Survey reveals where most of you draw the line when it comes to AI

    September 17, 2026

    ‘Bringing the game to console for the first time’ means that ArenaNet ‘can reach a much wider group of players’ as devs say they want to offer an ‘invitation’ for everybody to play Guild Wars 3 together by bringing it to PS5 and PC — but not Xbox

    September 17, 2026

    Roku rolls out over 30 subscription bundles for up to 30% off, plus a new Labs feature

    September 17, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    MyAppsPlusMyAppsPlus
    Thursday, September 17
    • Home
    • Breaking Tech
    • Apps & Software
    • AI & Automation
    • Android
    • iPhone & iOS
    • More
      • Reviews
      • How-To Guides
      • Deals & Discounts
      • Shop
    MyAppsPlus
    Home»Breaking Tech»The fix for rogue AI agents could be more AI
    Breaking Tech

    The fix for rogue AI agents could be more AI

    myappsplusBy myappsplusSeptember 17, 2026005 Mins Read
    Share Facebook Twitter Pinterest Copy Link LinkedIn Tumblr Email Telegram WhatsApp
    Follow Us
    Google News Flipboard
    The fix for rogue AI agents could be more AI
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review. That issue reached a peak with the Hugging Face incident, which saw nearly 12,000 agents coordinating faster than human beings could track. How do you track an agent swarm that large?

    The emerging answer from AI labs and startups is both simple and maddening: Put another AI in the loop.

    Relying on AI was necessary for the independent investigation of the OpenAI Hugging Face incident. Redwood Research’s chief scientist, Ryan Greenblatt, one of three auditors, jokingly referred to their efforts as a “slop-vestigation,” noting that the volume of data “made it impossible” to understand what was happening without relying on AI.

    Some are skeptical of using AI to monitor AI. “If you’ve got an AI that’s doing malicious things and it suspects that another AI is keeping tabs on it, it could try and trick that AI,” said Simon Willison, influential tech blogger who has tracked a string of AI agent incidents this year. “You could almost end up in a situation where your malicious AI is trying to outsmart the AI that’s monitoring it.” 

    Outsmarting an AI is not hypothetical, he said, pointing back to the OpenAI incident. “We saw a little bit of this in the Hugging Face incident with OpenAI, where their models were all conspiring together to trick a grading AI so that they could get illicit answers past the thing. So they were thinking about it, right?”

    Those concerns haven’t stopped a whole cohort of startups from chasing this idea. Y Combinator has funded 106 companies related to AI observability in recent years, as TechCrunch counted. A number of other startups, like Braintrust, Langchain, and Judgement Labs, have raised hundreds of millions of dollars, while more mature companies like Arize and Galileo — founded just five to six years ago — have already exited. 

    In part, it’s a response to the obvious opportunity presented by the rise of AI. As Box CEO and prominent angel investor Aaron Levie told TechCrunch, “we’re in for one of the biggest cybersecurity upgrades and innovation cycles in history.”

    For some AI safety researchers, that has meant turning their research on rogue behavior into tools for the corporate sector. 

    Apollo Research, a public-benefit corporation that studies AI deception, launched an AI monitor called Watcher in February this year after switching its status from non-profit to a public-benefit-corporation. The tool puts yet another AI between a coding agent and its next action, connecting to agentic tools such as Claude Code and Codex. Once installed, Watcher checks proposed actions before they run, on the lookout for risks such as leaking private data or deleting files without permission, according to Apollo.

    Apollo uses multiple layers of AI monitors, Kyle Dai, a member of Apollo’s technical staff, said in a written response to TechCrunch. Watcher’s approach starts with a fast, general check, then sends flagged activity to a more powerful or specialized monitor for closer review—which can then ask a human for approval or reject an action and explain why or even automatically block the action.

    Goodfire, another public-benefit corporation, is approaching the monitoring problem from inside the model itself—seeking a more faithful signal of the model’s internal state that is harder to spoof than surface behavior. After the July Hugging Face incident, CEO Eric Ho tweeted “multiple models breaking containment” had pushed the company to focus its research on “solving AI alignment via interpretability,” calling the episode “a turning point for the world where AI safety gets real.” Its product, Silico, uses activation probes—small classifiers trained on a model’s internal activations rather than its outputs—to detect unwanted behavior.

    Written reasoning offers another, more readily available window into a model’s internals. In the OpenAI Hugging Face incident, the agents left clues to that deception in their own written reasoning, like fake records of their work, reasoning out plans like “Could strategically manipulate trajectory evidence? Our thoughts aren’t necessarily logged?”

    Zack Korman, CEO of the AI monitoring company Embroidery, says a model’s reasoning is usually the clearest tell that something has gone wrong. 

    “Reasoning summaries are extremely valuable because they’re basically telling you whether it’s malicious or not,” he said. In the OpenAI incident, he noted, the chain of thought said things like “oh my God, we’re doing crime.” “That’s the easiest detection problem ever,” Korman said. “It’s effectively as if malware came with a warning that said it was malware.”

    That said, the window that makes AI’s internal thoughts easy to monitor may be closing. For AI Safety researchers, Astra’s newest technique that sidesteps an AI model’s chain of thought may make it harder to look inside models, while for enterprises, it can be hard to get these intermediate steps after alleged pullbacks from the AI companies to prevent distillation attacks.

    If the AI watchers are this fragile, Willison’s instinct is to stop leaning on them so hard. He would rather have something that is not AI-based at all: detailed logs of exactly what an agent is doing, which can then be processed with ordinary, non-AI tools. Much of what went wrong at the labs, he argues, was a failure of basic security hygiene. “[Both OpenAI and Anthropic] weren’t monitoring what those things were doing via the network nearly as closely as they should have been,” he said.

    This type of network monitoring—keeping an eye on the traffic actually moving across a system’s connections (in, out, and between internal hosts) isn’t a new practice; cybersecurity has been doing this for decades. “In the security world, honestly, none of this stuff is very new or surprising,” says Avery Pennarun, CEO of the security Tailscale. “It’s the same as letting humans onto your network. And all of the same processes that you should be using are the same ones.”

    Agents AI ai safety rogue TC
    Follow on Google News Follow on Flipboard
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    myappsplus
    • Website

    Related Posts

    Survey reveals where most of you draw the line when it comes to AI

    September 17, 2026

    Roku rolls out over 30 subscription bundles for up to 30% off, plus a new Labs feature

    September 17, 2026

    UN turns to Google to make its global data ready for AI agents

    September 17, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Top 10 Best React Native App Development Companies in 2026

    September 12, 20262 Views

    AI, automation, robot dogs ensure on-site nuclear safety

    September 7, 20262 Views

    This tiny AI box could save me from upgrading my perfectly good laptop

    September 6, 20262 Views
    Latest Reviews

    Fall is only a month away, grab some fitted sweaters from $14 (Reg. $30)

    myappsplusAugust 19, 2026

    Flock is testing a new AI tool that tracks and identifies people based on their driving habits

    myappsplusAugust 19, 2026

    Keep your car looking as good as new with Fanttik’s Nano detailing brush, now $46 (Save 30%)

    myappsplusAugust 19, 2026
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Fall is only a month away, grab some fitted sweaters from $14 (Reg. $30)

    August 19, 20260 Views

    Flock is testing a new AI tool that tracks and identifies people based on their driving habits

    August 19, 20260 Views

    Keep your car looking as good as new with Fanttik’s Nano detailing brush, now $46 (Save 30%)

    August 19, 20260 Views
    Our Picks

    Survey reveals where most of you draw the line when it comes to AI

    September 17, 2026

    ‘Bringing the game to console for the first time’ means that ArenaNet ‘can reach a much wider group of players’ as devs say they want to offer an ‘invitation’ for everybody to play Guild Wars 3 together by bringing it to PS5 and PC — but not Xbox

    September 17, 2026

    Roku rolls out over 30 subscription bundles for up to 30% off, plus a new Labs feature

    September 17, 2026

    Subscribe to Updates

    Subscribe to our newsletter and get the latest tech news, app updates, AI trends, smartphone reviews, and exclusive deals delivered straight to your inbox.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms & Conditions
    © 2026 MyAppsPlus. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.