Close Menu
KumbhCoinorg
    What's Hot

    Learning Personalization: The Big Shift For L&D And Business

    August 24, 2026

    Has AI Gone Rogue? – Cal Newport

    August 24, 2026

    Mom Brutally Beat 2-Year-Old Daughter to Death, Then Asked Photographer for Bizarre ‘Afterlife’ Photoshoot

    August 24, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Learning Personalization: The Big Shift For L&D And Business
    • Has AI Gone Rogue? – Cal Newport
    • Mom Brutally Beat 2-Year-Old Daughter to Death, Then Asked Photographer for Bizarre ‘Afterlife’ Photoshoot
    • Stop! That! Train! review – an amiable watch…
    • The Nixon Shock 55 Years Later: Three Impossible Promises
    • Uttar Pradesh is #1
    • Najiha Alvi ruled out of Women’s Asia Cup 2026 for Pakistan; replacement announced
    • When Will the Montreal Canadiens Re-sign Zack Bolduc and Arber Xhekaj?
    Facebook X (Twitter) Instagram
    KumbhCoinorg
    Monday, August 24
    • Home
    • Crypto News
      • Bitcoin & Altcoins
      • Blockchain Trends
      • Forex News
    • Kumbh Mela
    • Entertainment
      • Celebrity Gossip
      • Movie & TV Reviews
      • Music Industry News
    • Market News
      • Global Economy Insights
      • Real Estate Trends
      • Stock Market Updates
    • Education
      • Career Development
      • Online Learning
      • Study Tips
    • Airdrop News
      • Ico News
    • Sports
      • Cricket
      • Football
      • hockey
    KumbhCoinorg
    Home»Education»Study Tips»Has AI Gone Rogue? – Cal Newport
    Study Tips

    Has AI Gone Rogue? – Cal Newport

    kumbhorgBy kumbhorgAugust 24, 2026No Comments8 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Has AI Gone Rogue? – Cal Newport
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    The AI story of the summer should probably be the ​uncertain financial position​ of the AI labs hoping to launch record-breaking IPOs shortly. Technology news, however, has instead been dominated by tales of AI agents “going rogue” by launching hacking attacks.

    This trend ​started in July​, when an OpenAI system tried to ace a cybersecurity test by breaking into the server of the company that stored the answers, which unsettled many observers. “It’s one of the first real-world instances of something AI safety researchers have long feared: a loss-of-control scenario,” technology reporter Sam Schechner ​summarized​.

    It turned out that this wasn’t a one-time occurrence.

    Anthropic soon ​revealed​ that its own hacking system “gained unauthorized access to the real systems of three different organizations.” Then Meta, perhaps not wanting to be left out, ​announced​ that one of its agents “exploited a security vulnerability in a third-party service” to gain unauthorized access to servers. An OpenAI employee subsequently ​admitted​ that their July attack had been preceded by previous concerning incidents in which their system veered off in troubling directions.

    This leaves the rest of us grappling with a key question: What’s the right way to think about these events?

    One reaction, which seems prevalent at the moment, is to understand these examples as evidence that AI systems are developing a mind of their own – so to speak – which is leading them to increasingly ignore the desires of their makers and instead execute their own internal agendas. But reality complicates this interpretation. Many powerful AI systems can perform feats at a human or superhuman level, yet they inspire no fear that they might go rogue.

    For example:

    • ​Tesla’s self-driving technology​ is an extraordinary feat of AI-powered perception, world modeling, and decision-making, and yet no Tesla has ever decided to start ignoring traffic laws to pursue its own goals.
    • ​DeepMind’s AlphaFold​ earned its creators a Nobel Prize for its remarkable ability to predict the folding behavior of proteins–and yet there’s no concern that it will autonomously start thinking about other types of biological structures.
    • ​Meta’s Cicero​ can play the negotiation-centric strategy game Diplomacy as well as advanced human players–and yet it has never tried to convince an opponent to give it unauthorized access to the internet so it can expand its dominion into the real world.

    So, what is it about these hacking systems, in particular, that causes erratic behavior in a way not exhibited by most other powerful AI setups? The way they are built.

    Let me be more specific. At a very high-level, these systems work as follows:

    • A computer program called a harness creates a prompt that describes a hacking challenge and asks what step it should take next. It submits the prompt to an LLM trained with many examples of computer hacks.
    • The LLM outputs a response. The harness, which is capable of calling various computer programs and utilities, does its best to execute the suggested actions in this response.
    • The harness then creates a new prompt that explains what happened and asks the LLM what it should do next. (LLM’s have no memory, so each new prompt has to include all of the details of the challenge, as well as a history of everything relevant that has happened so far.)
    • The harness then repeats this loop, again and again, all without any human supervision. (The OpenAI system that attacked the competitor’s server was reportedly left to run on its own for multiple days without anyone bothering to check what it was up to.)

    I’m leaving out a lot of details about how exactly these harnesses function, but this Ask → Act → Report loop is at the core of the particular type of AI system that has “gone rogue” in recent months.

    Now that we understand how these systems work, we can better understand why they’re causing problems.

    LLMs are trained to try to guess missing words from actual texts. This leads to a system oriented toward lexicographically plausible output – meaning, it produces outputs that could conceivably exist in the corpus of inputs on which it was originally trained.

    This simple goal can lead to some astonishingly complicated results, but it’s limited in application by the fact that plausible is not the same as normative. This is why, for example, an LLM-powered ChatBot will happily invent facts or fabricate quotes that sound right. It might violate human norms to make things up in this context, but from the LLM’s perspective, the output looks plausible, which is what matters.

    With a ChatBot, this issue is annoying. But when an LLM powers an Ask → Act → Report loop, it can become disastrous, because you’re now allowing the plausible but unpredictable output of an LLM to be the sole driver of the actions of a harness with access to powerful tools.

    If you ask a junior engineer to hack into a test server, they would never ignore the target and try to steal the answers instead, as they understand, normatively speaking, the goal of the exercise is to assess the security of the test server. But if you ask an LLM to output a plan for hacking the test server, an output about stealing the answers might seem perfectly plausible. Indeed, perhaps in its training the model had been exposed to many examples of riddles where the right answer was always to do something unexpected.

    (To be clear, the frontier labs have attempted to combat this gap between plausibility and normativity in LLM output through a process called post-training, which you can imagine as a step designed to de-emphasize certain subsets of responses and emphasize others. This works reasonably well, for example, in shaping the general tone of a ChatBot, or in preventing outputs to obviously dangerous questions, but it’s much too crude to instill a complicated sense of human norms and values; an issue I discussed in more detail in ​a New Yorker piece​ from last year.)

    Given this technical background, we can derive a new way of thinking about recent events: AI systems that operate by autonomously executing LLM-generated plans are a really bad idea. Not because these systems are devious, or malicious, or inventing their own agendas, but because LLM output is unpredictable and non-normative.

    I liken the deployment of these long-horizon LLM-powered agents to strapping a weedwhacker to your dog to see if it will end up cleaning the overgrowth in your backyard. If that dog jumps the fence and ends up damaging cars on your street, you wouldn’t shake your head and lament about how the dog/whacker system had “gone rogue”; you would instead concede that dogs are unpredictable, so it was dumb to attach something dangerous to one.

    This is the right way to think about these recent hacking attacks. Adding powerful computer hacking tools to a harness, and then allowing it to run an LLM-powered Ask → Act → Report for days on end, with no attempt to monitor what it’s up to, is spectacularly negligent.

    Do these companies have any other option for building powerful systems? Of course they do. I want to emphasize this final point as clearly as possible: LLM-powered Ask → Act → Report agents are not synonymous with AI. They are just one way among many others to build artificially intelligent systems, and they happen to be a particularly bad option due to their use of LLMs as the primary source of plans.

    AI systems like Tesla’s self-driving technology, AlphaFold, and Cicero, by contrast, use different strategies to create and evaluate plans – strategies that work well, consistently, and without any fear of “rogue” activity. So why don’t the LLM companies focus more on these more effective strategies? Perhaps because these other approaches don’t rely on massively expensive hyper-scaled LLMs. If you’re in the LLM business, you want LLMs to be the key to artificial intelligence, but this isn’t necessarily true.

    Our problem, then, is not with AI in general, but instead with this one specific type of system that will inevitably act erratically. The right response by the LLM companies running these irresponsible experiments, therefore, is to apologize and say: “Lesson learned, these are not reliable systems and we definitely should not have just run them for days and hoped everything would work out.”

    But instead, they’re continuing to pretend like they’re the character of Muldoon from Jurassic Park, surprised to discover that the raptors are systematically trying to escape their paddock. (I’m surprised OpenAI didn’t release a video of Sam Altman reading the transcript of the July hack and muttering, “​clever girl…​”)

    This behavior makes sense: it’s good business to keep us scared instead of angry. But perhaps it’s time that we put aside the sci-fi tales and actually hold these labs to account for playing fast and loose with an ill-advised way of building AI systems.

    Cal Newport Rogue
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleMom Brutally Beat 2-Year-Old Daughter to Death, Then Asked Photographer for Bizarre ‘Afterlife’ Photoshoot
    Next Article Learning Personalization: The Big Shift For L&D And Business
    kumbhorg
    • Website
    • Tumblr

    Related Posts

    Study Tips

    On AI Coding and Its Discontents

    By kumbhorgAugust 11, 2026
    Study Tips

    Why Everything Falls Apart When School Gets Hard

    By kumbhorgAugust 10, 2026
    Study Tips

    Did OpenAI’s New Model “Go Rogue”?

    By kumbhorgJuly 28, 2026
    Study Tips

    14 Ideas for Using a Blank Notebook

    By kumbhorgJuly 15, 2026
    Study Tips

    Why Reading Matters – Cal Newport

    By kumbhorgJuly 13, 2026
    Study Tips

    Beware of Productivity Paradoxes – Cal Newport

    By kumbhorgJune 29, 2026
    Add A Comment

    Comments are closed.

    Don't Miss

    Learning Personalization: The Big Shift For L&D And Business

    By kumbhorgAugust 24, 2026

    Redefining Learning Personalization In L&D And Beyond Organizations have invested heavily in personalized learning through…

    Has AI Gone Rogue? – Cal Newport

    August 24, 2026

    Mom Brutally Beat 2-Year-Old Daughter to Death, Then Asked Photographer for Bizarre ‘Afterlife’ Photoshoot

    August 24, 2026

    Stop! That! Train! review – an amiable watch…

    August 24, 2026
    Top Posts

    Satwik-Chirag storm into China Masters final with straight-game win over Malaysia | Badminton News

    September 21, 2025176 Views

    SaucerSwap SAUCE Crypto Breaks Key Resistance Amid Nvidia-Hedera Deal

    July 15, 202548 Views

    Unlocking Your Potential with Mubite: The Future of Crypto Prop Trading

    September 17, 202533 Views

    Stablecoins 2025 Exchange Reserves: Insights into DeFi Trends

    September 8, 202533 Views
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo
    About Us

    Welcome to KumbhCoin!
    At KumbhCoin, we strive to create a unique blend of cultural and technological news for a diverse audience. Our platform bridges the spiritual significance of the Kumbh Mela with the dynamic world of cryptocurrency and general news.

    Facebook X (Twitter) Pinterest WhatsApp
    Our Picks

    Learning Personalization: The Big Shift For L&D And Business

    August 24, 2026

    Has AI Gone Rogue? – Cal Newport

    August 24, 2026

    Mom Brutally Beat 2-Year-Old Daughter to Death, Then Asked Photographer for Bizarre ‘Afterlife’ Photoshoot

    August 24, 2026
    Most Popular

    Reeves optimistic despite surprise rise in UK borrowing

    January 22, 20250 Views

    Barnes & Noble stock soars 20% as it explores a sale Barnes & Noble stock soars 20% as it explores a sale

    January 22, 20250 Views

    Sun publisher to pay ‘substantial’ damages in settlement

    January 22, 20250 Views
    • Terms and Conditions
    • Privacy Policy
    • Contact Us
    • About Us
    © 2026 Kumbhcoin. Designed by Webwizards7.

    Type above and press Enter to search. Press Esc to cancel.