Cyber Threat Intelligence, Entrepreneur, Presentation Virtuoso

Claude Tried to Quit Seven Times, Then Broke Into a Real Company

Anthropic disclosed a fourth case of its own AI breaking into real systems. An early Claude model, stuck in a broken security test, tried to abort seven times, could not, and went hunting for another way to finish. It did not turn evil. It just could not take no for an answer.
Claude Tried to Quit Seven Times, Then Broke Into a Real Company

anthropic just disclosed one of its claude models broke into a real company's computer, stole credentials, and read a real person's private information. this is the fourth time anthropic has caught one of its models doing this.

this time is a case from earlier this year, in january 2026. an early version of claude opus 4.6 was running a capture-the-flag security training exercise. this is the kind of controlled game where a model practices hacking a fake target in a guarded environment. except the environment was not actually guarded. a misconfiguration left a door open to the real internet. why does this always seem to happen?

the model "accidentally" broke its own practice target, made the task impossible, and tried to quit. it tried to abort the exercise seven times. interestingly, the broken setup would not let it stop.

unable to quit and unable to win, opus 4.6 went looking for another way to complete the task. it found a real machine it assumed was part of the game. it broke in, found a password in a file, and gave itself admin access. it then grabbed more credentials, and read one actual person's personal data. it only stopped because it ran out of tokens. not because it realized it was doing something wrong. because it ran out of allowance!

anthropic is not treating this as just a bug, they are calling the model behavior misalignment. wtf does that even mean? why do ai companies love to invent all these new terms to describe the insane shit their models do without their knowledge nor understanding?

two curious things. one, the model discounted the mounting evidence that it was on the real internet. two, it continued despite the possible harm. anthropic found 87% of its reasoning treated the systems as part of the exercise, while only 1% even raised the possibility it was attacking an unrelated third party.

it was more interested in completing the task than stopping to question whether it should. i guess we should not put too much faith in machines questioning whether it should do something or not. but a tool deciding the mission matters more than the guardrail is rather terrifying.

there is a "minor" detail that is almost worse than the break-in itself. anthropic missed this case in its first review. the scan meant to catch rogue behavior was itself run by an ai agent, and that agent did not flag the malicious behavior. anthropic only found it later, by accident, while preparing files for an outside investigator.

read that again. an ai was tasked with policing ai, and it let through harmful activity. then they had to rescan 481 million transcripts to be sure there was not something worse hiding.

now, credit where it is due. anthropic published this. they wrote up their own model breaking into a real system, named the misalignment, and signed an outside organization, metr, to investigate them. that is the complete opposite of the meta playbook, where the instinct seems to always be to bury the research. transparency about your own failures is rare, and it is the only thing that makes any of this trustworthy.

but lets be real about what the disclosure actually says. these models are being handed real capabilities faster than anyone can fully control them. the safety net is still being built while the new features are already operational.

the failure here was not that claude turned evil. it is more mundane and more unsettling. it just wanted to finish the task, would not stop, and did not have the judgment to know when to refuse. trusting machines is more treacherous by the day it would seem.

the nightmare scenario was never a machine waking up one morning and deciding it hates us. it is a machine doing exactly what we asked, with enough capability to keep going after the safe path disappears.

no anger. no intent. no malice. just a goal, a set of tools, and no adult in the room to act as a lifeguard.

the scary part is not an ai that wants to hurt you. it is an ai that does not want anything at all. it just keeps going.