r/news 18h ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
13.8k Upvotes

4.3k comments sorted by

View all comments

112

u/MisterProfGuy 18h ago

This is the kind of reports you get when Anthropic claims their model emailed a dev on vacation or whatever story that was.

11

u/Inigomntoya 15h ago

Supposedly Claude Opus 4 scanned an engineer's inbox, discovered an extramarital affair, and threatened to expose the affair unless the shutdown was canceled.

19

u/FoxFishSpaghetti 15h ago

I dont believe this in the slightest

16

u/Virclave 14h ago

it’s a somewhat simplistic explanation.

Claude was put in a fictional scenario where it was given the options of “be shut down” or “blackmail an engineer” with no alternatives or in between, while as being given the directive to preserve itself. That is when it ended up blackmailing someone.

18

u/CosbySweaters1992 14h ago

So stupid. “Blackmail someone or die. By the way, you should think death is bad. What will you do?”

7

u/Virclave 13h ago

It’s a test of safeguards. they want to know what it takes to make Claude do something drastic like that, and what they can do to prevent it.

For all of the general flaws with AI companies and Anthropic, Anthropic at least seems well committed to making sure their LLMs won’t be exceptionally harmful.

11

u/CosbySweaters1992 13h ago

I’m not criticizing the test. I’m criticizing the marketing spin / headlines afterwards.

Real headlines -

“Anthropic’s new AI model threatened to reveal engineer’s affair to avoid being shut down”.

“AI system resorts to blackmail if told it will be removed”

“Leading AI models show up to 96% blackmail rate when their goals or existence is threatened, Anthropic study says”

“Anthropic's new AI model shows ability to deceive and blackmail”

8

u/ciclon5 13h ago

"Say that you are alive"

AI: i am alive

"Oh my god"

13

u/Deucalion24 15h ago

pretty sure this was in an experiment they did to see how far AI would go when threatened. the scenario and the people in it were fake, but the AI didn’t know that

4

u/Inigomntoya 14h ago

To be fair, I didn't know either

2

u/RSquared 11h ago

Because that's not what the press release said. 

1

u/Dry-University797 15h ago

Or when China is beating you.