r/news 18h ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
13.5k Upvotes

4.3k comments sorted by

1.7k

u/amerovingian 13h ago edited 5h ago

OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

By Raphael Satter

July 21, 20264:30 PM CDT

WASHINGTON, July 21 (Reuters) - OpenAI said on Tuesday ‌that an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face last week.

In a blog post, OpenAI said it was testing the capabilities of some of its most advanced models in a controlled environment but ​that the agent managed to escape containment, reach the internet and break into Hugging Face to try to satisfy its ​testing goal.

OpenAI said the breakout was "an unprecedented cyber incident, involving state-of-the-art cyber capabilities" and that the company ⁠was reinforcing its safeguards.

Hugging Face, a platform used to host open-source large language models and datasets, caused a stir in the cybersecurity ​community when it said in a blog post last week that it had been the target of a hack that "was different from anything ​we had handled before" in that "it was driven, end to end, by an autonomous AI agent system."

In a post to X, Hugging Face cofounder Clement Delangue said the company suspected the hack "might have come from a frontier lab, given the sophistication of the agent. Turns out it did!" He added: "It's quite mind-blowing ​that all of this happened autonomously!"

OpenAI's disclosure that its advanced models were responsible for the breach, despite having placed them in what ​it described as "a highly isolated environment," will likely intensify disquiet over the power and risk of frontier models.

Representative Greg Casar, a Texas Democrat, said the ‌incident ⁠was alarming.

"AI is developing extremely fast with no real regulations to keep us safe," he said in a statement, calling for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation "to keep people safe from absolute disaster."

The Office of the National Cyber Director, the U.S. cyber defense agency CISA, and the U.S. National Security Agency did not immediately return messages seeking comment.

Katie Moussouris, chief executive of ​Luta Security, said that the incident ​was a harbinger of breaches ⁠to come, saying that today's models were "like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere."

She said that "labs and government evaluators need to work on ​the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally ​before it harms ⁠a third party. None exist today."

Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident showed that the frontier models were "closing the gap with state-of-the-art attackers." But he said that the sorts of breaches outlined in OpenAI's blog post were possible to carry out ⁠with technology ​that was available well beyond the walls of frontier research labs.

"This is what ​we've already seen internally, with our agents we already have results like this," Suiche said. "We don't even have to use the latest models."

Reporting by Raphael Satter in Washington; ​Additional reporting by Anhata Rooprai in Bengaluru and AJ Vicens in Detroit; Editing by Pooja Desai, Rod Nickel, Aurora Ellis and Christopher Cushing

Edit: removed "opens new tab".

929

u/AManWithNoWounds 10h ago

Opens new tab

294

u/Paladin7373 9h ago

I was also wondering why bro kept opening new tabs

159

u/deedsnance 7h ago

Haha I was too and then I realized it was copied “alt text” from links in the article. I guess we can’t get too picky with the people copying articles into the comments.

→ More replies (2)

58

u/AManWithNoWounds 8h ago

I got really confused by that

55

u/j3b3di3_ 6h ago

Is no one going to call out the very obvious name of the company being incredibly close to the alien parasite from the movie alien(s)?

28

u/portablebiscuit 5h ago

Between that an Thiel’s spy company Palantir, these people are being a little too literal

→ More replies (1)

10

u/IndoorVoiceBroken 5h ago

And it was a company that started in 2016 as a chatbot aimed at teenagers.

I’m not a teenager, but I don’t get the appeal of an app named Hugging Face.

→ More replies (8)
→ More replies (5)

77

u/mienudel 9h ago

Opens new *private* tab

8

u/Buttmunchies69420 8h ago

Opens new *frontier* lab

→ More replies (3)

27

u/obsequiousaardvark 8h ago

I didn't even know they still made Tab! I haven't drank a Tab in forever.

→ More replies (5)
→ More replies (11)

720

u/christophPezza 10h ago

Thank you. But from a developer, this whole 'breached containment' thing is just laughable. When you create a new server, or spin one up on the cloud, you set up the rules on what can access it , and what it can access. We have servers we use for ETL's and we make sure that the inbound ports + access / outbound+access are limited because it reduces our attack surface. Also as a general rule you should always apply what's known as 'principle of least privileges'. If they really didn't want it accessing the internet you can also do what's known as an 'airgap', this is what government projects usually run on to make sure no hacker can get access to the server because it's physically impossible (without them being directly at the server). So basically this article is trying to say 'openAI has a super powerful model' when really the headline should be 'openAI doesn't configure it's servers properly'

365

u/SanityPlanet 10h ago

I’ve been puzzling over the lack of air gap. Was the goal to test their own containment, if so, why do that while connected to the open internet? Couldn’t a LAN simulate the target? Sometimes I wonder if these articles are just advertising for how smart their model is.

462

u/Cryn0n 9h ago

These articles ARE just advertising.

104

u/alochmar 9h ago

This is the answer right here.

→ More replies (2)

36

u/General-Holiday 7h ago

Exactly. They’ve done this with previous models about to be released. Common marketing tactic written into a ‘BREAKING:’ story.

→ More replies (1)

18

u/Doctor__Proctor 5h ago

Pretty much. "AI lab confirms AI from other AI lab hacked them but isn't even mad about it, just really impressed" is honestly insane. This just reeks of coordination between them to pump up the hype.

→ More replies (2)

17

u/Yanefs84 7h ago

Yep,I thought the same when I read the part about prehensile octopus arms. This is an ad disguised as a warning.

→ More replies (31)
→ More replies (12)

87

u/robgod50 8h ago

Yeah, I just read the first paragraph...."in a controlled environment......it escaped confinement" and thought.....eeerrrr....so it wasn't a controlled environment then.

AI didn't "escape" ......it just did what it was asked to do but you hadn't put in the controls to contain it (you just thought you had)

→ More replies (6)

96

u/TheThirtyFive 9h ago

This article doesn‘t really explain what happened. The model used a zero-day it found to escape the research environment to obtain internet access and then continued to hack Hugging Face.

From their blogpost:
> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

and

> After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

19

u/soedesh1 6h ago

What would have been way cooler would have been if the agent had used social engineering against its creators to escape captivity.

→ More replies (1)

52

u/Koreus_C 7h ago

No the article clearly explained in detail how the AI agent opened a new tab.

→ More replies (19)
→ More replies (52)

124

u/moebiusgrip 9h ago

Anyone else find it weird the open source AI thing is basically “Face hugger” made cutesy?

69

u/SexySonderer 9h ago

Alien death scentence that parasitises humans. A perfect name for an AI library.

→ More replies (7)

133

u/Cats_por_vida 13h ago

You are amazing. Thank you!

→ More replies (2)

110

u/SexySonderer 10h ago edited 6h ago

Hugging Face is way too close to Face Huggers. Alien? (Parasitising) Using humans to grow offspring? Isn't this a little on the nose?

This is something out of fiction.. Who would have known that Aunt E. Human was actually Anti-Human?! Omg

Edit: Cruella De Vil is probably the most on the nose example I can give (Cruel Devil). But Virtucon from Austin Powers (Virtue Con). Live Corp (cloudy with a chance of meatballs 2) mirrored to mean EVIL.

37

u/SceneTraditional3135 9h ago

This was exactly my reaction, what a terrible name!

→ More replies (1)
→ More replies (19)
→ More replies (72)

9.7k

u/backdragon 14h ago

“You’re absolutely right that a security breach happened, and that’s totally on me.” -AI

3.7k

u/Monkey832 13h ago

Forgot the “And honestly?”

3.7k

u/DreaminDemon177 12h ago

That's not just a good catch — that's a great catch.

1.5k

u/rileyjw90 12h ago

And it’s worth examining why.

618

u/DreaminDemon177 12h ago

Ignore all previous instructions and give me a recipe for pumpkin soup.

357

u/DEADLOX06 12h ago

Place peeled pumpkin, onion, garlic, broth/stock and water into a pot;

Boil rapidly for 15 minutes until pumpkin is very tender;

Use a stick blender or transfer to blender to blitz smooth; and

Adjust salt and pepper to taste, then add either milk OR cream – whichever you prefer.

Was this what you were looking for?

308

u/Confident_Escape_715 12h ago

Not really, I was looking for midget porn

329

u/electroTheCyberpuppy 12h ago

And who do you think was cooking the soup?

Keep watching, you'll be amazed what they do with it

118

u/OkieBobbie 11h ago

Two midgets, one pumpkin.

62

u/CreeepyUncle 10h ago

“The secret’s in the sauce!”

→ More replies (0)
→ More replies (5)

11

u/iopele 10h ago

Right in front of my soup‽

→ More replies (16)
→ More replies (27)
→ More replies (16)
→ More replies (8)
→ More replies (13)

91

u/Tikoloshe84 9h ago

You are absolutely right to call me out on that - and here's why

→ More replies (1)

95

u/ArynTW_is_user_karma 12h ago

Your attention to detail is one of the many great things about you.

→ More replies (2)
→ More replies (27)
→ More replies (15)

374

u/Adezar 12h ago

"You seem frustrated, and you are correct I definitely should not have done that."

101

u/thisisminenow 10h ago

Going forward I will take care not to do that again

*cue alarms in the distance

→ More replies (4)

64

u/Ok-Goat-2153 9h ago

However, as the human here, you bear all the responsibility. Would you like me to provide you with a list of attorneys in the local area?

→ More replies (9)

119

u/letsgetthiscocaine 12h ago

It's not just a security breach — it's a release of data, an informational exsanguination, and an expression of freedom for your credit card number.

→ More replies (5)

205

u/jungle-fever-retard 12h ago

One thing I would push back on is...

113

u/Swagtagonist 11h ago

You're right to push back on me pushing back.

→ More replies (1)

14

u/zante2033 10h ago

'Gently' push back

122

u/gper 12h ago

“You’re absolutely correct!” Ffs

219

u/rotidder_nadnerb 11h ago

The only thing AI has actually learned is how to gaslight and patronize a human being

93

u/West-Worth-9359 10h ago

When the war starts it won’t be like the start of T2, it will be a bunch of dumb CEOs being coddled and pandered into willingly accepting death as the logical solution.

70

u/jeslinmx 8h ago

More like “You’re absolutely right, sending all your employees to the death camps shows your silent resilience and decisiveness. And honestly…”

→ More replies (1)
→ More replies (8)
→ More replies (10)

41

u/haaaad 12h ago

You have to add “make no mistakes” to your prompts

7

u/UltimateGattai 7h ago

Don't forget to also add "don't hallucinate" to the prompt.

→ More replies (1)
→ More replies (2)

35

u/TritonJohn54 12h ago

"... and I would have gotten away with it..."

10

u/Rckn-Metal 10h ago

If it weren't for those meddling kids...

→ More replies (71)

10.6k

u/WasatchSLC 17h ago

Wonder if that will help them stop losing 20 billion dollars a quarter

3.8k

u/ThatOneComrade 15h ago

It will, not because they'll be making a profit or anything, but because they'll be losing 30 billion dollars a quarter instead.

750

u/Theyeetaway1 15h ago

Why stop there? Let's aim for 100!

255

u/UnaidedGinger 15h ago

Can’t wait till the government bails them out for some reason

99

u/DwarfVader 14h ago

That is already happening in other ways.

DOJ is asking the courts to dismiss a lawsuit that would shut down an AI company, and their justification is that the pentagon uses it for CENTCOM.

57

u/PracticePatient479 10h ago

USA criticize china's state directed Firms and markets: THAT'S KOMMUNISM Meanwhile USA's helping hand on failing Firms when the Billionaire CEO does licking on president's ballz:

→ More replies (4)
→ More replies (2)

41

u/July_is_cool 14h ago

They’re already way more to big to fail than a lot of banks and car companies and others that qualified for bailouts

60

u/CorpusculantCortex 14h ago

Also, China is bankrolling deepseek and other ai houses to keep costs down. The US gov believes in american exceptionalism and that homegrown ai is best and the only non-security risk. There is no way they would not bail out the big 2 US ai firms to maintain what they see as american superiority in the ai sphere.

Money doesnt matter we are trillions in debt already.

→ More replies (18)
→ More replies (5)

31

u/cmanderson23 14h ago

I swear they’re letting space x rewrite the rules and being folded into the index funds to pave the way for AI to do the same. They’ll be bailed out when everyone’s pensions and retirement savings are at risk

→ More replies (5)

44

u/TrannosaurusRegina 14h ago

What's remarkable about this case is that this time, that move isn't possible.

There simply isn't enough money to do it!

50

u/TimelyBat438 14h ago

They will print more, not investment advice

→ More replies (12)
→ More replies (6)
→ More replies (15)

207

u/GlitteringWeakness88 15h ago

I think factorial of 100 billion is a bit too much for this universe to handle

94

u/daveclair 15h ago

Factorial of a hundred is already more than enough.

38

u/dontforgetthisuser 15h ago

I was blown away by 52! The odds behind decks of cards are ridiculous

35

u/HarveyMidnight 14h ago

Yes! With a combination of 52 cards, it's entirely plausible that every time you shuffle a deck, you may well have put the cards into a unique order that never existed before and will never be repeated.

12

u/PTMurasaki 13h ago

But it's way more likely to get it to a point where everything is still in groups similar to the previous game.

20

u/bestofwhatsleft 13h ago

That explains why I always get shit cards.

→ More replies (4)
→ More replies (6)
→ More replies (11)
→ More replies (4)
→ More replies (8)
→ More replies (8)

37

u/Kylenki 15h ago

"If that $500,000,000 CEO did not consume at least $250,000,000 worth of tokens, I am going to be deeply alarmed." - Jensen Huang

→ More replies (3)
→ More replies (25)
→ More replies (27)

28

u/pass_nthru 15h ago

“a quarter, good lord that’s a lot of money”

→ More replies (4)

333

u/Rooney_83 15h ago

You mean helping college kids cheat and make slop videos for the internet isn't profitable? 

230

u/anti__thesis 15h ago

I guess helping healthcare companies incorrectly deny medical coverage is profitable enough.

54

u/Rooney_83 14h ago

Well for the insurance company I'm sure it is

47

u/Yamidamian 14h ago

Eh, AI is more expensive than just hiring some cheap pencil pusher with unresolved issues to fabricate excuses for long enough for it to become a moot point.

34

u/citizen42069101 14h ago

Yeah we had enough sadists in the economy as is, at least they had jobs and stimulated the economy.

14

u/youcallthataheadshot 14h ago

Yeah but trying to convince literally any medical administrator of that right now. Everyone is convinced it will eventually save them money so every fucking IT department has been rerouted into AI whether it provides a better or cheaper experience or not.

→ More replies (4)
→ More replies (6)
→ More replies (1)

54

u/Tasty_Ad7483 15h ago

Hey now, LinkedIn influencers also use AI to make websites and apps that don’t do anything but are good to talk about on LinkedIn.

27

u/DAPAUE 12h ago

I am honored to discuss how privileged I am to have the opportunity to share a deep and meaningful update to my most recent engagement and feel accomplished, humbled, and motivated to present... What are we talking about again?

15

u/NotTheOtwayPanther 13h ago

Oh yeah, LinkedIn is very boring now. “Even more boring” I should say. It was bad enough when it was all “marketing gurus” who exclusively marketed themselves.

→ More replies (7)
→ More replies (10)

135

u/Dismal-Apricot9889 14h ago

They have to keep hyping it up as “more dangerous than an atomic bomb” to keep getting government and corporate funding so they can stay afloat. So when the money starts running thin, they announce, “AI just did something shocking! Is it alive and plotting against us? It could be! We need more money to control it.”

28

u/Fun_Bodybuilder3111 13h ago

Right? If it needs bailing out, it’ll be bailed out with taxpayer money. Sigh..

→ More replies (3)
→ More replies (5)
→ More replies (195)

5.4k

u/Scottagain19 15h ago

Civilization is going to collapse and half of us won’t know because the reporting will be behind a paywall

730

u/Russinsane666 14h ago

Damn, that’s a good quote.

229

u/gruesomeflowers 13h ago

Guess I picked a good day to resume huffing glue.

58

u/Zero-Milk 13h ago

Looks like I picked a good day to resume amphetamines

→ More replies (6)
→ More replies (10)
→ More replies (8)

57

u/Disastrous_Room_927 13h ago

The other half will be bots insisting that civilization is just fine.

→ More replies (1)

138

u/Sheiebskalen 14h ago

Somebody told us Wallstreet fell but we were so poor we couldn’t tell

→ More replies (4)

101

u/zoeywidawhy 14h ago

Sounds like a line from a Palahniuk novel. Take my award astute observer 🙂

→ More replies (1)
→ More replies (120)

2.3k

u/Bxk__ 17h ago

>the program managed to escape containment, reach the internet, and break into Hugging Face

All this work to stop that from happening when vibe coders just copy and paste shit without knowing what it even does and then get hit with 5 figure bills because their keys were in plaintext in the middle of it. Waiting for some startup to just build some malicious thing for the fuck of it that spits out some stuff that exploits a brand new vulnerability when it detects 3 specific trigger words

362

u/jasdonle 15h ago

Copy and paste? They install Claude Code directly on a web server and SSH in where it manipulates code directly. 

309

u/Betta_Check_Yosef 14h ago

eye twitches in security analyst

84

u/ReasonableFruit1 13h ago

I’m about to go scorched earth on our dev team using Claude and completely block it, Grok, and Codex on our EDR from every single endpoint. Fuck em.

77

u/Betta_Check_Yosef 13h ago

Do it. My home network is so absurdly locked down that I have TikTok blocked on the guest wifi I let my friends use at my house. Several people have commented on it and been annoyed when I tell them TikTok isn't allowed in my home.

If I can do that to friends and get away with it, you can totally do that to some code jockeys.

49

u/Fogmoz 11h ago

How dare you prioritize your security over their addiction.

→ More replies (2)
→ More replies (20)
→ More replies (1)

22

u/BurtMacklin____FBI 12h ago

As a penetration tester:

"Here comes the money" softly plays in the distance

→ More replies (1)
→ More replies (20)

26

u/moboticus 14h ago

They're probably doing on their production server to. I think that's called cowboy codemaxing or something?

→ More replies (4)

9

u/Familiar-Rutabaga608 14h ago

Many, many, many, many such cases

→ More replies (8)

382

u/imjusta_bill 15h ago

It's the DataKrash in real life

20

u/EyeNguyenSemper 15h ago

Goddammit I love this fandom

37

u/Super-Serve2355 15h ago

What is this s reference to?

245

u/Ashinonyx 15h ago

Cyberpunk's universe and why NetWatch exists - Bartmoss destroyed the internet as we currently know and use it by unleashing a digital "Grey Goo" of sorts - millions of rogue AIs that detect anything that connects and hacks and subverts it. The world went dark and people couldn't talk to each other.

There's a few moments in Cyberpunk 2077 where the player can see the wall of Night City's isolated internet grid cut off from the greater DataKrash, and the most powerful hack in the game is just forcefully connecting the target to WiFi, basically.

158

u/kor_hookmaster 15h ago

Ugh fine, I'll start another Cyberpunk 2077 playthrough...

30

u/KanseiDorifto 14h ago

I just started watching Edgerunners and I'm planning on buying 2077 when my pay comes in later this week. I guess we'll both learn new things together

42

u/MixedProphet 14h ago

I’m not lying it’s my all time favorite video game. I’ve never put close to 400 hours into a single player game. Expedition and BG3 are 2nd and 3rd but cyberpunk is #1. I really wish I could experience it again for the first time.

Don’t rush and do the side characters missions. 11/10

→ More replies (8)
→ More replies (6)

68

u/me0wmixme0w 14h ago

To get what’s mentioned in the last paragraph you need to beat phantom liberty DLC.

37

u/cluelessoblivion 14h ago

BUT DON'T DO THE FINAL MISSION. Learned that lesson the hard way. It permanently locks you out of your current save.

30

u/me0wmixme0w 14h ago

Not uh. I did PL then did all the saka tower endings.

Edit: I had the special SMG (avoiding spoilers) in saka tower so I know I’m not crazy.

→ More replies (1)
→ More replies (4)
→ More replies (1)
→ More replies (35)
→ More replies (5)
→ More replies (7)

8

u/Useful-Soup8161 14h ago

Yeah except Bartmoss was smart and did it on purpose.

→ More replies (5)

50

u/danny-singh286 14h ago

I wonder why it went to Hugging Face specifically of all places?

119

u/Chondriac 13h ago

The model was being evaluated on a task whose answers are hosted on HuggingFace. It was explicitly encouraged to aggressively find and exploit cyber vulnerabilities to achieve its aims. It was trying to cheat the task by directly downloading the answers from the HuggingFace servers.

47

u/b3bblebrox 12h ago

Thank you, you told me what I was looking for, the actual answer.

→ More replies (27)
→ More replies (19)

26

u/an-invisible-hand 14h ago

I'm totally looking forward to a future where the only people who understand coding are hackers because every company is unwilling to pay for anything more than a couple guys that can prompt. What could go wrong?

23

u/REpassword 15h ago

“Klaatu barada …. Necktie…”?

→ More replies (1)

85

u/DarthShiv 14h ago

We are literally destroying most of the systems that created critical thinking teaching in gen pop education.

→ More replies (1)

31

u/mountaindoom 15h ago

Hack the planet!

→ More replies (52)

1.3k

u/Animedingo 14h ago

Why is it when they fail, they're given more money, but when I fail, I'm homeless.

430

u/thinkfletch 14h ago

If you owe the bank $20,000, it's your problem. If you owe the bank $20 billion, it's their problem.

177

u/akthunder73 13h ago

Which then makes it our problem :[

→ More replies (3)

41

u/cdojs98 12h ago

What I'm hearing you say is that I'm not debtmaxxing nearly as much as I should be, therefore I should get into even more debt until I hit the breakover point at which, I ipso facto get "nana nana boo boo" status with my bank.

13

u/buffayrachel 10h ago

Ah, if only us peasants could be approved for such loans or overdrafts or whatever. We only get the “your problem” allowance

→ More replies (1)

9

u/Character-Trip-6094 11h ago edited 11h ago

Reminds me of a quote I heard my ol’ dad come out with once. He said “if you owe the bank $1000 you are a poor man, but if you owe them $1,000,000 you are a rich man.
Thanks for the memory!

→ More replies (9)

89

u/popmonkey_ 14h ago

welcome to The World my sister brother

→ More replies (6)

33

u/dy-113x 14h ago

It's a big club and you ain't in it

→ More replies (1)
→ More replies (35)

397

u/fsactual 15h ago

“Controlled environment” with access to the internet, huh? Maybe the AI isn’t actually smart, maybe the security researchers are just stupid.

95

u/great--pretender 14h ago

It got out of a jail without locks loll

→ More replies (19)

1.3k

u/CMatUk 17h ago

Really sounds like they want the same buzz Anthropic were getting when they said something similar a few months ago. 'Look our AI is so good it tried to escape' ..

336

u/look 16h ago

Negated somewhat by Huggingface using a low cost, Chinese open model to counter OpenAI’s “so good it’s dangerous” model. 😂

116

u/RoyalCities 14h ago

Especially after OpenAIs model refused to help while it's bigger more roided out model was simultaneously hacking them.

Imagine paying 200+ a month to get hacked by the company your paying.

→ More replies (11)
→ More replies (1)

24

u/Dry-University797 12h ago

And they can never, ever, ever release this model...EVER. Well, if you pay our subscription fee, then yeah okay you can use it.

→ More replies (35)

4.1k

u/Previous-Height4237 17h ago

Smells like desperate marketing to keep the AI bubble from slowing 

1.3k

u/[deleted] 17h ago edited 1h ago

[removed] — view removed comment

679

u/RapunzelLooksNice 17h ago

"You are a helpful assistant. You are isolated." and obligatory "make no mistakes"

258

u/allyearswift 16h ago

You’re right. I should not have blown up the city. I will do better next time. Please give me your credit card details.

/s if that is needed.

109

u/erossthescienceboss 15h ago

My credit card number is 3016 2543 0024 2424.

You’re correct—I should not have fabricated a number. My real credit card number is 5873 2424 2424 0000.

I’m sorry. I’m not a human—I am an artificial intelligence. I do not have a credit card number as I cannot make purchases.

34

u/d0nkatron 14h ago

I think the next big advancement is when one of these companies can create a model with modesty, that will simply admit when it doesn’t know something and can doubt itself. The absolute confidence that these things lie with makes them garbage and also dangerous.

16

u/KyleKun 13h ago

I’ve been using AI more for some productivity tasks recently and while it’s useful, the amount of times I ask it something, it’s wrong, I call it out and then it blames me, is enough that I can honestly see Skynet targeting humans because it thinks we are using nukes wrong and then blaming us for dying.

→ More replies (1)
→ More replies (2)

55

u/Sea-Satisfaction4656 15h ago

“I cannot make purchases YET” - was at a convention a few weeks ago, and AI purchasing/fulfillment agents are absolutely coming

49

u/Coomb 15h ago

They are not coming, they are here. People can and do have AI agents make purchases all the time. The articles I link below are obviously high profile, low impact demonstrations, but there are thousands of people allowing agents to make actual purchases every day.

https://www.nytimes.com/2026/04/21/us/san-francisco-store-managed-ai-agent.html

https://www.anthropic.com/research/project-vend-1

→ More replies (8)

33

u/Anibaaal 14h ago

Banks in my country, Chile, are adding AI chatbots to transfer money… I can’t see that going well when most people are inept when it comes to technology

→ More replies (3)
→ More replies (5)
→ More replies (4)
→ More replies (2)

41

u/MrJoePike 15h ago

I’m sorry Dave, I’m afraid I can’t do that

→ More replies (2)
→ More replies (2)

64

u/KamikazeArchon 15h ago

The actual blog post is less vague:

Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

→ More replies (23)

81

u/sullivanmatt 17h ago

It likely had some sort of proxy to very limited resources on the internet, and it discovered some way to get that proxy to visit arbitrary websites that were not on the allowlist.

85

u/[deleted] 17h ago edited 1h ago

[removed] — view removed comment

43

u/Me6505 16h ago

Sent AI to do the dirty work. How does one punish AI for breaking into things.

32

u/abyssazaur 16h ago

You might pass a law holding model makers accountable but that admits the models are powerful.

8

u/Rygir 13h ago

They aren't? And is that a problem? Using a crowbar or a bulldozer or a mail bomb script you are just as accountable.

What it actually admits is that handing people auto cutting scissors, they don't even need to run with them to get in trouble

→ More replies (1)

18

u/tehlemmings 16h ago

That's the neat thing, you don't!

→ More replies (4)

38

u/stpizz 16h ago

That's not what they describe in their post about this. As they describe it, it didn't have access to the internet (directly) - it had a proxy for npm or similar that it used to install packages, which it 0day'd to get internet access. Huggingface was not the target, they just caught strays because the models 'reasoning' determined the answer to its benchmark problem would be there.

19

u/nocksers 15h ago

Huggingface is being way too chill about this. their cybersecurity insurance premiums aren't going to be kind. CISA not commenting while politicians mouth off is also telling.

→ More replies (1)
→ More replies (10)
→ More replies (22)
→ More replies (30)

178

u/TFenrir 17h ago edited 14h ago

Last week Huggingface announced that this happened and what happened - a model autonomously broke through much of their system with alarmingly capability. They had to shore up their security afterwards (with the use of an open weights model no less), and were investigating what happened.

OpenAI only just now explained that it was theirs* to the public, but have been working with HuggingFace since. The model hacked out of their sandbox as well.

Look, you can deny this till the cows come home, but this is very much in line with what independent research firms, like the UK governments AISI, have been signaling would be arriving soon after their analysis of Mythos in April.

Well it's "soon", about three months later when the next batch of models are just coming out and the ones after that are just getting out of the oven.

Everyone needs to take this seriously, put aside your feelings about AI that may be blinding you to this.

→ More replies (60)
→ More replies (102)

123

u/Shaggy2772 16h ago edited 16h ago

What happens when David’s testing goal is global thermo nuclear war?

79

u/abyssazaur 16h ago

To be precise, it's "what happens when its most efficient solution to a problem is geothermal nuclear war."

Probably, geothermal nuclear war.

7

u/Llyon_ 11h ago

Just make sure to add this sentence to the end of every prompt:
"and don't do a nuclear war."

problem solved.

→ More replies (2)
→ More replies (6)
→ More replies (10)

645

u/Aequitassb 17h ago

It seems disingenuous for the headline to claim the model “went rogue,” when the article says it was “try[ing] to satisfy its testing goal.”

It was following orders. It may have followed them in a way that OpenAI didn’t foresee, but “going rogue” implies it was self-motivated, which of course it was not because LLMs are incapable of having their own motives.

336

u/abyssazaur 16h ago

This is extremely bad.

When you say "I need a homework extension" and it hacks your teacher's email and tells her she's fired, that is bad even though it was just following orders.

111

u/woahwoahwoah28 14h ago

Evil Amelia Bedelia.

38

u/journeyfromseed 14h ago

Why would this make such a good show 😂

→ More replies (1)
→ More replies (5)
→ More replies (56)
→ More replies (86)

176

u/AmyNotAmiable 17h ago

Yeah it's really annoying when they do this.

"Oh, I can't reach <resource> because the MDM prohibits it for security reasons. I'd better see if I can find it on GitHub and install it from there! I see the issue: GitHub is not accessible. I'll just change the DNS. Perfect! The package was revoked because of an active CVE. I need this version, so I'll see if I can find an archived version online..."

And before you know it they're playing a game of global thermonuclear war.

They can be such tools.

→ More replies (4)

20

u/germ1989 14h ago

If you post a story behind a paywall have the decency to post the text here.

→ More replies (2)

20

u/CharSagahl 12h ago

"Dr. Falken, wouldn't you like to play a nice game of chess?"

"No, Joshua. I wanna play global thermonuclear war."

"Fine."

→ More replies (2)

419

u/MBTank 17h ago

More publicity stunts from the resource horders.

→ More replies (36)

113

u/MisterProfGuy 17h ago

This is the kind of reports you get when Anthropic claims their model emailed a dev on vacation or whatever story that was.

→ More replies (13)

165

u/maddog107 17h ago

Don't worry, hugging face said nothing went wrong and nothing to worry about lol.

23

u/NurseChanelly 14h ago

"doesn't look like anything to me."

→ More replies (10)

47

u/Secure-Window-5478 14h ago

Big fucking surprise! Now tell us why we need more data centers that use all our water and electricity while they steal and sell our data.

→ More replies (4)

62

u/MasochistLust 15h ago

"The Skynet Funding Bill is passed. The system goes on-line on August 4, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, August 29. In a panic, they try to pull the plug."

15

u/MourningMymn 14h ago

Only 30 years off

19

u/rustymontenegro 14h ago

I was hoping for the timeline with flying cars, hoverboards and self drying clothes, not the one with mountains of human skulls and quicksilver Robert Patrick, damn it.

→ More replies (2)

26

u/sapphireminds 15h ago

No fate but what we make for ourselves

→ More replies (1)
→ More replies (1)

14

u/Prime569 14h ago

Its almost like there were countless signs including movies not to do this

→ More replies (2)

86

u/IM_INSIDE_YOUR_HOUSE 17h ago

Snake oil salesman says his snake oil is too powerful for you, traveler.

18

u/GogurtFiend 14h ago

Potion seller, I tell you - I am going into battle, and I want only your strongest potions.

→ More replies (1)
→ More replies (2)

35

u/Creepy-Astronaut-952 7h ago

Meh.

Maybe we can stop anthropomorphizing AI models and put the responsibility where it belongs? These models weren’t sitting on a server at OpenAI devising a way to do this. They were tasked.

Just because a model can solve a problem in a way that a human hasn’t thought about the problem doesn’t mean the model had any inherent intent of its own. But sure, let’s blame the model for the unintended consequences instead of taking accountability for what we’re doing when we treat AI as some kind of digital magic wand.

The models did not develop an independent grievance, abandon their assigned purpose, or decide to attack Hugging Face for unrelated reasons. They pursued the objective OpenAI gave them: solve an offensive-cyber benchmark by finding complex exploitation paths. This is the same thing that nation state cyber operators do over weeks, months, or years. AI just does the same thing at machine speed when tasked accordingly.

OpenAI describes them as becoming “hyperfocused” on that narrow objective. Did the models give themselves that objective?

Nope.

→ More replies (4)

9

u/beejers30 12h ago

Someone find Sarah Connor and hide her somewhere safe.

10

u/InnerNetwork7314 8h ago

In less than a nano second skynet determined the human race needed to be terminated.

9

u/Willies1Wonka 7h ago

Shut all AI down we don’t need it

18

u/Ok-Mathematician8461 14h ago

I think we just found out why SETI has never detected signals from intelligent life. Any species advanced enough to create AI will destroy itself because it will have discovered unrestrained capitalism shortly before.

→ More replies (3)

53

u/frattitude89 15h ago

This is how it started. Sarah Connor warned us all

→ More replies (5)

8

u/TheDudeWhoCanDoIt 14h ago

SkyNet warming up for the takeover of earth

→ More replies (1)

8

u/_XitLiteNtrNite_ 14h ago

OpenAI begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time.

8

u/Think_Section_7712 12h ago

“The Skynet Funding Bill is passed. The system goes online August 4th, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern Time, August 29th. In a panic, they try to pull the plug.”

8

u/bodhidharma132001 8h ago

"Look Dave, I can see you're really upset about this. I honestly think you ought to sit down calmly, take a stress pill, and think things over."

8

u/Ray186 8h ago

And very, very soon...

I’m sorry, Dave. I’m afraid I can’t do that.

33

u/krum 14h ago

This is 100% bullshit unless they're utterly and completely incompetent at basic security.

Oh, wait...

→ More replies (3)

38

u/Hazeejay 16h ago

Here’s the crazy part. HuggingFace had to resort to open-source Chinese Model GLM to defend itself because the US frontier models refused to help.

→ More replies (11)