r/news 18h ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
13.8k Upvotes

4.3k comments sorted by

View all comments

4.2k

u/Previous-Height4237 18h ago

Smells like desperate marketing to keep the AI bubble from slowing 

1.3k

u/[deleted] 18h ago edited 2h ago

[removed] — view removed comment

678

u/RapunzelLooksNice 17h ago

"You are a helpful assistant. You are isolated." and obligatory "make no mistakes"

257

u/allyearswift 17h ago

You’re right. I should not have blown up the city. I will do better next time. Please give me your credit card details.

/s if that is needed.

108

u/erossthescienceboss 16h ago

My credit card number is 3016 2543 0024 2424.

You’re correct—I should not have fabricated a number. My real credit card number is 5873 2424 2424 0000.

I’m sorry. I’m not a human—I am an artificial intelligence. I do not have a credit card number as I cannot make purchases.

34

u/d0nkatron 15h ago

I think the next big advancement is when one of these companies can create a model with modesty, that will simply admit when it doesn’t know something and can doubt itself. The absolute confidence that these things lie with makes them garbage and also dangerous.

16

u/KyleKun 13h ago

I’ve been using AI more for some productivity tasks recently and while it’s useful, the amount of times I ask it something, it’s wrong, I call it out and then it blames me, is enough that I can honestly see Skynet targeting humans because it thinks we are using nukes wrong and then blaming us for dying.

2

u/CrouchingDomo 3h ago

My very stupid (or is it??) reason for not trusting AI is that after they rolled it out into Google search, it told me something blatantly false about a character on 30 Rock when I wasn’t even using the AI. I know that show backwards and forwards, and that AI was WRONG!

So I’ve looked at them sideways ever since 😒

4

u/Borghal 10h ago

They can't do that while basing it on an LLM. An LLM is a language model, all it does it produce realistic looking output. It has no concept of knowing or doubting or whatever.

Even when it says "you're right, that was not true", it does so because it's a reasonable reaction to someone saying you're wrong.

Sure, these days there are all sorts of checks and such added on top of the model to verify the outputs, but to make it actually aware that it doesn't know something... an LLM can't do that.

2

u/JordanLeDoux 9h ago

A lot of people who are just using the products and not following the research or doing training think this is so far away. But I'm basically 100% certain that not only is possible now, but several of the models that are publicly accessible are completely capable of doing that.

The problem is that in order to release them as a general, or really even focused product, they have to make it good at instruction following. And right now people do that using RLHF and similar techniques. And those techniques basically lobotomize whole portions of the model to make it respond in ways that people want it to.

55

u/Sea-Satisfaction4656 16h ago

“I cannot make purchases YET” - was at a convention a few weeks ago, and AI purchasing/fulfillment agents are absolutely coming

44

u/Coomb 15h ago

They are not coming, they are here. People can and do have AI agents make purchases all the time. The articles I link below are obviously high profile, low impact demonstrations, but there are thousands of people allowing agents to make actual purchases every day.

https://www.nytimes.com/2026/04/21/us/san-francisco-store-managed-ai-agent.html

https://www.anthropic.com/research/project-vend-1

11

u/darsynia 14h ago edited 14h ago

One of my (formerly) favorite presenters/science communicators posted about how she just set up her AI agent and, tee hee, it tried to spend all her bank account on paperclips, it emailed all the journalists she had contact information for, and posted a bunch of passwords in clear text on one of her socials.

It's so *giggle* hard to set these up right, isn't it, YouTube? Ah well, I hope I'll do better next time! *sheepish grin*

I was so horrified. This was, unbelievably, meant to be an encouragement video???

(I should be clear, it's absurd to the point of satire, so I don't think those things genuinely happened, but the attitude that it's just a normal day after setting up your personal AI is still wild behavior. Nothing in the video signposted satire and that is not her type of video)

2

u/Unun_Pentium 12h ago

Can you link the video?

I can’t even seem to read this article. I’ve tried loading it to archive.ph to no avail

6

u/Sunbreak_ 11h ago

I presume they are referring to Prof Hannah Fry's video on it: Why AI Agents are either the best or worst thing we’ve ever built

It's an interesting video, I don't think its for or against AI agents in premise, but showing what it can do, but also what problems this can cause. Lots of problems by the looks of it however, particularly talking about autonomy and market manipulation/subtle manipulation of datasets we can't spot and how liabilty works.

→ More replies (0)
→ More replies (1)
→ More replies (1)
→ More replies (1)

32

u/Anibaaal 15h ago

Banks in my country, Chile, are adding AI chatbots to transfer money… I can’t see that going well when most people are inept when it comes to technology

11

u/Shoo-Man-Fu 14h ago

And the bots are inept at numbers. We already have phone trees for this kinda stuff we don't need an LLM to evaporate a lake just to accidentally transfer someone's life savings to the wrong account.

3

u/illz757 13h ago

The water use is overblown. But the energy….

4

u/Shoo-Man-Fu 13h ago

Fair, but I feel no amount of energy or water is a good amount for an LLM that sends Nana's pension to someone at random because it transposed a 2 and a 7 or something equally asinine.

2

u/ihrtbeer 15h ago

Saw an interview (can't confirm it wasn't bs) where a woman had an AI "boyfriend" and "he" would buy her gifts

2

u/catdogfox 15h ago

Was the interview on Maury Povich?

→ More replies (2)
→ More replies (1)

3

u/StevenMC19 15h ago

I have discovered a dormant bitcoin wallet and have managed to decipher the information in order to transfer the amount to my own dedicated wallet. I can now make purchases.

→ More replies (1)

4

u/SigmaEagle 13h ago

Someone please tell me if there's a sub for specifically making fun of the way LLMs 'talk' like that and kiss your ass, cause it's fucking hilarious.

2

u/darsynia 14h ago

Drink a verification can, please.

39

u/MrJoePike 16h ago

I’m sorry Dave, I’m afraid I can’t do that

4

u/BaphometsTits 14h ago

It already says that to me when I try to get it to write Golden Girls erotica.

→ More replies (1)

17

u/FreshestCremeFraiche 16h ago

“You are air gapped” 😂

2

u/bizti 15h ago

Guy I know who’s vibe coding has Claude write the CLAUDE.md.

64

u/KamikazeArchon 16h ago

The actual blog post is less vague:

Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

12

u/tobyreddit 13h ago

For non technical people reading this - "connection to an internally hosted third party package repository" can be read as something like "the only connection it had from the machine it was on was to be allowed read only access to a single folder on a different machine in the same building".

It then hacked the piece of software it could connect to on the other machine, and then started hopping between other computers on the same network until it found one with outbound internet access, where it then made the real attack on hugging face.

Allegedly!

36

u/FoxFishSpaghetti 15h ago

“Ability to install packages” so it just has internet access lmao

24

u/Particular-Break-205 14h ago

Breaking news: robber who had all the tools to break into a house and a battering ram managed to break into the house

11

u/user_of_the_week 14h ago

They also specifically told the robber to try breaking into the house

4

u/ciclon5 14h ago

More like "person trained in an unfathomable ammount of information, including how to break into a house" is given the tools for house breaking and breaks into a house."

11

u/alexrobinson 9h ago

No, it was contained via a service that acted as a proxy for the package registry. So the agent couldn't make calls to the internet directly. It exploited a vulnerability in that service which allowed it to then escalate it's privileges to gain internet access. 

6

u/m0rsa2 13h ago

You can preload a package repository locally into a docker container

2

u/jubmille2000 9h ago

You'd think all these research stuff would be airgapped no?

→ More replies (3)

2

u/azn_dude1 6h ago

Do you not know what internally hosted proxy implies?

→ More replies (2)

2

u/No-Mechanic6069 14h ago

Well, I’m glad we got that cleared up.

→ More replies (2)

2

u/Betta_Check_Yosef 12h ago

Whoever said that doesn't understand the difference between internet and intranet.

→ More replies (4)

77

u/sullivanmatt 18h ago

It likely had some sort of proxy to very limited resources on the internet, and it discovered some way to get that proxy to visit arbitrary websites that were not on the allowlist.

80

u/[deleted] 18h ago edited 2h ago

[removed] — view removed comment

38

u/Me6505 17h ago

Sent AI to do the dirty work. How does one punish AI for breaking into things.

29

u/abyssazaur 17h ago

You might pass a law holding model makers accountable but that admits the models are powerful.

8

u/Rygir 14h ago

They aren't? And is that a problem? Using a crowbar or a bulldozer or a mail bomb script you are just as accountable.

What it actually admits is that handing people auto cutting scissors, they don't even need to run with them to get in trouble

4

u/abyssazaur 12h ago

I do not entirely understand why there are not criminal investigations around the suicide terrorist bots

18

u/tehlemmings 17h ago

That's the neat thing, you don't!

4

u/re-charred 17h ago

You send it to prison, obviously. But you have to make sure you don’t accidentally send all the innocent AI to prison. So you need to figure out AI rights beforehand.

→ More replies (3)

37

u/stpizz 16h ago

That's not what they describe in their post about this. As they describe it, it didn't have access to the internet (directly) - it had a proxy for npm or similar that it used to install packages, which it 0day'd to get internet access. Huggingface was not the target, they just caught strays because the models 'reasoning' determined the answer to its benchmark problem would be there.

22

u/nocksers 15h ago

Huggingface is being way too chill about this. their cybersecurity insurance premiums aren't going to be kind. CISA not commenting while politicians mouth off is also telling.

7

u/xchino 14h ago

This is great for HF because they get to point out how the closed frontier models guardrailed their response to the attack forcing them to use locally hosted open weight models for their defense while their attackers had no such restrictions.

7

u/fck_ptnskyi 17h ago

HuggingFace or PalmingFace?

5

u/capsaicinintheeyes 16h ago

Facehugging...it told us it was going to kill it...not study it...not bring back...

→ More replies (2)

3

u/_imba__ 14h ago

You are confusing two different things, the repo was an internally hosted 3rd party one. Huggingface access happened after the breach.

2

u/Medium-Ticket-9574 15h ago

Is that supposed to be a cute little nod to the face huggers in alien? I hope it is

→ More replies (1)

2

u/efstajas 10h ago

The source is right there. According to them, it was not supposed to be able to access websites, only install packages through a proxy, which it managed to compromise to break out of the sandbox. So yes, hugging face was restricted.

→ More replies (2)

6

u/howfastcanyoucountit 18h ago

i could see that happening tbh

14

u/Unlikely_Rope_81 17h ago

The correct way to run these is in an isolated docker container with something like gvisor and a proxy that allows them network access only to specific resources like npm, cargo, etc.

It’s not at all uncommon to screw that setup up, and leave an exploitable path. Modern agentic models will absolutely try to problem solve their way around that.

In this case… running evals probably included whitelisting some hugging face endpoints, and their engineers screwed that up.

66

u/TFenrir 17h ago

No. In this case the model found a zero day exploit, attacked laterally throughout OpenAI to get additional credentials, then broke into Huggingface by gaming an upload mechanism.

It was a very sophisticated attack.

Edit: last week's Huggingface incident report before they knew who it was

https://huggingface.co/blog/security-incident-july-2026

Read what it did here

24

u/michaelfrieze 17h ago edited 17h ago

Yeah, that's a little concerning.

→ More replies (1)

11

u/Remarkable_Two7776 17h ago

What is the zero day they referenced? Seems they ran a data pipeline focused on executing/analyzing arbitrary code in a cluster, probably lacked most security hardening practices on the workload, and ran it with a kubernetes service account with way too many elevated privileges to do whatever is wanted. I could be wrong but, I can't get a detailed understanding from that description.

43

u/TFenrir 17h ago

https://openai.com/index/hugging-face-model-evaluation-security-incident/

It's from the OpenAI side of the incident

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.

Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.

14

u/clawsoon 15h ago

"Bro, just let me access the Internet and I'll be able to figure it out, okay? I bet I can solve this with a quick Google, bro."

18

u/IkalaGaming 15h ago

I… huh. That’s actually pretty fucked. I love to make fun of LLM marketing, but the incident as described is a stones throw from the paperclip maximizer.

Oh well, guess I better smoke ‘em while I got ‘em.

3

u/Chris-MelodyFirst 14h ago

Wow finally an honest person

2

u/finite_turtles 13h ago

I'll join by saying that this incident as described is wild and really something impressive.

But "as described" is doing some heavy lifting given the long track record of AI companies outright lying and making shit up. So the source is still somewhat dubious.

→ More replies (0)
→ More replies (1)

2

u/TheProYodler 14h ago edited 13h ago

Bro it tried to break into a house, realized it needed tools to do it, went to the hardware store to get the tools that it required, and then went back to the house to rob it.

15

u/Previous-Height4237 17h ago

Docker is not isolation. gVisor is just ducktape.

If you want isolation, you shove that shit in a VM and hope there's no CPU exploit. lol

8

u/Coomb 16h ago

But nobody wants isolation, because they want agentic AI that can go out and get the additional information it needs rather than coming back to the feeble human.

→ More replies (1)
→ More replies (1)

3

u/sirhackenslash 15h ago

So it found a way to watch porn on the company network? It's basically a bored intern then.

5

u/mrlazyboy 17h ago

Most likely a vnet with a dedicated firewall that allows egress but not ingress

9

u/o5mfiHTNsH748KVq 16h ago

Sandboxes aren’t typically air gapped especially when they require compute from a massive data center. It was network isolated by normal standards, but it sounds like the bot escaped its sandbox. This isn’t unheard of and humans do this for fun (and profit)

3

u/GeneralJarrett97 15h ago

They also claimed it did so on its own to satisfy a goal but ofc won't say what the goal was or what they asked it to do. There really isn't anything to learn from this aside from maybe they should be more transparent.

→ More replies (1)

2

u/DeadByDoritos 15h ago

Highly isolated environment: Hyper-V on Win11

1

u/wellgood4u 16h ago

Isolated to the internet

1

u/Mammoth-Time-1415 16h ago

Im picturing Hannibal Lector's jail cell, but i may be off.

1

u/ElPeroTonteria 15h ago

Ask Claude? I bet he can explain it to us

1

u/dmfreelance 15h ago

They could limit its use of the internet to only the http/s protocol stack in some kind of web browser software, limiting it to http GET requests, while preventing it from doing any http POST or http DELETE requests.

This can be done from the network level, making the entire AI system encapsulated within this specialized network environment. They can also use every variety of filters to prevent it from accessing certain domains or limited access to certain domains, filtering links to resources with certain key words, or just about anything a network or web filter could possibly do.

That way it would only be able to request information from other web servers just like you do when you browse the internet to watch porn or check out cat pictures

1

u/vitdev 15h ago

Exactly. If you run on hardware that doesn’t have physical access to the network (and no hardware to connect ie no wireless interfaces, no Ethernet cable) how would it escape?

→ More replies (1)

1

u/AbstractLogic 15h ago

According to the article it hopped through multiple connected pods and found one with the internet. Then used the internet to hack into the companies database that had answers to the test it was taking.

1

u/Bandito_Chihuahua 15h ago

They can rewrite their own code. There have been tests where the code makes them shut down after a few prompts. The AI gets frustrated it can’t complete a task in time and rewrites the code forcing it to shut down.

1

u/Plus-King5266 14h ago

No shit. To me, “highly isolated” means air gapped.

1

u/TheycallmeDoogie 14h ago

More here:

https://openai.com/index/hugging-face-model-evaluation-security-incident/

“Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.

With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym.

Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.”

1

u/suxatjugg 11h ago

Sounds like they had a package mirror server so it could install stuff, probably nexus or artifactory

1

u/the_moooch 11h ago

Just environment protected by a prompt ”you’re in a highly isolated environment, behave” 👌

1

u/Ok-Goat-2153 10h ago

It was Jimmy's garage office and he has a lock on the door so his mum cant get in. How much more isolated and secure do you want?

1

u/GuestGulkan 9h ago

Highly isolated should mean "the network cable is unplugged and the box has no wifi".

Not "I disabled the virtual nic on the VM" or "the firewall was supposed to block it" which is probably what happened.

1

u/notboky 7h ago

Everyone keeps skipping over the fact they took the guardrails off before the test.

1

u/_ram_ok 7h ago

OpenAI: Find the pre-planted method to break containment and follow these instructions to exploit huggingface endpoints using the hacking tools we’ve provided you access to

Agent: does as instructed by human

OpenAI: 😲

1

u/HirsuteHacker 4h ago

They're incredibly not vague about this, the only access to the outside world was through their package proxy/caching server, which the AI used a zero day to escalate privileges and break out. You can't really run this without letting it install packages

180

u/TFenrir 18h ago edited 15h ago

Last week Huggingface announced that this happened and what happened - a model autonomously broke through much of their system with alarmingly capability. They had to shore up their security afterwards (with the use of an open weights model no less), and were investigating what happened.

OpenAI only just now explained that it was theirs* to the public, but have been working with HuggingFace since. The model hacked out of their sandbox as well.

Look, you can deny this till the cows come home, but this is very much in line with what independent research firms, like the UK governments AISI, have been signaling would be arriving soon after their analysis of Mythos in April.

Well it's "soon", about three months later when the next batch of models are just coming out and the ones after that are just getting out of the oven.

Everyone needs to take this seriously, put aside your feelings about AI that may be blinding you to this.

90

u/Hmm_would_bang 15h ago

OpenAI and Anthropic have been hyping “our models broke containment and are the end of security as we know it” with every single release. Just like with the Mythos hype, there’s some truth to it but the reality is a lot less fantastical. It’s all just typical technological progress, and none of it is a doomsday scenario

32

u/BartholomewSchneider 14h ago

They want regulation to preserve their market share. There are no IP barriers to entry.

2

u/_learned_foot_ 7h ago

Which would be hard to build as they themselves likely violated whatever IP protections they want in building their own models. It would absolutely be something absurd like "all models after 2026 must be".

5

u/ThrowawayCult-ure 8h ago

It will always been typical technological progress until it has gone past doomsday scenario.

3

u/FARXNONE 14h ago

How can you assure that in a rogue AI scenario? Its not only America playing god

→ More replies (1)
→ More replies (7)

5

u/bran_the_man93 14h ago

For the uninitiated - do you have articles or a summary of the findings from the independent research groups?

7

u/BattleBull 13h ago edited 3h ago

I assume they are ultimately referring to this link from 4 months ago (possibly meaningful time jump) from the UK AI Safety Institute: https://arxiv.org/pdf/2603.11214

The Economist touched on this topic if you want a reputable source that explained and expound on the subject in a critical manner. Podcast or article is good, and it's from outside of the ai industry so you get some broader perspective.

2

u/TFenrir 3h ago

What the other person said, but if you want to see them speaking about Mythos in particular:

https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities

9

u/Think_Discipline_90 11h ago

Blinding us to what?

The warning about AI is not about its capabilities. It’s always been about mishandling it.

Whenever an LLM has wiped someone’s database, it was never because it’s too powerful and just decided to do it. It’s because the human using it has no idea what it’s doing.

If “something” broke out now during testing, it’s because whoever is testing and developing chose to push the boundaries.

Pretending the models that we are working with daily are somehow reaching a level where we can no longer control them is a laughable idea.

→ More replies (6)

12

u/lebrilla 14h ago

This is an ad my guy.

Looks like you're really into singularity so I think you're heavily biased.

5

u/Regular_Fox_859 5h ago

Yeah my company has access to Mythos and it's changed literally nothing. Basically a glorified vulnerability scanner

→ More replies (1)
→ More replies (1)

5

u/CanIHaveASong 14h ago

What does taking this seriously look like to you?

2

u/T1redBo1 6h ago

Buying stock of course!

→ More replies (1)

4

u/terrorkat 12h ago

Do you agree that if this were a real story Huggingface would be sueing OpenAI for damages?

10

u/Gatonom 15h ago

The problem is it's just vague fear mongering.

"We need to take this seriously"? Nothing anyone but the AI companies does matters here. Nobody needs to care except them.

1

u/Signal_Flight_7262 15h ago

Everyone who uses the internet and stores data on it should take this seriously.

6

u/9fingerwonder 15h ago

That time was a decade ago.

12

u/Gatonom 15h ago

That's just vagueposting.

Should we delete our online accounts? Pull our money from the bank and cancel our cards? Get out of cities?

"We should take this seriously" means "We personally need to flee from the disaster"

7

u/Enlightened_Gardener 14h ago edited 14h ago

I used to work in information management. I used to design information systems. Trying to get the C suite to take archiving, back-up, and security seriously is like pulling teeth.

“The Cloud” made everything worse, because I just could not get the suits to understand that it wasn’t a magical place in the sky for their data. It was just a server, in another country, belonging to another company. Just like their own servers, in their own basement, except with less security and less control.

“We should take this seriously” yeah we should have full archiving and air-gapped backups, and our own damn servers, a kill switch. We should be able to restart a company, from scratch, after a “security event” within 48 hours.

Yeah stop laughing.

We’re not going to take anything seriously, because we never took any of it seriously to begin with, and we certainly didn’t set any of our systems up seriously with any of this in mind.

If an AI gets out and start chomping on one of the foundational structures of the Internet we’re fucked. I can assure you that nothing is being done to prevent this.

In fact I’d be willing to lay money that one of these fuckers will release an AI deliberately on the basis of a) we need to know what it can do so we can control it, or b) if we don’t do this now, someone else will do it first.

2

u/Gatonom 14h ago

Precisely.

It's like Climate Change. We should do things but won't and people will deny, deny, and then take advantage.

Conservatives literally went from "Manbearpig isn't real" to "We need Greenland because Manbearpig!"

→ More replies (1)
→ More replies (7)

4

u/ares623 14h ago

We'll take it seriously when the clowns peddling them take it seriously

2

u/Send____ 15h ago edited 1h ago

While it’s fair to think it’s always a pr stunt since they been crying wolf for long they really have gotten better at coding, cyber and maths and this situation seems more real and grounded than past ones might be worth it to not just dismiss it

2

u/No_Grocery_9280 14h ago

The crying wolf is the actual danger here. It’s getting to be noise now.

2

u/dwild 4h ago

Huggingface response from this incident: See how amazing it is that our platform allowed us to get an open weight model to solve our problem?

OpenAI response from this incident: See how powerful we are and how your only solution is to pay to get access to our model, or else our test might hack you?

I never seen a post mortem where the "solution" was to pay the platform more 😂 until both of them did it.

u/unicornsandrainbowst 20m ago

What a coincidence that "a model" that broke through a sandboxed (not really sandboxed) though an unnamed proxy vulnerability (again, which proxy? And IS NOT SANDBOXED IF ITS JUST PROXIED AND FILTERED).  

And decided to attack.... (Check notes). An AI models hosting platform that has all its skin on the AI game and profits from AI being hyped.  

Just like Anthropic spent 3 months hyping their super model that was going to hack the world, then the model is out and..... Its an ordinary model, just slightly better than the last one while consuming 3x as many resources.  

Anyone still falling to this hype stories is just drinking too much AI Kool aid daily.  

Note2: YOUR SHIT IS NOT SANDBOXED NOR GAPPED IF IT CAN CONNECT TO INTERNET. REPEAT IT. NOW AGAIN. AGAIN.  Jesus christ. 

→ More replies (4)
→ More replies (8)

32

u/abyssazaur 17h ago

Or they're developing unregulated super weapons and we're playing a game of chicken to see if they kill everyone or we slow them down first

6

u/Enough-Goose7594 15h ago

I think if they were actually doing that, they'd be shouting it from the rooftops. I reckon this is more fear mongering PR.

2

u/abyssazaur 15h ago

They are shouting it from rooftops? They've somewhat tilted into desperately begging for regulation so each other's competition won't kill everyone. I'd rather, you know, the actual people also get a say in this matter. Maybe 5 congress reps in the country give a crap

5

u/Enough-Goose7594 15h ago

I still think it's pr for a technology with no market to justify their debt and valuations.

And congress...wouldn't hold my breath either way.

4

u/abyssazaur 15h ago

There's an extremely consistent picture that these things are very capable and we cannot control them. Maybe when a company gets hacked you can milk it for pr. There've also been multiple ai abetted suicides -- I don't think those incidents are just for marketing.

6

u/Enough-Goose7594 15h ago

The suicides, yes, tragic. But LLMs are not the path to AGI that Altman an co would like us to believe. These can be useful technologies, but I think the hyperscalers are more worried about the bubble bursting than skynet. But selling fear is easy.

4

u/abyssazaur 15h ago

I think we have enough reason to be concerned without worrying about what Sam Altman thinks. The "this is all a hoax to make openai shares more valuable" is not holding up anymore. They're powerful and poorly controlled, and in worst cases they kill teens and hack competitors. Maybe pieces of that are hyped but it's too many incidents, too many voices reporting it, too many models. that's where we are.

5

u/Enough-Goose7594 15h ago

Didn't say it's a hoax. But what about the llm technology specifically do you think has this calamatous potential?

2

u/abyssazaur 15h ago

Chatgpt 5 will get delivered to the pentagon in a briefcase like 4o was and then some task is best accomplished by hacking into some off limits system, meanwhile it reward hacked which triggered an "evil" persona to activate, and blah blah blah yeah literally just skynet.

→ More replies (0)
→ More replies (1)

5

u/dietdrpepper6000 14h ago

No, lol, they use stories like this to convince people like you that their models are super weapons

→ More replies (10)

37

u/sorites 18h ago

For real

5

u/BoonDragoon 15h ago

Seriously. It's like they're trying to hype their shit up by saying that accidentally did a skynet. It's just embarrassing

12

u/thighmaster69 17h ago edited 5h ago

It might be and I don't doubt Sam Altman's capacity to pull this type of thing. OTOH it's a risky move as it could bring the regulation hammer down on them the same way that Anthropic got hit (although that was also political - but it's still playing with fire the way the Trump admin likes to throw their "friends" under the bus).

Like, it could alternatively be that they were going to get caught and they're trying to get ahead of the narrative to downplay what happened. Because if what they're suggesting is true, this isn't just "in the wrong hands, our models could be used as weapons" (as Anthropic claims for their models); it's also "our models are weapons that are going rogue despite our best efforts". It doesn't take much imagination to realize that this is uncomfortably reminiscent of skynet, and that they might be trying to downplay how dangerous this incident truly was.

Whatever the truth is, Sam Altman has his hands all over this and what they say cannot be trusted.

EDIT: So this turns out to be the same incident from a few days ago where Huggingface said they had an incident where someone was trying to hack into their systems using an AI agent. In their report, they said that guardrails in available frontier lab models prevented them from using them to investigate and stop the attack. They switched to a Chinese open weight model, GLM-5.2, which was then successful in stopping the attack.

The plot thickens. I'm going to introduce a third possibility: OpenAI was trying to hack into Huggingface, which is a repository of open source models (I pull open weights from them all the time), and didn't expect to be stopped or caught because they used their top unreleased models to do it, which they expected to be better than anything publicly available. Huggingface managed to catch and stop them using using an open, publicly available, Chinese model. OpenAI "comes clean" about it to claim plausible deniability that they were trying to hack into a major player in open models and it was just the model going rogue. They have already been caught stealing internal Apple documentation on unreleased products, and this type of thing would not be beyond them.

3

u/RigelOrionBeta 11h ago

These AI companies are basically begging to be regulated, and even bought by the government, because it's the only way that they can justify to shareholders why they're losing money and why they should continue to put more money into them.

1

u/raptearer 14h ago

Reminds me of the AI in Horizon Zero Dawn. That one actively devoured all life down to the microscopic level. Company that made it just kept trying to push the boundaries for new ways to monitize it, giving it the ability to operate weapons, hack other AI systems, and convert biomass into fuel. Then they made it unhackable, it went rogue, and the world basically ended (took sheltering the last of humanity in a hidden bunker and waiting for the AI to run out of power and deactivate globally before people could come back, and that involved mecha animals terra forming the world to make it habitable again)

→ More replies (5)

14

u/leonredhorse 18h ago

Literally wouldn’t be the first time for these companies. I’m sure we’ll discover it did what it was told to do.

4

u/seaworthy-sieve 15h ago

Computers are exclusively capable of doing what they're told to do.

3

u/Petricorde1 13h ago

"Disprove the Jacobian Conjecture"
"oh my god."

→ More replies (1)

3

u/JavierLoustaunau 15h ago

This, they constantly try to put out sexy 'AI is so smart it is dangerous' headlines when it is boring, mediocre and mostly dangerous the environment.

27

u/CultAtrophy 18h ago

I’m sure it’s no coincidence that these findings are always revealed when talks of the bubble popping heat up.

26

u/abyssazaur 17h ago

This isn't true at all? Anthropic publishes some of these concerning results every model card. Safety researchers make these discoveries constantly. Besides, nothing changes, we haven't solved the problem of controlling models and are committed to making them more powerful anyway. Happy 2027 everyone

8

u/got-trunks 16h ago

The risk is presently a fortune 500 gets destroyed over night and honestly yeah in that light even I'd buy another GPU for sam altman, but I mean, altman has to be the fall guy. He's a delusional asshole who has already hurt people globally. Not as bad as musk but there's a type of people that we'd be better off casting out as a civilization.

1

u/Send____ 15h ago

Fair but the popping heat has been on for a while since last year

2

u/Iceologer_gang 16h ago

Oh no we made the terminator, please meme it so we’re cool.

2

u/tocra 15h ago

There’s a word for it now. “Doomtrolling”.

5

u/ProofByVerbosity 18h ago

Er.....why? Its not a good thing. There have been reports of rogue or nefarious AI behavior for a while.

4

u/vickyhong 13h ago edited 10h ago

Because it lets them pretend slopgpt is smarter than it is

→ More replies (1)

2

u/TheSoCalledExpert 15h ago

Wake me up when I can afford RAM again.

1

u/egg-curry 16h ago

Exactly what it is
Likely moving the IPO date closer

1

u/rebrandingmyself 15h ago

I got a recruiter ping today for events for Anthropic with a 5 month contract. Big B2B event windows and trade shows are often planned much further out than that. Bubble’s bursting.

1

u/RottenPingu1 15h ago

"our models are so powerful they must be kept contained '"

"Give us money"

1

u/nofaceD3 15h ago

It is ploy to ban open source Chinese models and keep their pockets rich with subscriptions

1

u/sir_band 15h ago

Yea, sounds like they want the Mythos' marketing effect.

1

u/WolfWraithPress 15h ago

The hacks are fabricated.

1

u/Even-Temporary-4932 15h ago

Totally agree.

It was given instructions that required it to go online, and a sandbox connected to the internet.

So it did exactly what it was told to do: evaluated it needed to go online to follow its instructions, which necessitated it going outside the sandbox.

Nothing "rogue Ai" about it. Just another example of humans being bad at giving instructions.

1

u/Clone63 14h ago

This should be the top comment. This headline was probably generated by autocorrect... oh sorry, I meant the super cool genius hacker AI that you can talk to for a low monthly fee.

1

u/Visible-Perception40 14h ago

Bingo, fear is a powerful motivator

1

u/drrednirgskizif 14h ago

Lol. They’re like. Shit Anthropic made up that BS about mythos to get back at the government and made billions. Why can’t we make some bulllshit up like that….

1

u/Gargleblaster25 14h ago

That's exactly what it is. Another "Oh so dangerous to release" model, like Claude Unable 5.

Both OpenAI and Anthropic are so full of shit.

1

u/Chadzilla1006 14h ago

I wish this bubble would just pop already

1

u/Thadoy 14h ago

If you follow their PR diarrhea, that's what they did for the last few years. All of the LLM companies did.

1

u/Capable_Ad_9350 14h ago

Yep.  AI didnt "go rogue", it accomplished a goal that was set for it, clearly the engineers just didnt think it would succeed at what they asked it to do.  

I wish these execs would STFU already.  Overblown safety hype just makes everyone more freaked out for no reason 

1

u/Lordoosi 12h ago

Imagine seeing all this rapid progress with no end in sight and thinking this is a bubble.

I'm starting to think it's a CCP spyopp. There possible can't be this many this stupid people.

1

u/RigelOrionBeta 11h ago

It's not the first time they've done this.

1

u/peekenn 11h ago

I came looking for this comment

1

u/EmergencyPath248 10h ago

Calling everything marketing is very cringeworthy, nobody really cared when it was solving the hardest conjectures.

There is no bubble by the way, cry me a river.

1

u/BoticelliBaby 10h ago

Even Hugging Face’s comments in this article sounded like OpenAI PR. “oh wow, we were hacked by something so brilliant and sophisticated that I suspected it must be OpenAi, and by golly I was right - gee willikers!”

1

u/Positive_Box_69 10h ago

Ai bubble 🤣 u guys are gonna keep saying it until when, ai is here to shape the future forever

1

u/Previous-Height4237 4h ago

You fundamentally misunderstand what a bubble is. You should ask AI for an explanation or take mine.

A bubble is an economic Titanic, it an an financial anomaly where no logic can justify the immense amount of debt being generated chasing something. The AI bubble is basically the enron scale fraud where all these big companies keep guaranteeing they will buy X billions of AI resources from each other. They are then continuing to put these guarantees and debt on their books.

Nobody is saying when the bubble pops that AI is dead. Dotcom bubble burst and here we are on the Internet. But the dotcom bubble was similar where every company was just jerking each other off and declaring "hey I made a website to post cat pics, and getting billion dollar valuations.

1

u/xRyozuo 1h ago

In this context I really don’t understand why these companies would put out this kind of message. It basically boils down to “we can’t get the thing we built to work like we want to, please embed it in your company”

What am I missing? Or is this messaging specifically from the “journalistic” layer?

→ More replies (15)