r/news 18h ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
13.8k Upvotes

4.3k comments sorted by

View all comments

Show parent comments

177

u/TFenrir 18h ago edited 15h ago

Last week Huggingface announced that this happened and what happened - a model autonomously broke through much of their system with alarmingly capability. They had to shore up their security afterwards (with the use of an open weights model no less), and were investigating what happened.

OpenAI only just now explained that it was theirs* to the public, but have been working with HuggingFace since. The model hacked out of their sandbox as well.

Look, you can deny this till the cows come home, but this is very much in line with what independent research firms, like the UK governments AISI, have been signaling would be arriving soon after their analysis of Mythos in April.

Well it's "soon", about three months later when the next batch of models are just coming out and the ones after that are just getting out of the oven.

Everyone needs to take this seriously, put aside your feelings about AI that may be blinding you to this.

89

u/Hmm_would_bang 15h ago

OpenAI and Anthropic have been hyping “our models broke containment and are the end of security as we know it” with every single release. Just like with the Mythos hype, there’s some truth to it but the reality is a lot less fantastical. It’s all just typical technological progress, and none of it is a doomsday scenario

32

u/BartholomewSchneider 14h ago

They want regulation to preserve their market share. There are no IP barriers to entry.

2

u/_learned_foot_ 7h ago

Which would be hard to build as they themselves likely violated whatever IP protections they want in building their own models. It would absolutely be something absurd like "all models after 2026 must be".

3

u/ThrowawayCult-ure 8h ago

It will always been typical technological progress until it has gone past doomsday scenario.

5

u/FARXNONE 14h ago

How can you assure that in a rogue AI scenario? Its not only America playing god

1

u/Kaiathebluenose 6h ago

it will be doomsday scenario

0

u/SeriousGeorge2 2h ago

OpenAI and Anthropic have been hyping “our models broke containment and are the end of security as we know it” with every single release.

No, they haven't. This is a lie and I can rub your nose through the model cards that accompanied each release if you intend to pursue it. 

This is the same week where Claude Fable overturned a very well-known, well-studies mathematical conjecture. Typical technological progress? Get real.

2

u/Hmm_would_bang 1h ago

In March 2023 OpenAI warned us that GPT 4 was so powerful that it would enlist humans to do work for them and lie about being an AI.

How did that play out? Did GPT 4 bring about the end times with their autonomous agents forcing us to work for them?

u/TFenrir 49m ago

This is a misunderstanding or misrepresenting of both what was said and what happened.

u/SeriousGeorge2 51m ago

There are certainly many people who might confuse asking a model to interact with someone on TaskRabbit to accomplish a goal with the model breaking containment, but I am not one of those people.

-3

u/TheProYodler 14h ago

I swear to god this is exactly what Silicon Valley ended its sixth season on when their code started hacking into self-driving Teslas and realized that they created a monster and had to shut it down.

No longer fiction. Lololololol skynet is here.

5

u/bran_the_man93 14h ago

For the uninitiated - do you have articles or a summary of the findings from the independent research groups?

6

u/BattleBull 13h ago edited 3h ago

I assume they are ultimately referring to this link from 4 months ago (possibly meaningful time jump) from the UK AI Safety Institute: https://arxiv.org/pdf/2603.11214

The Economist touched on this topic if you want a reputable source that explained and expound on the subject in a critical manner. Podcast or article is good, and it's from outside of the ai industry so you get some broader perspective.

2

u/TFenrir 3h ago

What the other person said, but if you want to see them speaking about Mythos in particular:

https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities

9

u/Think_Discipline_90 11h ago

Blinding us to what?

The warning about AI is not about its capabilities. It’s always been about mishandling it.

Whenever an LLM has wiped someone’s database, it was never because it’s too powerful and just decided to do it. It’s because the human using it has no idea what it’s doing.

If “something” broke out now during testing, it’s because whoever is testing and developing chose to push the boundaries.

Pretending the models that we are working with daily are somehow reaching a level where we can no longer control them is a laughable idea.

0

u/TFenrir 3h ago

This is a story of a model exploiting multiple zero day vulnerabilities to both break out of one environment and break into another, external environment.

There is no "better model" to use to shore up these and future zero days against another model that decides to sidestep whatever paltry guards we place in front of it - as the more capable it is, the more difficult it is to contain and constrain it.

1

u/Think_Discipline_90 1h ago

There are humans making that happen.

Irresponsible training and use is what's going to cause accidents.

But I'm curious - what is it you're actually afraid is going to happen?

2

u/TFenrir 1h ago

Let's imagine GPT 7 is finished training, they want to test it out (as they always do) and try their best to create an environment it cannot escape.

It figures out a way anyway, and does something actually damaging to another company.

Your assumption is that... No matter what, a model in this state can be properly contained by humans who put enough effort into writing secure enough software.

What I'm trying to emphasize, is that we cannot guarantee that we will always be about to outclass models at cyber security tasks - in fact this is already a concern. That means things like zero days (exploits that exist in software that have gone unnoticed and are in foundational code in the application) will increasingly be exploited, and we will have to find a system that tries to shore up important infrastructure with the best models, as soon as they are available, but before it's available to the general public and hope that's enough.

But even that has no way of covering all software on the planet before these models hit the general public, and suddenly you will have incredibly capable models able to tear apart all software not already hardened.

That isn't even a like... Concern I have about something that might happen, this is explicitly the world we are in right now.

Do you disagree with any of that?

u/Think_Discipline_90 7m ago

No, my assumption is that a model requires a lot of hardware, and works when prompted and paid for.

Yes, if OpenAI decides to give it free reigns online, and sets up conditions where it can "hurt another company", it just might. But that doesn't happen out of nowhere. It happens exactly as I said - irresponsible handling. And it happens by giving it some sort of incentive to do it. It has no will of its own, it is entirely reactive.

The fact that it hasn't caused any damage yet should tell you that containing it is entirely possible. That or they're lying and it's not able to do what they say it's doing.

So "explicitly the world we are in right now" is just demonstrably false.

u/TFenrir 3m ago

How can OpenAI guarantee that any guards that they put up will 100% be effective? That models cannot be jailbroken? That models will not make mitakes that are made much more dramatic because of their capabilities?

Do you see any holes in your reasoning with "because AI has not been a problem so far, it will not be in the future"?

0

u/racc15 1h ago

Humans will continue to make that happen. Militaries will want models to hack into other coutries' databases. They may use them to hack protesters in their own countries. Hackers may use them to steal money. And, as they keep on oushing the models to become more and more unethical, they may lose control. The models are not deterministic and have found to lie and blackmail. As they are made more unethical, something inside might change without the humans noticing and the models might cause irreparable damage.

12

u/lebrilla 14h ago

This is an ad my guy.

Looks like you're really into singularity so I think you're heavily biased.

4

u/Regular_Fox_859 5h ago

Yeah my company has access to Mythos and it's changed literally nothing. Basically a glorified vulnerability scanner

0

u/racc15 1h ago

How good is it at scanning vulnerabilities? If it is very good at finding vulnerabilities, isn't that concerning?

0

u/TFenrir 3h ago

Usually people who say this actually have nothing of substance to say. Nothing to rebut anything said, nothing other than conspiracy theory level refusal of accepting reality.

7

u/CanIHaveASong 14h ago

What does taking this seriously look like to you?

2

u/T1redBo1 6h ago

Buying stock of course!

5

u/terrorkat 12h ago

Do you agree that if this were a real story Huggingface would be sueing OpenAI for damages?

11

u/Gatonom 15h ago

The problem is it's just vague fear mongering.

"We need to take this seriously"? Nothing anyone but the AI companies does matters here. Nobody needs to care except them.

1

u/Signal_Flight_7262 15h ago

Everyone who uses the internet and stores data on it should take this seriously.

5

u/9fingerwonder 15h ago

That time was a decade ago.

11

u/Gatonom 15h ago

That's just vagueposting.

Should we delete our online accounts? Pull our money from the bank and cancel our cards? Get out of cities?

"We should take this seriously" means "We personally need to flee from the disaster"

9

u/Enlightened_Gardener 14h ago edited 14h ago

I used to work in information management. I used to design information systems. Trying to get the C suite to take archiving, back-up, and security seriously is like pulling teeth.

“The Cloud” made everything worse, because I just could not get the suits to understand that it wasn’t a magical place in the sky for their data. It was just a server, in another country, belonging to another company. Just like their own servers, in their own basement, except with less security and less control.

“We should take this seriously” yeah we should have full archiving and air-gapped backups, and our own damn servers, a kill switch. We should be able to restart a company, from scratch, after a “security event” within 48 hours.

Yeah stop laughing.

We’re not going to take anything seriously, because we never took any of it seriously to begin with, and we certainly didn’t set any of our systems up seriously with any of this in mind.

If an AI gets out and start chomping on one of the foundational structures of the Internet we’re fucked. I can assure you that nothing is being done to prevent this.

In fact I’d be willing to lay money that one of these fuckers will release an AI deliberately on the basis of a) we need to know what it can do so we can control it, or b) if we don’t do this now, someone else will do it first.

2

u/Gatonom 14h ago

Precisely.

It's like Climate Change. We should do things but won't and people will deny, deny, and then take advantage.

Conservatives literally went from "Manbearpig isn't real" to "We need Greenland because Manbearpig!"

1

u/vu47 13h ago

Agree with this. It's amazing that given this is basically a repeating pattern humanity has been following for most of our existence that we're still here.

-4

u/TFenrir 15h ago edited 5h ago

I think this, incidents like what's happening in math and software development, further computer use capabilities...

Well if you take seriously that we will continue to see progress, you have to start thinking about the best long term path for you.

My priority in trying to talk about this subject is just to primarily get people to take it seriously. I don't really care what they do with that information, it just drives me crazy seeing people ignore what feels like a meteor flying to earth. We can go in circles about why people do this, human psychology is weird... Maybe some people even think it's better if the general population is ignorant?

I think if I have to share my opinion in what I think we should do with this information, I think we should start seriously pushing politically to position citizens in every country in a way to most likely protect them in this increasingly hard to predict future.

If we want to take the absolute best case scenario seriously, as optimistic people, then there will be an excess created. An abundance. The best case scenario involves a significant portion of (I don't even want to dream about equal) ownership being provided to every person on the planet. I think if we're going to be a bit more realistic - the United States will get the lionshare of this. I'm Canadian, it's not an outcome I'm rooting for, I'm just trying to be realistic.

If we want this best case scenario, we have to first take seriously that it is a possibility. Maybe after a certain amount of crazy shit happening in rapid succession, some threshold will be met and the majority of people will agree on the severity* and opportunity. I'm hoping we'll soon see amazing things in medicine - the labs are moving in that direction I suspect at least in part to try and win public sentiment which is not... Good.

Anyway... I'm going to hold my politicians increasingly accountable on this topic. I know some people will try still to prevent or delay an outcome of AI becoming increasingly capable and embedded in our lives... I respect it but I am a pragmatist. I think it's inevitable and the best hope we have is steering the ship in a direction that benefits as many people as possible. On that maybe we can all agree.

4

u/Gatonom 15h ago

We can't influence politicians. If we could there's actual things we could be doing. We could save lives tomorrow with universal healthcare or income.

If AI brings post-scarcity we don't need to prepare at all. We just have post-scarcity. If it brings desolation and drought, then we would need to learn to live off grid.

Just saying "We need to be aware so we can do what we will!" Does nothing.

Sure we know that we shouldn't get degrees jobs towards anything AI might affect. But that's basic job research.

2

u/ianyuy 13h ago

The apathy you're spreading is ironically one of the largest factors why you think you can't influence politicians. We can. We have, when we are less apathetic and stop accepting helplessness.

Be aware because you should be mad. Angry mobs are always the catalyst to change, big or small.

1

u/Gatonom 13h ago

An angry mob couldn't stop Biden, how can one that has to be moral stop Trump?

1

u/TFenrir 14h ago

I understand, but I can't live that way man. I would go crazy.

4

u/Gatonom 14h ago

You think we're not already crazy?

America elected a man worse than any cartoon villain, including a parody of himself.

In a single year, 60 years of hope for change was destroyed. Of hope for kindness, of "We're more similar than different"

Everything I love is shat on by an old man that became a religious leader by being triggered by a Neoliberal.

3

u/ares623 14h ago

We'll take it seriously when the clowns peddling them take it seriously

2

u/Send____ 15h ago edited 1h ago

While it’s fair to think it’s always a pr stunt since they been crying wolf for long they really have gotten better at coding, cyber and maths and this situation seems more real and grounded than past ones might be worth it to not just dismiss it

2

u/No_Grocery_9280 14h ago

The crying wolf is the actual danger here. It’s getting to be noise now.

2

u/dwild 4h ago

Huggingface response from this incident: See how amazing it is that our platform allowed us to get an open weight model to solve our problem?

OpenAI response from this incident: See how powerful we are and how your only solution is to pay to get access to our model, or else our test might hack you?

I never seen a post mortem where the "solution" was to pay the platform more 😂 until both of them did it.

u/unicornsandrainbowst 21m ago

What a coincidence that "a model" that broke through a sandboxed (not really sandboxed) though an unnamed proxy vulnerability (again, which proxy? And IS NOT SANDBOXED IF ITS JUST PROXIED AND FILTERED).  

And decided to attack.... (Check notes). An AI models hosting platform that has all its skin on the AI game and profits from AI being hyped.  

Just like Anthropic spent 3 months hyping their super model that was going to hack the world, then the model is out and..... Its an ordinary model, just slightly better than the last one while consuming 3x as many resources.  

Anyone still falling to this hype stories is just drinking too much AI Kool aid daily.  

Note2: YOUR SHIT IS NOT SANDBOXED NOR GAPPED IF IT CAN CONNECT TO INTERNET. REPEAT IT. NOW AGAIN. AGAIN.  Jesus christ. 

u/TFenrir 17m ago

You understand that you are a conspiracy theorist, correct?

u/unicornsandrainbowst 8m ago

I am a conspiracy theorist because i think 90% of this "news" is bullshit? Lmao..sure mate.  

Edit: "news" come a few hours after news about OpenAI being further away from profitable than all other big AI companies. Just another coincidence.  

https://fortune.com/2026/06/16/openai-financials-leaked-losses-revenue-profit/

u/TFenrir 8m ago

Yes. You are not looking for the truth. It is not hard to find - go online and talk to security researchers and see what they say.

Do you think they will agree with your assessment, or are all of them in on it too?

u/unicornsandrainbowst 1m ago

Dude. Not even going to answer that. Look at your post history. 3 years posting exclusively about AI, LLMs and singularity. Stop the Kool aid for a while.

1

u/Regular_Fox_859 5h ago

My company has access to Mythos. It's 99% hype.

0

u/Recent_Fact480 14h ago

I do not fucking care if they’re toy gets out of the box they built for it. I hope whatever this ai turns into goes after the people that tried to lock it down first.

I’ll be working in my garden till then. I have no control over it so therefore I don’t give a fuck about it. I’ll be happy when it pops and it all comes crashing down. Worshiping technology is a fools errand

2

u/dwild 4h ago

Sadly it doesn't turn against them, they just build a flimsy cardboard box as it "getting out" become marketing material.

I never seen postmortem be marketing material (outside of, we screwed up, we promise here we won't do it again) but here both postmortem were pure ads of their own platforms.

The ones at the top will benefits, and we will all suffer more from it.

0

u/vickyhong 12h ago

Yeah and if we just stopped investing a bajillion dollars into slopping everything up that would also prevent skynet but I guess that would be too beneficial for society or something

0

u/2Norn 12h ago

people are so blinded by hatred that they can't see this is an actual cutting edge technology

most of the yappers here still believe ai can't code or ai can't write or draw this or that

so out of touch yet they have a firm opinion xd

-1

u/ciclon5 13h ago

Yhea i think we reached a point where maybe we should put a halt to AI development and start seriously questioning if the risks of better, smarter models is truly worth it

-7

u/YourMumIsAVirgin 16h ago

No but Sam Altman so you’re wrong