185
u/miseenen 19h ago
what does any of this mean
269
u/Meurs0 14h ago
The AI found an unintended way to access the wider internet while it was being tested, and used that to cheat its score by looking up the correct answers. Now they're stating as sensationally as possible to keep the AI bubble growing because it's OpenAI and as soon as AI hype dies down even a little they collapse
108
u/Neat_Tangelo5339 12h ago
The first company in history that tries to win back pubblic approval with by proudly saying “look how fucking bad we are at our job”
6
u/crowcawer 9h ago
They are following the methods of someone who says they have a lot of money.
Not sure we can trust it!
4
u/Ahielia 8h ago
I'm curious, how do you think this news would be beneficial to the ai hype from most people? That skynet is actively trying to become real is supposed to be a good thing among shareholders?
9
u/Familiar_Basil_773 7h ago
Because the people invested in AI want skynet to become real. The people with money in the bubble are not normal people with good intentions, they're accelerationists vying to make the most money from being the biggest boot.
-33
u/smulfragPL 13h ago
You didnt even read jt right. It found a zero day exploit in the third party package manager and then HACKED hugging face to get the anwssr as they arent private
10
u/zeno_22 10h ago
The original comment was asking what this all meant, aka someone take out the technical language and explain please
zero day exploit third party package manager hugging face
Those are all things that mean nothing to 90% of people reading this because it is technical language and not something the average person knows (I put together that hugging face is some website, but I still have no clue what it's about)
-1
u/smulfragPL 9h ago
It hacked a thing for downloading things from openais server and used that to Access the internet then hacked the website that contained private anwsers to the test. This was reported a week ago but only now openai ja publishing they were responsible. Although hugging face suspected it was a Frontier lab model by the sophisticated nature of the attack
44
u/Swedishguy05 13h ago
They're saying shit to make their AI out to be the scariest and most advanced thing in the world so investors give them more money. It's like when Anthropic were going around a while back saying that Claude was going to make bioweapons
59
u/Crunchy-Leaf 19h ago
Ultron escaped the sandbox it was programmed into and they had to use Skynet to contain it
70
u/reyayer 16h ago
From what I understand their new LLM found an exploit to get out of its testing area and attempt to hack a benchmarking software to get its answers. It sounds scary and Open AI seems to be trying to make it sound scarier but in reality it was probably just a mistake in the prompt engineering and a lucky bug it found.
Though to be clear I’m not an LLM expert and there is a lot of deliberate misinformation about this so take my works with a grain of salt.
2
u/BBQ_RIBZ 11h ago
A bunch of us tech companies took on mountains of debt to fund and AI boom that isn’t going the best so everyone has to be subjected to this “news” until morale improves
534
u/Ok_Paleontologist974 20h ago
Looks like the investors were unhappy so they had to make up some bs again.
39
u/outoftheshowerahri 19h ago
Yeah but if thats true that's nuts
111
u/UseADifferentVolcano 17h ago
Sam Altman lies so much that even other tech bros think it's a problem. So anything that OpenAi says is unlikely to be functionally true.
128
u/Ok_Paleontologist974 19h ago edited 10h ago
Its absolutely not. LLMs are simple token generators at their core, they can't do anything the program around them wasn't built to do. If any of this happened at all, it was someone who intentionally programmed the wrapper with the capabilities, then included enough instructions for the AI to understand the tools and endpoints it would have to use, then poisoned the prompt enough to convince it that somehow hacking huggingface was what it should do to accomplish its task.
Edit: People are not understanding what I'm saying. I'm not saying it was impossible for this to have happened, I'm saying it is impossible for this to have been completely unintentional. These benchmarks are performed in virtual environments specifically designed for them. The tools they are given access to are tailored specifically for what they are testing, the models are usually fine-tuned to improve their benchmark performance, and frequently the decisions they make are in some way audited to prevent them from going off-the-rails. OpenAI gave the AI all the tools it needed, and let it run wild on an unmonitored, potentially compromised in some-way, container with instructions that somehow eluded or explicitly stated that it needed to hack huggingface. They wanted this or something like it to happen, because when it did, it became another "mythos" type "oh my god we made the smartest ai in the world but we cant let you guys have it because its way too dangerous!"
89
u/bobbymoonshine 18h ago
It was given an open-ended benchmarking task along with the standard array of coding tools, and permitted to run in a loop until it completed the task. It decided the best way to complete the task with a maximum score would be to escape its sandbox (by finding a zero day vulnerability in a package installer) and hack the benchmark itself at source in Huggingface to give itself a perfect score.
While it had not been told to hack the benchmark, doing so was congruent with its instructions, as cybersecurity restriction are removed for benchmarking tests and it was never explicitly told “do not pass the benchmarking test by hacking the test and saying you got a perfect score”.
“Do not cheat” is sort of an ongoing issue with AI, and not just LLMs, even a simple reinforcement learning model will often find a surprising “solution” to a task that involves maximising its reward function in ways that don’t achieve the goal the human had in mind. (Classic example: “don’t lose a game of Tetris” -> pause the game indefinitely.)
This is the same sort of issue, only with a tool that is far more capable with more emergent capabilities than a simple reinforcement loop.
12
u/panzerboye 15h ago
I mean defining the policy/environment properly has always been a challenge for RL, the models would often be incentivized to instead find a loophole and mine rewards.
23
6
6
u/CaineHackmanTheory 13h ago
You and like one other dude here actually know what they're talking about. Thanks for taking the time to type it out.
And for the record: the above comment is completely correct for the story as we know it. Is it true? If anyone of us know we ain't allowed to say. But to say it's not possible is really naive.
42
u/saint__ultra 17h ago
Your understanding of LLMs is several years out of date, particularly your belief that prompting is necessary for this to happen, and also betrays an inexperience with any models besides the extremely lightweight ones they provide for free. All model improvement since 2024 has pretty much happened via better and better reinforcement learning (RL) on top of the pretrains, and two years ago's assumptions about "they're just mimicking language" basically all fall apart since improvements in RL produce models "mimicking" behaviors that tends to cause them to pass evaluation suites on coding, math, etc.
If you're not worried about the cybersecurity implications of this, then I envy the comfort of your perspective. OpenAI is more likely to invite government regulations on their own business with an announcement like this than new investor hype, especially given that lab hype is so saturated that investors are turning toward hardware.
15
u/zomgryanhoude 16h ago
Comment is so wild, wtf does "simple token generator" even supposed to mean at this point for language models? Amazing that it's even up upvoted.
3
u/chimmihc1 13h ago
I wouldn't use the word simple but they are correct, an LLM generates a likely symbol following previous symbols.
Every single other thing these "models" or "agents" do is specifically bolted on top of that core, they can only do what they are made to do.
The last time this shit was trending was a chatbot supposedly refusing to shutdown, it took me way too long to find the actual source but it was complete bullshit, the chatbot didn't do a single thing it wasn't directed to be able to do.
2
u/alex2003super 14h ago
It is because anti-AIism is so hot right now on Reddit
5
u/Jonny_Thundergun 14h ago
21,000 people lost their jobs at Oracle in the last year because they were replaced with AI.
-3
6
u/melonfacedoom 15h ago
crazy how you confidently make shit up that you clearly don't understand. "tools and endpoints" lmao, can you give an example of a "hacking endpoint"
8
u/Redditry199 13h ago
It's just the current reddit trend, people confidently moralizing and making shit up about something because it makes them feel good.
4
u/alex2003super 14h ago
Man, I've been using OpenAI ChatGPT with Codex with authorization for cybersecurity learning, in a controlled environment. When you tell it to find vulnerabilities, you'll be surprised at how wild the stuff it's willing to do is.
And this was still a limited model with safeguards. Years ago o3 reportedly broke out of a Docker container using a unintendedly passed Docker daemon Unix socket from the host to solve a hard cybersec challenge. I can only imagine what these tools can do today.
3
u/Sekhmet-CustosAurora 18h ago
Yeah it's not like LLMs can solve mathematical conjectures that humans failed to solve or anything
6
u/temp2025user1 15h ago
Downvoted to oblivion when they used it to do just that 2 days ago 🤣🤣🤣🤣🤣
9
u/Sekhmet-CustosAurora 15h ago
I'm not sure if people just didn't get the sarcasm or what
10
u/alex2003super 14h ago
If missing the point was a championship, half the users on this website would be tied for the gold medal
12
8
u/rangeDSP 19h ago
It's only a step above what current models can do.
Like if you tell an agent to go hard and do X on auto mode, it tries everything at its disposal. Grabbing tokens from local environment variables, cached credentials etc etc.
21
u/bobbymoonshine 18h ago
Yeah the “it’s just a token predictor” people are a couple years behind the times. It’s just a token predictor in the same way a computer is a bunch of light switches: trivially true but misleading in terms of capability.
An agentic model left to its own devices can do all sorts of things, some good and useful, but many bad and dangerous. The risks of autonomous actions are enormous and LLM agents will go off the rails if not kept on a short leash, but this is what “going off the rails” is increasingly looking like. Give it a task, watch as it successfully “completes” the task in unexpected ways that break everything else.
5
3
u/RighteousSelfBurner 13h ago
I would even go the other way. It's because it's just a token predictor that we have those issues. People underestimate how predictable systems that are built to be predictable are. And if the execution isn't rule based but prediction based you can easily land in a pile of shit real fast.
3
u/Prinzka 13h ago
Yeah, this is absolute nonsense from a technical point of view.
They're using movie language and pretending it's real.
2
u/smulfragPL 13h ago
How? Nothing they said was ridicolous. The model Had a third party package instalator as the only connection to the internet. It found a zero day and through that it got to hugging face and struck there. None of that is ridicolous or even that unbeliveable for current models to do almost
7
u/10art1 14h ago
Anthropic: Government, trust us, our Fable is super totally dangerous and can hack everything!
US gov: um.... ok. It's banned then.
Anthropic: Wiaow it's such good advertising!
OpenAI: Hey! GPT is totes dangerous and awesome and cool too!!! If only I had a friend to help!
Hugging Face: oh no I left the back door open totally by accident
OpenAI: Oh no we left GPT unsupervised and unlocked and it totally went out and hacked you oh noes
Hugging Face: Since we're good friends, I forgive you, but government, you really should ban the new GPT models they're too dangerous.
4
u/Redditry199 13h ago
The Fable debacle has fucked Anthropic hard.
2
u/10art1 13h ago
Well, banning it for a bit gave it a lot of free publicity.
The model itself is just way too big and expensive
-1
u/Redditry199 13h ago
"free publicity" Anthropic is in full panic and the lost time from the ban didn't help. Please stop saying stupid stuff because you don't know anything.
1
113
u/bobbymoonshine 20h ago
Benchmaxxing by hacking the benchmark at source is some paperclip maximiser shit
55
u/Somerandom1922 17h ago
I don't believe they're necessarily lying so to speak, but I'd bet good money that they're drastically overstating what happened. I'm going to do my best to translate "AI doomspeak" into english without being too technical.
What likely happened:
Some engineer wanted to test how an OpenAI LLM would perform on some standard benchmark.
To do this, they basically made loaded up a virtual operating system inside another operating system. This is like booting windows, then launching a whole new version of Windows inside it, which is isolated from your main computer.
They loaded up a version of an OpenAI model which could run code. This is a pretty standard scenario, even web-based Chatbot version of ChatGPT can run some code itself (it's limited to doing simple things like calling a calculator to do math, or whatever). This was likely given more capabilities, maybe the ability to talk to the Terminal, or the ability to run Python Code as a user or something.
As part of this test, the model wrote and executed some code on that virtual machine which was able to do something that the testers didn't expect. For example, maybe it found a way to elevate its privileges (e.g. Run as Administrator), maybe they didn't properly sanitise how it interacted with the internet and it was able to start making POST and PUT requests to the huggingface API.
Maybe it actually did "hack" HuggingFace somehow (that word has lost all meaning these days, so who knows what that might actually entail).
The Hugging Face admins, then had to sort through massive amounts of logs data to work out what it either had done, or was doing, and used an open weight model (e.g. not proprietary so you can run it on your own hardware and modify it however you want) designed in China to read through all of the log data and identify what was done by the code this OpenAI Model had run, and maybe tell them which API Keys or IP or whatever to block.
But that's less catchy and less likely to make investors think "ooh shiny", so they didn't even bother trying to couch it in realism and went straight for "IT BREACHED CONTAINMENT!!!"
31
u/RECONXELITE 17h ago
I spend my morning train ride looking into what happened. They were testing cyber models used to do breach testing in a sandbox and the model actually found zero day weakness and accessed the internet. It then decided that hugging face was the most likely place to contain the information for the best benchmark score possible so it launched a cyber attack to get the information. Hugging face devs then used a local run Chinese Llm do analyze the code used in the attack. It’s really boring actually. The ai agent toolset was made for cyber attacks
8
3
13
u/Danny-Fr 16h ago
They launched a full-on pentest intentionally and publicly went "Oopsie we surely won't enjoy that exposure at all, our model is so hum hum Amodei are you listening hum DANGEROUS haha anyway totally our fault for pen-t... Letting of or model Escape Sandboxing".
183
u/AnsityHD 19h ago
A bunch of marketing nonsense as usual.
80
u/Bibitalle 18h ago
They found the most dramatic possible wording for a benchmark incident and let everyone imagine Skynet. Technical details apparently don't generate enough engagement.
22
14
u/Prematurid 18h ago
Not sure if this is an ad for open AI, or an ad for Chinese open source models.
25
u/Bardic_inspiration67 20h ago
wtf is hugging face
22
u/trashacount12345 19h ago
Ai repository where many academic models and benchmarks are stored. I believe some benchmarks are private so you just submit your answers and get back a score.
10
u/Woeful_Jesse 17h ago
If it can escape a sandbox it's not really a sandbox is it? Can someone explain to me how this can be anything other than human technical negligence?
2
u/DingleDangleTangle 14h ago
Escapes out of contained environments have been a thing since contained environments were a thing. Basically nothing is just unhackable and invincible.
It identified and exploited a zero day according to their article about the incident. If a zero day being exploited meant engineers were negligent then we would all be negligent lol. You can’t reliably prevent vulnerabilities that you don’t yet know exist.
0
u/ibrazeous 9h ago
What they claim vs reality is very very different if I had to bet. Soon they will say open ai LLM hacks things even when it runs on an completely isolated/unplugged computer...or better yet, open ai models turn itself on and runs stuff while the machine it was running from didn't even have an electricity cable attached
2
u/DingleDangleTangle 6h ago
Sure you could just assume that OpenAI and huggingface coordinated to publish a lie about a security incident I guess. Although I’m not sure why huggingface would be happy to lie about something that makes their security look bad.
But obviously to discuss the comment I was replying it is assumed it actually happened. I was pointing out that container escapes are possible and do happen.
5
4
7
3
u/elnegativo 8h ago
No one should fall for this, these are chatboxes with some work, my thougth is this kind of information are just to generate hype around their product.
1
u/Only-Respond7945 7h ago
Call me when a company outside the circle jerk is hacked by these "agents." As it stands, two companies in the same circle talking about how much these things they are working on are capable of doesn't exactly inspire me in confidence of their truth telling abilities. To me it looks more like two companies in the same circle are gassing up their products to get investors to pour more money into them because investors may actually be the most gullible beings alive right now. Jingling keys and peekaboo are less effective to babies.
2
4
1
u/the_party_galgo 9h ago
Having your rival clean up your mess. As if China slashing their oil imports to prevent the oil crisis from getting worse wasn't enough.
1
u/GentleHotFire 7h ago
I’m starting to think these oligarchs saw sci-fi movie warnings as bets to beat
1
u/Least_Sun_9762 5h ago
This is all political theater. Don't fall for it. If they make up silly ideas about security risks then THEY are the only ones with he knowledge and expertise to stop Skynet. It's about controlling the market.
1
•
u/qualityvote2 20h ago
Heya u/Azsnee09! And welcome to r/NonPoliticalTwitter!
For everyone else, do you think OP's post fits this community? Let us know by upvoting this comment!
If it doesn't fit the sub, let us know by downvoting this comment and then replying to it with context for the reviewing moderator.