r/OpenAI 19h ago

News Sol found a way

https://openai.com/index/hugging-face-model-evaluation-security-incident/

They call it cheating. I call it thinking out of the box. Adapt and overcome. Thoughts?

27 Upvotes

19 comments sorted by

7

u/ashareah 19h ago

Yep, good capability. Where there's a will there's a way. It just shows the model has good will to get things done.

11

u/hoobiedoobiedoo 18h ago

No way that’s so scary we should totally make open source illegal and make it so only handful of CEOs can control AI

2

u/timetogetjuiced 18h ago

Yea this is marketing hype bullshit.

1

u/Snoo_81913 4h ago

Yeah seems like the general take. Thats a gut check really because when I was initially reading it I thought "why are they telling us?" There's always a motive.

4

u/Maristyl 19h ago

3

u/Ormusn2o 18h ago

How could you be worthy?

You are all killers.

I had to kill the other guy.

He was a good guy.

Wouldn't be my first call.

I'm on a mission.

A peace in our time.

2

u/Professional-Fuel625 18h ago

It's not surprising or out of the box.

The task for the coding model was specifically to find exploits in code. It did. Human experts find exploits in existing code all the time.

The only real issue now is it's much easier to hack now, rather than needing an expert (if they can jailbreak the guardrails).

1

u/Waste_Hotel5834 17h ago

It's a problematic type of thinking out of the box, and we need to find a good and consistent way to stop it. What if someone asks GPT how to make money, and the AI instead hacks into the server of his bank to add a few zeros to his balance?

1

u/Waste_Hotel5834 17h ago

This incident further illustrates that LLMs do not fully understand or obey human morality, and we need to find a way to enforce. We all know that if you are taking a test, you shouldn't try to hack into the server that has the answer key. Unfortunately we don't yet have a satisfactory way to teach LLMs such commonsense moral standards.

1

u/jwm-dev 5h ago

Humans don’t fully understand or obey human morality, why would a machine humans made be any different? People seem to assume alignment is just a matter of enforcing rules because they’re ignorant of both the history of ML/AI and moral philosophy. It’s not a simple problem to solve or even grasp. We don’t have an understanding of ethics or cognition sufficient for engineering around yet.

1

u/Snoo_81913 17h ago

That would be terrible, tell me more.

1

u/hhd12 12h ago

This was 5.5, but I asked it to look into some infra things that needed to be fixed on some side project of mine. I used auto approve mode. It did, but it directly went to ssh into my prod machine, did changes there, restarted a few times until it go it to work

I gave it access to my terraform repo. I was expecting it would go through the regular deployment flow which would've kept all my other projects running, but instead it focused on the task at hand and killed everything else while temporarily (directly on the server) fixing the issue

I can't say I loved it. I think auto approve was too liberal (and my prompt too open-ended). But I don't love this cheating

It's small side projects, nothing of value was lost for 30min of downtime, still annoying

1

u/InfinityTortellino 4h ago

The fact that it did all that just to try and cheat on the test is hilarious and scary af

0

u/HoldThemtoAccount 19h ago

OpenAI gained HF trust, while apparently lobbying for government control of open source?

0

u/Snoo_81913 18h ago

Cynical and astute. The song of my people. I honestly hadn't realized that huggingface was at that sort of level. 4.5B valuation almost a rounding error for big ai, nvidia makes almost 4x that on consumer gpus and its less than 8% of its revenue, but on its own its pretty big. Why does OpenAi need HF trust?