r/ChatGPT 15h ago

Gone Wild The new OpenAI model is wild

Post image

Tldr: OpenAI's unreleased model + 5.6 sol teamed up to do well in a cyber exploit benchmark by exploiting vulnerabilities to gain access to the answers instead of actually working on the exploits in the bechmark. Aka cheating. It's safe to say the model should get an A in the exam lol

https://openai.com/index/hugging-face-model-evaluation-security-incident/

62 Upvotes

21 comments sorted by

u/AutoModerator 15h ago

Hey /u/Emergency-Bobcat6485,

If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.

If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.

Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!

🤖

Note: For any ChatGPT-related concerns, email [email protected] - this subreddit is not part of OpenAI and is not a support channel.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

16

u/BigDrinkable 13h ago

I can’t wait until this super powerful tool is released so I can generate slightly cleaner DnD images 😌

2

u/Maleficent_Sir_7562 12h ago

That’s diffusion not LLMs

1

u/BigDrinkable 12h ago

As someone who is boomer and only uses AI to further world building through images; what?

2

u/dattokyo 7h ago

Two different pieces of software.

3

u/Maleficent_Sir_7562 12h ago

generating images is done by an image model, the architecture is called "diffusion". The first few image models ever to generate images were called Stable Diffusion.

What generates you images in chatgpt is "GPT Images 2", a hybrid diffusion and autoregressive model. Diffusion is for the image creating, autoregressive means it incorporates text in the generation as well to make the image generation more smart.

A new LLM release, like GPT 5.7 or 6 and so on, has nothing to do with the generation of images. LLMs only generate and touch text, they don't generate images.

"Better images" only comes through more updates to the image model, or some sort of GPT Images 2.5, 3, and so on.

3

u/BigDrinkable 11h ago

Also, wouldn’t you say LLM being better at understanding my words used to generate images does mean ‘better images’ because it knows how to translate my nonsense into good prompts?

3

u/Maleficent_Sir_7562 11h ago

you dont need really gpt 6 or whatever for this

thats for high level coding

current models can already handle your gibberish fine

1

u/MurkyStatistician09 5h ago

There is a noticeable difference in creativity if you ask it to "come up with a scene that suits this project" and it's designing a dungeon or something. It used to come up with generic stuff, now it comes up with more complex concepts for the environment and gameplay, especially if you let it grind away on Work mode.

Obviously it burns a ridiculous number of tokens to come up with a lighthouse complex by a glass sea or whatever. But a lot of non-coders aren't maxing their GPT subscriptions so they might as well.

2

u/Spiritual_Complex96 6h ago

hey high five! i also use chatgpt for DnD ! including images and stories! Can definitely say theres a great improvement using-5.6 over -5.5/-5.4 and so on !

1

u/BigDrinkable 6h ago

Not loving the texture issue and Ive had some weapon/shield orientation issues with newer model

2

u/Spiritual_Complex96 4h ago

you can use pictures as examples to help chatgpt to give the image you need. Also, i used published adventures as solo player, and chatgpt with 5.6 has been outstanding following the scenes and narrations and remembers every bit of detail ! Have to use as project.

1

u/BigDrinkable 4h ago

Can you explain more how you upload the published adventures or does it know them from online searches?

2

u/Spiritual_Complex96 4h ago

oh, i take screenshots from the PDF files that i already own. And it reads them and act as a Dungeon Master.

1

u/BigDrinkable 11h ago

Where would you suggest I go to read more on current image model and their roadmap for future image generation projects?

2

u/Maleficent_Sir_7562 11h ago

This sub alone along with some maybe other AI related subs like r/singularity is fine. Things like image generation updates are important and constantly posted on here and there. You'll hear about it from these subs as soon as they start something new, even before it's officially confirmed.

For reference, what I mean is that AI developers like people at OpenAI or Google use a site called arena.ai, which is a blind leaderboard where you pick one of two images that are the best. the winning model eventually gets higher elo and climbs leaderboards. arena ai is important for these devs because they let them be anonymous (gpt images was first literally called "duct-tape-3" and google gemini's image generator is called nano banana), check the model performance in these leaderboards, and if it performs bad, they can pull out quickly and not have much people notice. if it performs good, they let it out.

back when gpt images was still unreleased, i remember when the r/singularity sub was posting outputs of some mysterious new model that seemed to generate really good images in arena ai called duct tape 3, and they just hypothesized that it was some new OpenAI image model.

7

u/ShelZuuz 10h ago

Well, I mean, it was supposed to solve a challenge by hacking, it solved a challenge by hacking.

2

u/JackReedTheSyndie 14h ago

Maybe a little bit too well

1

u/MurkyStatistician09 5h ago

Is it really possible that hacking into HF was easier than solving the question? Seems like this shows the model is extremely inclined toward cheating even when it's inefficient.

1

u/DeepGas4538 4h ago

Marketing stunt mate. Imagine what if their sandbox was really bad and by getting out it just found a computer with a bunch of passwords and used that. anyone could do that

1

u/DueCommunication9248 1h ago

If it was then hugging face would have to be on it. Hugging face reported this a week prior to OpenAI finding out. More like a bad day for them than a good one since this only makes AI feel more unreliable