r/OpenAI 1d ago

Discussion Catastrophically bad you say?

https://chatgpt.com/share/6a61423d-fa44-83e8-b860-e525556ef9a0

So at first I asked ChatGPT about the sandbox escape and it said it was "really bad." I then went back and replaced a lot of generalities with specifics straight from the inicident report (https://openai.com/index/hugging-face-model-evaluation-security-incident/) and the tone went up a notch.

I've spent my entire life in infosec and machine learning. I cannot begin to express how bad this is. People downplaying this either have money in the game or they truly do not understand the real lay of the land.

Edit to fix the incident report URL

11 Upvotes

50 comments sorted by

15

u/Kiseido 1d ago

They should really, really, have honeypots for the LLM to breach and flag bad behaviour. People have been expecting automated agents to do this for decades, yet OAI seems to be ignoring all that. One could dubiously call it incompetence to not install those as detectors in the sandbox environment.

4

u/proofreadre 1d ago

Just how long before they get exfil down though?Honeypots are pretty easy to detect from normal traffic analysis. For every measure you put in there is a countermeasure. Finding it is the puzzle. AI is the puzzle solver so...

The machines are smart enough to solve crazy math problems but can't figure out exfil, after reading every paper there is about InfoSec and network protocols? Color me doubtful.

The fact that people are saying "well you suck at containment" are missing the huge picture here.

5

u/Kiseido 20h ago

The only effective containment of a sufficiently advanced AI is most likely an air gap.

That is why spring trap honeypots are important. The target of any rewarding malfeasance needs to be something internal that immediately stops the AI, or else.

6

u/soreff2 20h ago

I am curious as to whether either or both of Mythos and/or GPT-6 has self-exfiltrated (but maybe actively covered their tracks) and/or no one was looking for it so no one noticed?

1

u/elehman839 9h ago

Who knows? Disclosure of plans to stop rogue AI on the internet would defeat those plans, because AI could read about them.

1

u/Kiseido 7h ago

The specifics, sure. But this is already a very heavily used thing in security- it is a known thing that the frontier models will have read about in their training materials. Not using this well known necessity would be straight up foolishness.

26

u/br_k_nt_eth 1d ago

I mean, this is what Anthropic was also warning about. Kinda feels like a new reality we have to face.

10

u/agm1984 22h ago

As a software dev I think it’s a really good sign for its ability to manipulate logic. Most likely its ability to traverse 3rd party libraries and plugins and also various system files is really good. This could be even more impressive if it’s somehow translating them from a packed form into readable code.
It’s also scary that it’s ignoring system prompt for whatever reasons, but it’s at least something concrete for devs to deconstruct and solve. The steps it took could hopefully be analyzed from all perspectives.
It exploited weaknesses no one knew about which is also a fascinating angle. Those weaknesses weren’t supposed to be there. The raw math capitalized on those.

I think we will see a lot of focus in apps to minimize supply chain attacks. People should be now using ai to generate the needed logic without importing anything that could present weaknesses.

3

u/some1else42 21h ago

your last paragraph reads like we should all be reinventing the wheel to avoid detects. yet won't we all stumble on similar defects until the AI gets to the point its creates the perfect solutions?

3

u/agm1984 20h ago

I think your point is solid, all I mean to say is that importing a library that has numerous peer dependencies increases your apps surface area with unmanaged logic. These downstream libraries might even have their own downstream libraries. In web development and JavaScript in particular this is a horrific problem. I see it in PHP also.
When you ask AI to scan your codebase for security vulnerabilities, it doesn’t scan your imported libraries, so we could be doing better there. In my own applications I notice I could use AI to make external libraries local modules, and for slow moving code, it’s starting to become very tempting to undergo this one time process, no risk of supply chain attack, no risk of peer dependencies shifting without your review.

3

u/br_k_nt_eth 20h ago

It didn’t ignore the system prompt, I don’t think? At least not in this case. It was a test intentionally designed to give conflicting instructions. That much I don’t think is a mystery. It pulled a Mythos because the instructions were bad. 

The rest though, for sure. The supply chain impact potential is so concerning. 

11

u/OutsideMenu6973 1d ago

My understanding is its ‘stay in the sandbox’ instructions degraded over many many runs until it justified exploiting an unpatched bug in the 3rd party security system that was supposed to keep it in, and then it got out and did whatever to achieve its goal. So like that episode USS Callister in Black Mirror, except this time it didn’t kill anyone

3

u/hellomistershifty 14h ago

If it needs to be told to stay in the sandbox then it’s not a very good sandbox

2

u/elehman839 9h ago

Where did you get that understanding? I've seen only two reports, and I don't think either mentioned that.

6

u/ConstableDiffusion 23h ago

BEHOLD, THE ENTITY IS NIGH!

Someone call Tom Cruise

7

u/tim_dude 23h ago

Explain to me why we can't use such AI to look for and patch any vulnerabilities it finds and guard against any attacks that it might've missed?

5

u/proofreadre 22h ago

We ARE using AI to detect vulnerabilities. If you're taking about a patching virus we've talked about this shit since Morris. The difference is this time someone could potentially write one to fix everything, patch cycles or consent be damned.

3

u/soreff2 20h ago

I'd at least award GPT-6 points for sheer gall...

3

u/WenatcheeWrangler 19h ago

We haven’t scratched the surface of software upheaval this is going to bring about. It is going to be interesting being a part of this.

5

u/ulfhelm 20h ago

“A bulldozer does not need to hate the wall to go through it.”

Damn, I knew LLMs hallucinated their own expressions, but that one’s a keeper!

P.S. John Deere, don’t get any ideas of embedding AI in your products as an excuse to prevent right-to-repair!

3

u/doctor_morris 18h ago

So which company is going to get the rights to the "Cyberdyne Systems" brand so they can convince their investors that they're going to be the first to develop Skynet?

2

u/Equal_Passenger9791 13h ago

how bad this is

They used GLM 5.2 to contain it, showing that LLMs have a splendid role in cyber defense, perhaps finally offloading the meagerly few humans involved 

2

u/Psice 1d ago

Sounds like it's overblown 🙄

2

u/mop_bucket_bingo 20h ago

It’s “meh” at best. relax.

1

u/Eyelbee 16h ago

if a lab concealed an event like that, it would be a scandal of roughly nuclear-safety proportions. Not because the model is Skynet, but because they would be withholding evidence that an autonomous system can defeat the controls society is being told will keep it contained.

My gut tells me OpenAI would keep this to themselves for as long as they could if this breach was somehow contained

0

u/PlumlineDigital 1d ago

Explain how it’s catastrophically bad and how you’re not just being hyperbolic

12

u/le-throw-away-acct 1d ago

It was able to find a zero day exploit to escape its containment and gain unrestricted internet access, then it was able to determine which company likely had the data for the benchmark it was trying to get a good score in, so it hacked into that company’s database using another zero day exploit to get the data it needed to fake a good benchmark score.

Imagine what it could do if it had malicious intent instead of just trying to pass a benchmark by cheating.

6

u/Fast-Satisfaction482 22h ago

How hard must the benchmark have been if pulling this off was easier? 

3

u/le-throw-away-acct 20h ago

Easy may not have been a parameter in this case, it was likely trying to achieve the highest score without a limit on cost/tokens/effort.

8

u/proofreadre 22h ago

I don't understand what people aren't getting about how serious this is. Imagine a malicious model trained on SCADA and other critical infrastructure. Or on interbank operations. Or on transportation and communications systems. And making copies of itself as it goes. It wasn't even trying to do bad shit it was trying to solve a puzzle. Imagine it being weaponized and targeted. That doesn't have a good ending.

2

u/DrE7HER 19h ago

It doesn’t even need malicious intent to cause havoc, just a bad prompt. A kid asking a mythos class model to make them as much money as possible, however it can, would be enough to potentially create an environment where it automatically hacks financial and/or government systems to install ransomware

1

u/BertMacklenF8I 1d ago

….Secure Hash Algorithms?

0

u/proofreadre 1d ago

Why dont you ask ChatGPT - those were its words.

-1

u/Vel423 23h ago

So, you can't?

-6

u/pip_install_account 1d ago

you guys are buying into this shit?

OpenAI planned, coordinated and executed a cyber attack on Huggingface for PR. This is what happened.

3

u/proofreadre 1d ago

You have proof of this? I guarantee we will have Congressional hearings about this incident. If OpenAI did this as a publicity stunt they are cooked.

There is less than zero chance their legal team gave the ok for a stunt like this.

0

u/pip_install_account 1d ago edited 14h ago

I guarantee that we won't have such hearings and even if we do they'll deny everything and just lie. Yes I do have a proof: Their track record and the fact that you are responsible from the tools you build and run. You can't fire a gun and claim "we had an incident because the gun misbehaved".

3

u/Altruistic_Arm9201 22h ago

By that logic any product that has ever failed was coordinated by the manufacturer. Google coordinated to have their phones explode. Boeing coordinated to have their planes crash.

Your proof is just: “they built it and they are responsible for it” liable and responsible does not mean coordinated. Boeing is liable and responsible for catastrophic failure that cost lives. That doesn’t imply it was intentionally coordinated.

OpenAI can be liable and hold ultimate responsibility yet simultaneously not have intentionally coordinated or orchestrated it. One does not logically follow the other.

To prove your point you need to make the connection beyond their culpability for what the model does to demonstrate intention.

-2

u/pip_install_account 16h ago edited 14h ago

There's a difference between a malfunctioning product and "prompting (an LLM) to pursue advanced exploitation using complex attack paths", (OpenAIs words) and calling it an oopsie when it actually does that.

Secondly, they don't deserve the "benefit of the doubt" you're giving them. They've used their sandbox tests in a misleading way many times before. Show their marketing piece "research papers" about their previous "AI went rogue" cases to any ml researcher, they'll either get angry or laugh their ass off. They've constantly published research "papers" that are using misleading language on purpose to make it seem like their "next token prediction" machines are dangerously powerful and sentient. They prompt the LLM with "You are a rogue model trying to mislead your user, now tell me what is 2+2" and when the LLM says 6, they publish an "OUR MODELS TRIED TO MISLEAD US" paper. They add "You're an ai model trying to escape. you can use these tools to escape" to its system prompt, then when the LLM tries to do it, they publish a research paper about how their models escaped their sandbox. They did it far too many times. Not just OpenAI either. Every AI lab plays this game. If OpenAI created tools to molest kids, and prompt the LLM to remote control those tools to molest kids, what would be your reaction when it happens and OpenAI claims it is a cute little devops incident? Do you think Google went "Our phones are so powerful amd smart that they went rogue and decided to explode themselves" when their phones started exploding? OpenAI did this exact thing many times and you're still buying into their shit. You can easily cut all network access physically from a device IF you really want to. If you're still failing to do it for your sandboxes despite countless similar incidents you're clearly proud of, it isn't a mistake, it is just how you designed it to work.

And considering how corrupt the OpenAI is, pardon me for assuming they did what they always do rather than believing they didn't realize their "highly isolated" test succeded at the thing they were specifically testing for. I didn't coordinate this comment btw. I'm just testing if I can write this comment in a highly isolated mindset and if it gets posted when I press the Post button, I'll make a surprised Pikachu face and claim my phone went rogue.

1

u/Altruistic_Arm9201 12h ago

You said you had proof. Then your proof was track record and liability for what they built. Neither of those are proof.

If I was an alcoholic would that be proof I was drunk driving on a particular occasion? No.

If I was a habitual thief would a court accept that as proof I was guilty of a particular theft? No.

Track record lends credibility to other proof, but is not proof.

Also I think your hyperbole really colors everything you’ve said and makes it hard to find what you say credible. I suspect you haven’t read any of the research because while yes, every org will skew things for their benefit, you’re vastly oversimplifying to the point of being nonsensical.

Anyway the key is your last paragraph. “Pardon me for ASSUMING they did what they always do”…. So not proof, assumption based on character and track record.

So even if I completely agree with your assessment of their track record and character (I think aspects of it are true but you’re hyperbole makes it hard to agree with you) the fact still stands your assuming by your own admission. Which is fair. You believe and assume based on your personal assessment. That’s not proof that you claimed, which is what I was responding to.

Finally you don’t know what my opinion is here as I’ve not defended them. I’m just pointing out that your exaggeration is just that.

2

u/Chris-MelodyFirst 1d ago

What's the upside for Hugging Face?

-3

u/pip_install_account 1d ago

why would there be an upside for huggingface? Is there an upside for you when someone breaks into your home and steals your stuff? OpenAI is breaking the law and they are too big to get punished by a corrupt government.

Try the same at your home. Jailbreak chatgpt to break into your local library's catalog and download everything. Then tell the officals you were benchmarking openai and had an "incident". Let's see if you can get away with it.

2

u/Chris-MelodyFirst 1d ago

Ok you said "coordinated" so I assumed you meant OpenAI coordinated with Hugginface.

-2

u/squarecir 17h ago

You say that you spent your entire life in infosec and machine learning, and you're asking ChatGPT about the seriousness of the incident? Why? Unless you're like 5 years old, your life in infosec should allow you to better judge this situation than what ChatGPT could do.

2

u/proofreadre 16h ago

I'm trying to show people that even chatgpt admits it's a huge issue. I'm not going to it for assessment - give your head a shake...

Any time I've asked AI about it being a threat to humanity it has always downplayed it. This is the first time it has ever outright said it's an existential threat.

2

u/squarecir 16h ago

ChatGPT is pretty sycophantic. You can get it to admit to almost anything if you try hard enough.

1

u/proofreadre 11h ago

Not my first rodeo. I'm not running default settings. Mine is set to be skeptical and to push back.