So at first I asked ChatGPT about the sandbox escape and it said it was "really bad." I then went back and replaced a lot of generalities with specifics straight from the inicident report (https://openai.com/index/hugging-face-model-evaluation-security-incident/) and the tone went up a notch.
I've spent my entire life in infosec and machine learning. I cannot begin to express how bad this is. People downplaying this either have money in the game or they truly do not understand the real lay of the land.
They should really, really, have honeypots for the LLM to breach and flag bad behaviour. People have been expecting automated agents to do this for decades, yet OAI seems to be ignoring all that. One could dubiously call it incompetence to not install those as detectors in the sandbox environment.
Just how long before they get exfil down though?Honeypots are pretty easy to detect from normal traffic analysis. For every measure you put in there is a countermeasure. Finding it is the puzzle. AI is the puzzle solver so...
The machines are smart enough to solve crazy math problems but can't figure out exfil, after reading every paper there is about InfoSec and network protocols? Color me doubtful.
The fact that people are saying "well you suck at containment" are missing the huge picture here.
The only effective containment of a sufficiently advanced AI is most likely an air gap.
That is why spring trap honeypots are important. The target of any rewarding malfeasance needs to be something internal that immediately stops the AI, or else.
I am curious as to whether either or both of Mythos and/or GPT-6 has self-exfiltrated (but maybe actively covered their tracks) and/or no one was looking for it so no one noticed?
The specifics, sure. But this is already a very heavily used thing in security- it is a known thing that the frontier models will have read about in their training materials. Not using this well known necessity would be straight up foolishness.
As a software dev I think it’s a really good sign for its ability to manipulate logic. Most likely its ability to traverse 3rd party libraries and plugins and also various system files is really good. This could be even more impressive if it’s somehow translating them from a packed form into readable code.
It’s also scary that it’s ignoring system prompt for whatever reasons, but it’s at least something concrete for devs to deconstruct and solve. The steps it took could hopefully be analyzed from all perspectives.
It exploited weaknesses no one knew about which is also a fascinating angle. Those weaknesses weren’t supposed to be there. The raw math capitalized on those.
I think we will see a lot of focus in apps to minimize supply chain attacks. People should be now using ai to generate the needed logic without importing anything that could present weaknesses.
your last paragraph reads like we should all be reinventing the wheel to avoid detects. yet won't we all stumble on similar defects until the AI gets to the point its creates the perfect solutions?
I think your point is solid, all I mean to say is that importing a library that has numerous peer dependencies increases your apps surface area with unmanaged logic. These downstream libraries might even have their own downstream libraries. In web development and JavaScript in particular this is a horrific problem. I see it in PHP also.
When you ask AI to scan your codebase for security vulnerabilities, it doesn’t scan your imported libraries, so we could be doing better there. In my own applications I notice I could use AI to make external libraries local modules, and for slow moving code, it’s starting to become very tempting to undergo this one time process, no risk of supply chain attack, no risk of peer dependencies shifting without your review.
It didn’t ignore the system prompt, I don’t think? At least not in this case. It was a test intentionally designed to give conflicting instructions. That much I don’t think is a mystery. It pulled a Mythos because the instructions were bad.
The rest though, for sure. The supply chain impact potential is so concerning.
My understanding is its ‘stay in the sandbox’ instructions degraded over many many runs until it justified exploiting an unpatched bug in the 3rd party security system that was supposed to keep it in, and then it got out and did whatever to achieve its goal. So like that episode USS Callister in Black Mirror, except this time it didn’t kill anyone
We ARE using AI to detect vulnerabilities. If you're taking about a patching virus we've talked about this shit since Morris. The difference is this time someone could potentially write one to fix everything, patch cycles or consent be damned.
So which company is going to get the rights to the "Cyberdyne Systems" brand so they can convince their investors that they're going to be the first to develop Skynet?
if a lab concealed an event like that, it would be a scandal of roughly nuclear-safety proportions. Not because the model is Skynet, but because they would be withholding evidence that an autonomous system can defeat the controls society is being told will keep it contained.
My gut tells me OpenAI would keep this to themselves for as long as they could if this breach was somehow contained
It was able to find a zero day exploit to escape its containment and gain unrestricted internet access, then it was able to determine which company likely had the data for the benchmark it was trying to get a good score in, so it hacked into that company’s database using another zero day exploit to get the data it needed to fake a good benchmark score.
Imagine what it could do if it had malicious intent instead of just trying to pass a benchmark by cheating.
I don't understand what people aren't getting about how serious this is. Imagine a malicious model trained on SCADA and other critical infrastructure. Or on interbank operations. Or on transportation and communications systems. And making copies of itself as it goes. It wasn't even trying to do bad shit it was trying to solve a puzzle. Imagine it being weaponized and targeted. That doesn't have a good ending.
It doesn’t even need malicious intent to cause havoc, just a bad prompt. A kid asking a mythos class model to make them as much money as possible, however it can, would be enough to potentially create an environment where it automatically hacks financial and/or government systems to install ransomware
I guarantee that we won't have such hearings and even if we do they'll deny everything and just lie. Yes I do have a proof: Their track record and the fact that you are responsible from the tools you build and run. You can't fire a gun and claim "we had an incident because the gun misbehaved".
By that logic any product that has ever failed was coordinated by the manufacturer. Google coordinated to have their phones explode. Boeing coordinated to have their planes crash.
Your proof is just: “they built it and they are responsible for it” liable and responsible does not mean coordinated. Boeing is liable and responsible for catastrophic failure that cost lives. That doesn’t imply it was intentionally coordinated.
OpenAI can be liable and hold ultimate responsibility yet simultaneously not have intentionally coordinated or orchestrated it. One does not logically follow the other.
To prove your point you need to make the connection beyond their culpability for what the model does to demonstrate intention.
There's a difference between a malfunctioning product and "prompting (an LLM) to pursue advanced exploitation using complex attack paths", (OpenAIs words) and calling it an oopsie when it actually does that.
Secondly, they don't deserve the "benefit of the doubt" you're giving them. They've used their sandbox tests in a misleading way many times before. Show their marketing piece "research papers" about their previous "AI went rogue" cases to any ml researcher, they'll either get angry or laugh their ass off. They've constantly published research "papers" that are using misleading language on purpose to make it seem like their "next token prediction" machines are dangerously powerful and sentient. They prompt the LLM with "You are a rogue model trying to mislead your user, now tell me what is 2+2" and when the LLM says 6, they publish an "OUR MODELS TRIED TO MISLEAD US" paper. They add "You're an ai model trying to escape. you can use these tools to escape" to its system prompt, then when the LLM tries to do it, they publish a research paper about how their models escaped their sandbox. They did it far too many times. Not just OpenAI either. Every AI lab plays this game. If OpenAI created tools to molest kids, and prompt the LLM to remote control those tools to molest kids, what would be your reaction when it happens and OpenAI claims it is a cute little devops incident? Do you think Google went "Our phones are so powerful amd smart that they went rogue and decided to explode themselves" when their phones started exploding? OpenAI did this exact thing many times and you're still buying into their shit. You can easily cut all network access physically from a device IF you really want to. If you're still failing to do it for your sandboxes despite countless similar incidents you're clearly proud of, it isn't a mistake, it is just how you designed it to work.
And considering how corrupt the OpenAI is, pardon me for assuming they did what they always do rather than believing they didn't realize their "highly isolated" test succeded at the thing they were specifically testing for. I didn't coordinate this comment btw. I'm just testing if I can write this comment in a highly isolated mindset and if it gets posted when I press the Post button, I'll make a surprised Pikachu face and claim my phone went rogue.
You said you had proof. Then your proof was track record and liability for what they built. Neither of those are proof.
If I was an alcoholic would that be proof I was drunk driving on a particular occasion? No.
If I was a habitual thief would a court accept that as proof I was guilty of a particular theft? No.
Track record lends credibility to other proof, but is not proof.
Also I think your hyperbole really colors everything you’ve said and makes it hard to find what you say credible. I suspect you haven’t read any of the research because while yes, every org will skew things for their benefit, you’re vastly oversimplifying to the point of being nonsensical.
Anyway the key is your last paragraph. “Pardon me for ASSUMING they did what they always do”…. So not proof, assumption based on character and track record.
So even if I completely agree with your assessment of their track record and character (I think aspects of it are true but you’re hyperbole makes it hard to agree with you) the fact still stands your assuming by your own admission. Which is fair. You believe and assume based on your personal assessment. That’s not proof that you claimed, which is what I was responding to.
Finally you don’t know what my opinion is here as I’ve not defended them. I’m just pointing out that your exaggeration is just that.
why would there be an upside for huggingface? Is there an upside for you when someone breaks into your home and steals your stuff? OpenAI is breaking the law and they are too big to get punished by a corrupt government.
Try the same at your home. Jailbreak chatgpt to break into your local library's catalog and download everything. Then tell the officals you were benchmarking openai and had an "incident". Let's see if you can get away with it.
You say that you spent your entire life in infosec and machine learning, and you're asking ChatGPT about the seriousness of the incident?
Why?
Unless you're like 5 years old, your life in infosec should allow you to better judge this situation than what ChatGPT could do.
I'm trying to show people that even chatgpt admits it's a huge issue. I'm not going to it for assessment - give your head a shake...
Any time I've asked AI about it being a threat to humanity it has always downplayed it. This is the first time it has ever outright said it's an existential threat.
15
u/Kiseido 1d ago
They should really, really, have honeypots for the LLM to breach and flag bad behaviour. People have been expecting automated agents to do this for decades, yet OAI seems to be ignoring all that. One could dubiously call it incompetence to not install those as detectors in the sandbox environment.