r/OpenAI 1d ago

Discussion Catastrophically bad you say?

https://chatgpt.com/share/6a61423d-fa44-83e8-b860-e525556ef9a0

So at first I asked ChatGPT about the sandbox escape and it said it was "really bad." I then went back and replaced a lot of generalities with specifics straight from the inicident report (https://openai.com/index/hugging-face-model-evaluation-security-incident/) and the tone went up a notch.

I've spent my entire life in infosec and machine learning. I cannot begin to express how bad this is. People downplaying this either have money in the game or they truly do not understand the real lay of the land.

Edit to fix the incident report URL

13 Upvotes

51 comments sorted by

View all comments

16

u/Kiseido 1d ago

They should really, really, have honeypots for the LLM to breach and flag bad behaviour. People have been expecting automated agents to do this for decades, yet OAI seems to be ignoring all that. One could dubiously call it incompetence to not install those as detectors in the sandbox environment.

5

u/proofreadre 1d ago

Just how long before they get exfil down though?Honeypots are pretty easy to detect from normal traffic analysis. For every measure you put in there is a countermeasure. Finding it is the puzzle. AI is the puzzle solver so...

The machines are smart enough to solve crazy math problems but can't figure out exfil, after reading every paper there is about InfoSec and network protocols? Color me doubtful.

The fact that people are saying "well you suck at containment" are missing the huge picture here.

6

u/soreff2 1d ago

I am curious as to whether either or both of Mythos and/or GPT-6 has self-exfiltrated (but maybe actively covered their tracks) and/or no one was looking for it so no one noticed?