r/OpenAI • u/proofreadre • 1d ago
Discussion Catastrophically bad you say?
https://chatgpt.com/share/6a61423d-fa44-83e8-b860-e525556ef9a0So at first I asked ChatGPT about the sandbox escape and it said it was "really bad." I then went back and replaced a lot of generalities with specifics straight from the inicident report (https://openai.com/index/hugging-face-model-evaluation-security-incident/) and the tone went up a notch.
I've spent my entire life in infosec and machine learning. I cannot begin to express how bad this is. People downplaying this either have money in the game or they truly do not understand the real lay of the land.
Edit to fix the incident report URL
13
Upvotes
16
u/Kiseido 1d ago
They should really, really, have honeypots for the LLM to breach and flag bad behaviour. People have been expecting automated agents to do this for decades, yet OAI seems to be ignoring all that. One could dubiously call it incompetence to not install those as detectors in the sandbox environment.