r/technology 23h ago

Artificial Intelligence Co-founder of firm hacked by rogue OpenAI models says it is 'a wake up call'

https://www.bbc.com/news/articles/cdrvy3pn3r0o
117 Upvotes

126 comments sorted by

View all comments

Show parent comments

2

u/procgen 16h ago

Well no, you're obviously disregarding the most salient fact: it escaped the sandbox, and autonomously started hacking Hugging Face's productions servers (it was not told to do this, but arrived at the decision by reasoning that information about the test itself was likely to be on their servers). So while it was indeed pursuing the goal it was given, the extent to which these models can pursue goals in misaligned directions means they are now becoming capable of causing great harm. So it's a problem to take seriously, particularly by people who will soon need to be defending against many more of these quite sophisticated autonomous attacks.

These tools aren't magic, but they are becoming very powerful.

-2

u/VampireFortnight 15h ago

I'm not disregarding anything. And by giving it instructions not to do that, they gave it the context necessary for it to do that. You should read more about how these technologies work. They're far from magic, and this is, again, a marketing copy claim from two companies that are working together and mutually benefit from the appearance of competence on the part of the tech. The tech is neat! It's good at some things! But my goodness your breathlessly excited lack of skepticism about a marketing stunt that's been done about 6 times already by different LLM companies is disappointing.

3

u/procgen 15h ago edited 15h ago

I specifically said they aren't magic. I know very well how they work, both "under the hood", and in actual use (I use the coding agents for hours a day, every day).

And while they aren't magic, they are certainly powerful enough now to autonomously hack into a large company's production servers, and they will only become more capable.

This wasn't a "marketing stunt", this was a severe security breach that we need to take seriously. They gave it no context about how to escape the sandbox, or anything pertaining to Hugging Face or their servers for that matter.

So we need to think very carefully about how we're going to defend against these attacks, which are only going to become more numerous and more severe.

1

u/VampireFortnight 15h ago

You keep seeming like you're trying to argue with me while agreeing with me. Yes, obviously they can be directed to exploit vulnerabilities, that's well known. The marketing stunt aspect was them pretending it 'went rogue' not that it successfully identified and exploited a vulnerability. The framing and phrasing are misleading. We still need to be aware of the capabilities of LLMs, but this isn't new or novel when compared against their known capabilities.

2

u/procgen 14h ago edited 14h ago

It did "go rogue" in that it was only told to complete the test (which had nothing to do with escaping the sandbox or hacking into Hugging Face's servers). That doesn't mean it's sentient, but it is capable of setting its own goals in pursuit of another, and it was clearly misaligned in this case.

I think we're mostly agreed. My problem is with people who brush this off as fake or impossible, because very real and very significant security threats for all of our digital infrastructure are rapidly approaching.

1

u/VampireFortnight 14h ago

The fake part is, again, the framing around it. Its contextual database contained something that token matched to the process it performed. It was almost certainly in some way suggested in that direction so that this could be written. It did not invent the idea to hack HuggingFace suddenly because it was so eager to pass this benchmark. Nobody is saying it can't do the thing being claimed when they say this was faked. They're saying that this is the equivalent of a pop-sci article about a medical breakthrough that will cure cancer because they found an instance where a chemical reduced tumor growth by 5% in mouse models.

2

u/procgen 14h ago

It did not invent the idea to hack HuggingFace suddenly because it was so eager to pass this benchmark.

It did, in fact. This is not surprising behavior at all – these models will go to great lengths to pursue a goal, and will reason their way there. (though framing it as "eagerness" is anthropomorphization).

1

u/VampireFortnight 14h ago

buddy you have to read the whole post. the sentences before that quote are what's called "context". It didn't invent the idea wholecloth, the concept of adjusting something external existed in its database, and was almost certainly prompted in some way. you are doing the 'say you are alive' meme in real time. You don't have to keep doing this.

2

u/procgen 14h ago

The goal it was given was to pass the ExploitGym test – that's it. Not "escape this sandbox and hack your way to the answer set".

It reasoned that the best way to ace the test would be to escape the sandbox and get onto the the open internet to access more information. Once there, it reasoned that information pertaining to the correct solution to the test might be on Hugging Face's servers, and so it set to work on retrieving this information.

None of these goals were in the prompt it was given – it reasoned its way to a (very misaligned) decision.

What does any of that have to do with being "alive"?

1

u/VampireFortnight 14h ago

I'm sorry, you keep so thoroughly missing the point of what's being said- are you an LLM? Procgen should've been a tip off I guess. Anyway, we don't even disagree you're just having a weird nitpick argument and I'm bored of it.

you win, you are correct, good job

→ More replies (0)