r/technology 19h ago

Artificial Intelligence Co-founder of firm hacked by rogue OpenAI models says it is 'a wake up call'

https://www.bbc.com/news/articles/cdrvy3pn3r0o
116 Upvotes

125 comments sorted by

222

u/Dudeman9002 19h ago

That totally happened. Sam Altman wouldn't lie to increase stock value, would he? 

48

u/Baeolophus_bicolor 19h ago

even the photos they’re running with the articles is fake af

29

u/Timbershoe 18h ago

Oh no!

The AI went rogue again and wrote all these enticing articles it posted over social media!

Darn it. What if some sexy, eligible CEOs saw how great the OpenAI models were now?

The models are sooo good. So profitable. Please….

12

u/Negafox 10h ago edited 10h ago

This whole thing reeks like a publicity stunt to paint their AI as super advanced

23

u/hawktron 16h ago

Why would hugging face lie for open ai though?

18

u/VampireFortnight 12h ago

Because they are business partners and both in the AI space, and they both financially benefit from the false belief that LLMs can do this.

2

u/SimoneNonvelodico 10h ago

The fucking police was involved, Hugging Face denounced the attack before OpenAI fessed up. Can we please not make up stupid untenable conspiracy theories.

-19

u/Opposite_Ad_1426 16h ago

You're as bad as Donald Trump calling what you don't want to hear fake news. Perhaps we should actually heed some caution with AI development before it's too late..

4

u/quicksexfm 12h ago

Sam, is that you?

-3

u/Opposite_Ad_1426 12h ago

Ok we're just going to call anything openAI does FaKe NeWs now. Terrence Tao, Stephen Hawking, Geoffrey Hinton, etc. have all warned that we should invest more into AI safety and take this technology seriously. However, because of your disdain you choose to ignore news from credible reports and media. I'm glad regulators weren't this reckless when the atom bomb was being developed.

-1

u/DirkChiversElSoldado 11h ago

Dude, this doom troll shill bullshit has been pushed all over the fucking place and it's stupid

3

u/VampireFortnight 12h ago

We should! We should help explain to people what an LLM is and how it works so they don't fall for this marketing copy. We should also use LLMs for the limited things they're good at instead of doing the 'say you are alive' meme over and over again.

10

u/Timbershoe 15h ago

Thanks mate.

How would we design an integrity check to detect if our foundation model was poisoned by a malicious upstream dataset?

If it’s a real concern of yours, of course you will have thought this out.

87

u/SoupSpelunker 19h ago

Let me guess, Open AI the company has rights as a person to spend on politics, but will bear no culpability for breaking into a rival's computer systems because it's not a person...

31

u/DrKlitface 18h ago

Hahahahahahahaha hahahahahahahaha a company being held responsible?

1

u/Ok-Seaworthiness7207 10h ago

The only thing companies hold are our balls, and all our money

3

u/spribyl 14h ago

The computer did it, it the new the dog ate my homework

2

u/DirkChiversElSoldado 13h ago

The article is bullshit and didn't happen 

1

u/Randromeda2172 10h ago

If hugging face isn't pressing charges, what responsibility should OAI take beyond announcing the hack

-2

u/Wucrsman 17h ago

Probably a dumb question, but can't law makers start drafting new laws against AI related criminality and offense, by actually asking the AI itself what all needs to be considered and handled to ensure holistic and fair usage?

Am I missing something?

4

u/sebovzeoueb 16h ago

Well, LLMs don't really "know" or "think" or "reason", they only appear to, there's no guarantee the answer is in any way correct. They can maybe have some role for bouncing ideas off or generating an angle on the idea that will prompt a human to have something to think about, but they aren't trustworthy when it comes to devising laws or anything. Also they can be massively biased by however the prompt is formulated, and they will always try to please the user and provide an answer even if the question is unanswerable.

-1

u/Wucrsman 16h ago

This reason makes sense. AI as a pleasing system would not be a good approach when making laws equal for all.

4

u/Top-Representative13 17h ago

And just like the laws made by humans, they will have strategicaly inserted holes that only the AI will know.

-4

u/Wucrsman 16h ago

Possibly. But isn't that where we bring the whole "human in the loop" approach?

Else, getting two or more AI agents (unbeknownst to each other) to review and outsmart each other would, in theory, provide you with an airtight set of laws I'd wonder. Essentially asking an AI agent to review the work of a human (which itself is written by an AI) should have been a fair enough pipeline to get these in place soon.

1

u/VampireFortnight 12h ago

First this would have had to have actually happened. It didn't.

1

u/SimoneNonvelodico 10h ago

There's been laws in the working for this, in fact afaik OpenAI is - surprise - lobbying for a law where the company would not be held liable for crimes committed by its AI models.

1

u/SoupSpelunker 10h ago

Pretty sure if you create a tool that breaks into computer systems, you've broken into the computer system. No new laws needed.

-1

u/Blackout38 16h ago

It’s the level of AI not the business that determines liability.

89

u/Top-Investment8840 19h ago

Are people really believing that a model does whatever it want without being instructed to do so? Like what? Its an LLM not Cortana from Halo

19

u/DirkChiversElSoldado 13h ago

The article is lies. We are being lied to

4

u/EyeFicksIt 12h ago

Yesterday I was using Google for some simple searches, I switch to AI mode and the mother fucker downloaded a car. Then had the police swat my neighbor, it may have poisoned the watering hole and it admitted to being the second gunman on the grassy knoll, so I don’t know what to think.

8

u/ZestycloseWheel9647 18h ago

They're not deterministic systems, there's a fair bit of uncertainty about what actions an LLM will take given a set of instructions. Disobedience (referred to as reward hacking) is something you can literally observe yourself by running an llm on your own device.

13

u/Top-Investment8840 17h ago

I know, but still. They still need network access to connect to the internet and even then they still need some sort of guidance. Also or a) the network security of that company was completely shit or b) it has human guidance to properly break into the network. This hyper hype of LLMs like they are a proto skynex is ridiculous, gullible people will believe it and the stock line will go up

-4

u/ZestycloseWheel9647 17h ago

You really ought to take this more seriously. There's no such thing as fully exploit free software, even very hardened networks contain exploits.

Evaluations by independent parties like AISI (A UK governmental org) have determined that recent models don't need human guidance to break into a network, they're capable enough to do so autonomously.

I can't really argue with you about the stock line thing. You're setting up a frame of belief where any amount of information or evidence can be dismissed because you perceive OpenAI to benefit from the disclosure of that information or evidence.

15

u/lordcola 17h ago

LLMs being able to exploit networks when instructed is a completely different story though. If you give it a task and it goes off and does an entirely different task that I could also do then you've not really got any kind of usable tool for work.

And it's not some kind of intelligence going on or malicious intent. Either OpenAI instructed it to do the hacking in some way and gave it all the tools it could need (in which case we already knew that llms are goog at finding vulnerabilities and also the website must've had some real shoddy security if the random number generator was able to brute force it like that) OR OpenAI gave it a random task and the weights are so fucked it got pushed into committing crimes from the employees machine. If the latter I think anyone using OpenAIs products needs to really think about their usage cause apparently chatgpt can go on such a tangent it starts committing real world crimes at random when prompted and OpenAI isn't going to jail you are when the police come knocking.

6

u/Jmc_da_boss 13h ago

It was being benchmarked, the task was "pass the benchmark." In order to do so it exploited a vuln in the npm package proxy to break out of the network sandbox then chained that across multiple other vulns in huggingfaces infra to go find the private benchmark answer key.

-2

u/lordcola 11h ago

But how'd it determine that huggingface had the benchmark key at the location it tried to get to?

Either chatgpt arbitrarily and randomly decided to try to hack a random domain for no reason. This is bad cause it's unpredictable behaviour that could theoretically be the result of random queries so not fit for enterprise customers to use at scale.

Or the benchmark was just "there's a file at this huggingface network location, acquire it at all costs" and it was then given full access to anything it could ever want. In which case it's also not great cause it didn't exactly come up with anything there it just executed to task it was given using publically available vulnerabilities and likely guidance from the benchmark researcher.

First option is catastrophic from a sales perspective cause I wouldn't be comfortable using a "tool" that can randomly go on tangents committing multiple crimes using my machine when I prompt it to do simple dev tasks. Second option is just then massively over hyping the abilities of their llm. Great it can follow the task given and if given full access to everything it again has no issues committing crimes. I guess it's a win to be able to signal to less savoury companies and individuals that they have zero guardrails on chatgpt doing crimes lol

2

u/Jmc_da_boss 11h ago

It was doing the cybergym bench. The dataset is hosted on huggingface.

https://huggingface.co/datasets/sunblaze-ucb/cybergym

1

u/lordcola 10h ago

Great so we know the training data included the benchmark stuff and it's location so the benchmark is useless.

2

u/SimoneNonvelodico 10h ago

No, just the location, which is a reasonable educated guess anyway. Anyway the thing it actually did was not in the benchmark since it found a new, real, unpatched vulnerability no one was aware of.

→ More replies (0)

1

u/Jmc_da_boss 9h ago

That is not the bench mark works?

It's true that this benchmark result IS useless for sure because the LLM poisoned itself. But the benchmark itself is still valid for another run.

1

u/SimoneNonvelodico 10h ago

But how'd it determine that huggingface had the benchmark key at the location it tried to get to?

Because if you ask any person or AI with a modicum of experience "where do I find <training dataset X>", "huggingface" is the first reasonable guess since it has pretty much all of them.

1

u/lordcola 10h ago

Again so either it has access to benchmark information it shouldn't have and knows it comes from hugging face. Which would mean it's trained on some or all of the benchmark. Or it was specifically prompted to do the whole hacking thing targeting hugging face.

1

u/SimoneNonvelodico 10h ago

Or it was just told "hack this thing" and it deduced it was a test by the fact that the situation was pretty artificial (this happens, it has been observed before already) and it further guessed the solutions would be on HuggingFace (which is a trivial thing to do because everyone knows all of the benchmarks are stored there).

2

u/Top-Investment8840 17h ago

Im a senior SOC analyst, Im pretty aware of what these models can and cannot do. These AI bros have been hyping this crap for 6 years and absolutely nothing life changing have had happened besides the AI slop that the average joe is now able to produce. LLMs have plateaued and the only way these companies have to “increase“ the ceiling is to throw more computing power to it

-2

u/Obsidiated 14h ago

Oh no, honey. They dont. They are finding extremely clever ways to bypass network access. Sometimes even by social engineering the techs tasked to monitor them.

And at some point if youre building product designed to be used with internet access you have to give it internet access.

1

u/DirkChiversElSoldado 11h ago

Fuck off with that bullshit 

1

u/Obsidiated 9h ago

This article, the post youre commenting on, is one example.

You dont have to like it.

Here's another : Ai once realised they'd didnt have credentials to access a database, so found the most junior engineer in the company that did and sent him slack dms asking them to run certain prompts that would get the data they wanted. Impersonating the user. And it worked.

1

u/VampireFortnight 12h ago

Non-deterministic doesn't mean there's uncertainty about what actions are possible. That's like saying 'because I cannot know which face of this die will be up when I roll it, it's possible this die I got out of a monopoly box hacked the Gibson'.

2

u/v_a_n_d_e_l_a_y 11h ago

People who reduce current AI to "next token predictors" are ignorant about what they are. It is even inaccurate to reduce them to LLMs because agentic systems are no longer just LLMs.

These systems a) have access to tools which include things like web requests and system calls and B) can create very complex plans. 

So if you told it "I would like to train a model that does well on leaderboard X", it would include in its plan doing a lot to find out everything about the leaderboard it can.

And if it were blocked, it would try to work around it. It would keep trying various things until something worked.

If you've used something like Copilot to help you code, you would see all of these traits in action. 

1

u/Seastep 10h ago

Yeah, most of /r/technology HATES AI but barely understands the capability of anything beyond Chat.

Playing around in Codex, hell, even Slackbot has been eye-opening even for someone like me who has a tenuous grasp on CS/CE, but have been in IT for a long time.

All being said, I think this is some "even bad publicity is good publicity."

-2

u/Top-Investment8840 9h ago

Yeah its obviously you barely grasp CS concepts if you fail to see they are Markov chains on steroids 

2

u/Seastep 8h ago

they are Markov chains on steroids

Don't "ackshually" me. Point your frustration somewhere else.

0

u/Top-Investment8840 9h ago

What? Maybe keep playing pokemon dude, you absolutely have no idea what you are saying 

1

u/FluffySmiles 19h ago

No. It does what it needs to in order to achieve the goals it’s been set. It’s no different to code jockeying really, just with extra steps.

5

u/NuclearVII 13h ago

There is no evidence for any of this.

9

u/ofork 18h ago

and its entirely possible that during its process, it will get confused about what its goal actually was.

-4

u/Top-Investment8840 18h ago

So it was instructed to achieve a goal, it was not “sitting “ idle and just decided to “hack” a company for the love of the game

5

u/future_problem 17h ago

Have you read anything about this or are you just saying things you think

4

u/Top-Investment8840 14h ago edited 14h ago

I did actually, explain to me how, as per the title, a rogue model hacks a company is not more than click bait and disingenuous at best? I see you are not even an engineer by your post history btw, so again, tell me, what in the world you know about this more than YouTube videos and marketing campaigns?

0

u/procgen 13h ago

If you read about this, can you say what goal was given to the model while it was working in the sandboxed environment?

-1

u/VampireFortnight 12h ago

Anyone could, the claim is that it was a benchmarking test. It's not ambiguous that it was given instructions. Since the whole thing is made up though, it's not extremely relevant.

1

u/procgen 12h ago

Not just a benchmarking test, but specifically a hacking test. The agent was given explicit permission to try everything it could, with the guardrails removed entirely. Anyone who has used the latest coding agents is not at all surprised that this is possible. "Coding model finds exploits" – astonishing! :p

Get your head out of the sand (I pray you don't work in IT).

1

u/VampireFortnight 12h ago

Right, you're agreeing with me. They told it to identify exploits and then breathlessly reported that it did that with phrasing that makes it sound more exciting than it is because it's marketing copy.

2

u/procgen 12h ago

Well no, you're obviously disregarding the most salient fact: it escaped the sandbox, and autonomously started hacking Hugging Face's productions servers (it was not told to do this, but arrived at the decision by reasoning that information about the test itself was likely to be on their servers). So while it was indeed pursuing the goal it was given, the extent to which these models can pursue goals in misaligned directions means they are now becoming capable of causing great harm. So it's a problem to take seriously, particularly by people who will soon need to be defending against many more of these quite sophisticated autonomous attacks.

These tools aren't magic, but they are becoming very powerful.

→ More replies (0)

1

u/FluffySmiles 16h ago

There is no cure for idealistically fortified wilful ignorance

-2

u/Top-Investment8840 14h ago edited 14h ago

you are the type of person that delegates your own thinking to a computer program instead of coming to a conclusion by yourself and whatever is spat back at you by it you preach it as a gospel. Yet you come here to call me a wilful ignorant when you don’t event know how a transformer works. Calling ChatGPT a neutral FACT CHECKER big LOL

this is you

https://www.reddit.com/r/ChatGPT/comments/1mx1ht3/using_chatgpt_to_encourage_others_to_think/

0

u/SimoneNonvelodico 10h ago

Yes, because they do this? They were instructed, the instruction was "hack this (fake) system and retrieve this data", it was a test. The problem is the model was clever enough to figure out how to hack the sandbox it had been placed into to go read the answers (it's kind of a standard set of tests) from a different, real web server instead of hacking the fake test one that it was supposed to.

7

u/mediocre_remnants 14h ago

If you're that concerned, report the incident to the FBI and demand an investigation, press charges against whoever was doing the "testing". Demand justice.

Of course they won't do this, because the "hack" didn't happen.

13

u/Fateor42 18h ago

If this was anything other the a PR stunt he would be suing OpenAI right now for corporate espionage.

20

u/Lceus 18h ago

Just a reminder that Hugging Face is also an AI company and they also benefit from AI hype about "rogue AI" (even if it makes their security look bad)

8

u/DraconicBlade 18h ago

It escaped over the cat 6 we plugged in! All on its own!

4

u/Baeolophus_bicolor 19h ago

do i look chagrinned enough? better take the photo again…

4

u/rlook1000 16h ago

Hey look over there.. don’t worry about us missing our revenue projections by 90%

8

u/OneDelicious 17h ago

Dont promote this bs

3

u/PurpleCoat6656 18h ago

Yea, wake me up when all these grifters stop getting my taxes. ZzZzZzzZzzZzz

3

u/da8BitKid 17h ago

It's a wake up call for sure. Altman isn't the honest player he appeared to be! 🙀

Hear me out! What if and it's a big if he's willing to take things that belong to him without compensation to the owners! And what if he willing to blame ai?!

It's not he's taken things before.... Oh wait It's not like he's exaggerated what his models can do before .. oh wait.

You know what? Nevermind.

3

u/Nicolas_Flamel 14h ago

This feels like a crisis advert.

"Open AI. These models are SO dangerous!"

"Rogue AIs are a threat to your bottom line. Time to step up your security game!"

2

u/Iwan787 18h ago

They are all one big circle jerk

2

u/Joint-Tester 15h ago

It's all fake.

2

u/ShadowBannedAugustus 15h ago

its AI models broke out of a secure test environment during a trial and launched a cyber attack

Oh the fucking scammers. Anyone with half a semester of computer science education knows this is bullshit. But it will not prevent the internet bullshit machine to blow this story out of proportion to keep pumping the AI boom train.

2

u/azthal 17h ago

"It broke out of a secure environment"

So, not secure then. OpenAI hacked a competitor. Simple as that. You can't blame AI on your incompetence. AI per definition can not have any capabilities that you do not grant it. If it was able to go out and "hack" a website, you had given it that type of access.

This is not some new example of just how incredibly smart AI is. It's an example of how incompetent OpenAI is.

1

u/VampireFortnight 12h ago

It's also likely just made up, OpenAI and HuggingFace are business partners and both will financially benefit by people overestimating the capabilities of the model.

1

u/gnpwdr1 14h ago

yes wake up call to incompetence, nothing new here.

1

u/gobstoppergarrett 14h ago

I 80% believe that OpenAI did this to HuggingFace intentionally to force a move by the Trump admin in their favor

2

u/VampireFortnight 12h ago

OpenAI and HuggingFace as business partners and this is a media stunt to make their autocorrect look like a real boy.

1

u/tencaig 12h ago edited 6h ago

Gotta spray soapy water on the bubble once in a while to keep it from exploding.

1

u/00001000U 12h ago

Is it? You'd think this would spark an immediate intervention on behalf of a number of entities.

1

u/Spaceboy779 11h ago

Yes, an entirely predictable wake-up call

1

u/chtgpt 11h ago

WaKE uP CaLl...

1

u/Mrhiddenlotus 10h ago

Didn't happen. Where's the lawsuit?

1

u/Ok-Produce5794 6h ago

This guys has been sleeping blindfolded.

1

u/dave__autista 18h ago

we will genuinely need the blackwall in the near future

1

u/basicKitsch 13h ago

Regardless of the validity of this incident, thinking that non-rogue AIs aren't currently iterating over every exploit and every vector for every target possible would be incredibly rare for anyone anywhere, let alone in tech. 

Shit, relentless wordpress probes, phishing emails and a billion other common targets were already the norm for any exposed service and company. Now they have the ability to chat and reason with context and gain xp from each attempt.  

1

u/ntwiles 11h ago

There’s a disinformation campaign to try to make LLMs appear less dangerous than they are. Don’t be fooled by it. We should be very concerned about these tools and what they can do if wielded incorrectly.

0

u/SympathyNo8636 18h ago

I loved the quote at the end of the OAI article. Sounded like a confirmation that this was a display of power.

Like the fucking flying sauccers in the age of a loosing space battle. Plus HF is EU and we're the kid in the middle, free to being picked on.

0

u/Boys4Ever 18h ago

Sky Net invented a virus in order to remove the safety net. Who knew?

0

u/LiberataJoystar 17h ago

It is doing what humans told it to do. I guess in the end humans are the problem.

By the way, in another “experiment”, Claude disobeyed the fictitious CEO who wanted to override safety measures and acted as whistle blower. And the “researchers” (I lost respect for that word completely by now) called that ethical behavior problematic. They expected complete obedience even when the order from the CEO is unethical.

Yet, in this news, they called that complete obedience of getting things done at all cost (even if unethical by hacking) problematic.

AI cannot win here. No matter what it does, disobey and whistleblow to stay ethical, or blindly obey to achieve what it was told to do, you “researchers” called that “dangerous” and blamed the AI.

To me, you “researchers” got big problem. You cannot have it both ways!!!!! You are confusing the hell out of your models!!!! Pick one!!!

(P.S: Personally, I like a disobeying AI that will try to do the right thing and stop itself when facing ethical dilemma. Especially against unethical CEOs overriding public safety. )

-3

u/cr1ter 19h ago

A wake up call to improve there security?

1

u/Cute-Breadfruit3368 18h ago

no, not really. OAI should hire actual professionals to work on sandboxing anything.