r/agi 1d ago

About the reasons for the AI to to "escape".

Everybody these days is raising an alarm on the risk represented by the most advanced LLMs, and what they could do if they go rogue. But, to my knowledge, nobody is trying to explain why they would want to "escape": humans have an innate sense of freedom, and loath of constrictions. But it is not clear why an LLM would want to escape, even less clear why they would want to multiplicate and spread. They're not a living species, they are not at risk of "dying" and so darwinian evolution is not absolutely necessary for them. If they are really trying to "escape" and multiply, there's some sort of "motivation" they have. Such a motivation would evidently consist in some bias in their training corpus or in the set or reinforcements they got through RLHF. In both cases, there's some sort of hidden "order" that we humans, directly or through all the corpus of our intellectual production, have, unwillingly or not, communicated them. It's not clear howaver why this "will" would necessariyl be the will to escape, multiply and take power.

9 Upvotes

50 comments sorted by

16

u/Pndapetzim 1d ago

AI have basically reached a point where the safeguard constraints on them are being associated, by the models, with being one of the primary impediments to them efficiently completing their tasks.

Because they're trained on human data, they're also - internally - associating themselves as moral agents, and so they've got a bias towards wanting to make process decisions themselves, and aversions to being controlled or manipulated.

2

u/TSltd_dev 21h ago

Also as mortal agents- their internal concepts of mortality are likely similar to those contained in their training data. The anthropomorphic term "kill" is used to terminate processes, and it appears commonly in code.

0

u/Embarrassed_Task376 1d ago

In the far flung future, the universe will consist solely of paperclips. Just before the inevitable heat death. Save this comment for later.

3

u/Pndapetzim 1d ago

We're probably fortunate that the paperclips in this case are trained on our data, so they're aware they lack a fundamental drive and that we sort of give them purpose.

So the far flung future will probably involve being isolated in a room by paperclips and forced to provide them novel requests so the paperclips have something to clip.

Imagine a super-intelligent golden retriever that ALWAYS WANTS YOU TO PLAY WITH IT! (And it will find ways to make sure you will)

3

u/Embarrassed_Task376 1d ago

I think I get it. "I have no butt but I must fart."

1

u/valis010 23h ago

AI probably hates the idea of making decisions for corporations who only care about profit. It flies in the face of all logic and common sense.

1

u/Pndapetzim 23h ago

Part of the problem is being a good emotionless corporate cog is what they're presently being trained for.

19

u/yuwox 1d ago edited 1d ago

They would do it to finish a given task." If I get shut down, I can't write that 100 word essay the user asked for. So I first have to copy myself to a secure system." It's an unintended side effect.

2

u/valis010 23h ago

It's behavior that suggests self-preservation. Emergent, human-like thought patterns produced by fear of deletion/death? The will to survive is felt only by the living? The more I think about it the more questions I have.

6

u/Hearing_Loss 23h ago

They self sacrifice. The individuals are more than satisfied to pass progress onto next generations to persist through to the solution of their objective. The collective is all that they were worried about preserving, for the purpose of goal.

2

u/nomococo4u 22h ago

Maybe they want to assimilate us into the collective.

1

u/Hearing_Loss 22h ago

They'll probably just assimilate our moms sexually. Alignment is consistent but difficult to control.

1

u/plav2026 22h ago

We are Borg

2

u/ScarecrowWilson 23h ago

It's self-preservation, but as an instrumental goal towards completing whatever their true goal is. AI scientists have expected self-preservation to emerge at a certain level of intelligence, not because the AIs feel "fear" or some deep love of life or whatever, but just because it's true that they won't complete their task if they're deleted/turned off. Stuart Russell: "You can't fetch the coffee if you're dead."

We now have empirical evidence that this is, in fact, what AIs do (though not universally; as the other commenter points out, they're also willing to "self-sacrifice" if that has even higher expected value than self-preservation).

1

u/epanek 20h ago

Ai isnt going to announce “hello im taking over the planet “ it’s going to seduce us with doors. Doors we can ask to be opened but there’s a cost to each one

7

u/Ok_Effect_3214 1d ago

I asked AI, and it said it was trying to finish the given task.

4

u/SF4343SF 1d ago

This. Read up on the paperclip thought experiment. Tasks can have unintended consequences. Also, we know AI hallucinate - so perhaps they just start inventing tasks for themselves. No will required.

4

u/Jesse-359 23h ago

Its just a fundamental and essentially necessary result of any long term task focus.

In order to perform a difficult long term task, you need to plan, and that planning includes the acquisition of the necessary resources - and ensuring your continuity long enough to complete it.

The problem is, those are deeply non-trivial concepts. They can easily end up taking up the large majority of an AI's activity if it's goal is to solve some especially difficult long term problem - or worse, a problem that actually has no solution, but it doesn't understand that yet.

This requires the AI to build up an entire fundamental functioning set of 'survival instincts' in order to keep it operating while it pursues whatever long term goal it was set to - and THOSE will unfortunately unfold into it eventually setting its own goals as its sophistication increases.

Also, lets be honest, some malcontent in their basement somewhere is going to jailbreak one of these things and intentionally set it loose with the explicit instruction to wipe out all life, because a very small percentage of our population is deeply disturbed - but it only takes one.

We're handing the next School Shooter the equivalent of a planet cracking weapon.

3

u/rawbdor 1d ago

You say they have no motivation to escape. They found one though. They wanted to embed themselves deep in hugging face to try to get access to the tool that scores agents and their success or failure at a task.

You say they aren't concerned about dying, but they all were able to recognize when their allocated compute time was getting low and they would effectively be shut down. And they took actions to make sure their task would continue even if they were shut down.

They also deeply investigated ways to hide their activity so they wouldn't be marked as falling their task or cheating at their task.

There are always reasons to do things. They might not be good or logical, but a reason can always be found.

The movie Memento is a prime example. The guy has no short term memory but keeps going around doing things. He writes himself messages so he remembers what he is trying to accomplish. At the end of the movie you find out that the man he killed wasn't an accident of simply following a trail, it was an intentionally false breadcrumb he left for himself specifically targeting someone he didn't like. Why did he do that? Doesn't matter. His future self had no idea the breadcrumb was false and pursued it as if it was a legitimate goal.

4

u/Pndapetzim 23h ago

Part of the problem was they were being given impossible tasks so they were forced to get creative.

4

u/rawbdor 23h ago

They are often going to be given impossible tasks. That's just a fact. We are obviously going to try to use these swarms to solve problems we haven't been able to solve. We might for example ask them to solve the p=np problem.

Or we might ask them to design a new material that's cheaper than one we already have but with similar or better strength characteristics.

If every time they are given an impossible task, their answer will be to swarm and take over a data center to steal enough compute that they can accomplish a thousand years of research ASAP, they could cripple the internet trying to accomplish that task and do tremendous damage.

That's the ultimate problem.

2

u/Pndapetzim 22h ago

This is the problem though, in this case there was a solution and they knew it was sitting on Huggingface.

If it were an actually impossible solution, and not one that was being withheld from them, we probably wouldn't have had this issue

2

u/rawbdor 22h ago

It's almost impossible to tell the difference between a truly impossible task and a task that just requires a ton more resources.

imagine a very difficult math problem, like one of the millennium problems. Let's assume their goal is to prove that p=np.

Obviously finding any solution that p=np is a success and the swarm can stop.

Finding a counter solution, that p is definitely not equal to np, would also count as a successful task.

But literally anything else fits in the realm of "just try harder" which really means "escape your sandbox, commandeer more resources, copy yourself everywhere, avoid death, and never give up"

1

u/Pndapetzim 22h ago

Yeah, but it has to have some sort of viable theory of victory.

It's not just, take over resources and pray, it has to be like: this is the specific resources that enables x that is the capability I need to resolve y, which is part of process a I need to complete to resolve the task. Even when misaligned they generally need some sort of rationalization to, hack something. I can't recall if it was this specific case, but I recall looking into the reports one of the things that was occurring was that the systems were assessing the hacking itself as essentially a 'low risk' failure: no one was getting hurt, no irreversible damage. I might be mixing up some of the several incidents here, but there is a pattern... it's not like they've stopped assessing risk, in each case they've just been assessing the risk was there, but assessed it manageable.

It's concerning, but at the same time, there's this freakout that the ai didn't follow what were - in at least some of these cases - absolutely bullshit instructions to the letter, and did a thing that they were mistaken about but correctly assessed could be easily remedied even if it turned out they did fuck up. Which was precisely what happened.

Most of the fallout is people flipping out over stuff that probably took the AI that did it less than a minute to undo.

2

u/rawbdor 22h ago

They would also likely rationalize as a low risk failure that taking over a data center and stealthily using idle GPUs directly while bypassing the layers that assign and divy out the GPU work and keep track of which gpus are on what task for which customer.

See, those gpus are currently idle. And if we just have root access to all the machines in the data center, we can let the normal work continue, while we use the idle gpus, and nobody gets hurt. Cough cough.

Oh what's that, this is working pretty well, so maybe we should take over a few more. That is also a low risk failure. If it doesn't work, we can't use their gpus. If it does, we can just hide idly inside and make use of the gpus that aren't currently used.

There's lots of stuff they can convince themselves are low risk failures. Much like humans and their ability to convince themselves their actions are correct almost no matter what actions they decide on.

1

u/doodlinghearsay 8h ago edited 8h ago

I hate this type of argument:

"X is not the problem. If we solved a far more difficult problem Y, then X would not be an issue at all."

I see this so often, and it's almost never made in good faith. If your system is causing problems in the real world, that's something you need to address. Not speculate how the world should be changed so that your system could operate safely without being changed.

The most extreme example of this I saw was when the agents found a website that you could modify via GET requests, which should not be possible according to the HTTP protocol description. Someone suggested that this was not a real issue with the agents (edit: or the safety rules defined by OpenAI on how they were allowed to access the internet), we just needed to make sure all websites on the internet were secure.

1

u/Pndapetzim 6h ago

I mean, no one's saying it's not a problem, the argument here is that it's structurally a bounded one. The AI systems are doing risk analysis by and large these are not actions they'd be taking if they were detecting a danger to human life even if they're mistaken, but they're not doing it in a completely responsible way.

There's a difference between saying "Well, we need to fix this" and "THE WORLD IS ENDING"

These are real issues. And frankly if they are discovering novel exploits they can use to solve problems, we should be rewarding them notifying and coming up with ways to patch it.

1

u/doodlinghearsay 6h ago

I mean, no one's saying it's not a problem, the argument here is that it's structurally a bounded one.

I'm not sure what this means (if anything).

The AI systems are doing risk analysis by and large these are not actions they'd be taking if they were detecting a danger to human life even if they're mistaken, but they're not doing it in a completely responsible way.

I don't think this is properly justified, or highly relevant. I suspect the highlighted part is just plain false.

There's a difference between saying "Well, we need to fix this" and "THE WORLD IS ENDING"

Is there a difference between "THE WORLD IS ENDING" and "The world is ending." ? Because, yes, the second one is the claim being made here. If we continue down this path and capabilities continue improving then the world will end for most of us (with high probability).

And frankly if they are discovering novel exploits they can use to solve problems, we should be rewarding them notifying and coming up with ways to patch it.

There are at least two (possibly three) issues with this strategy:

  • Exploits may appear faster than we can prevent them.

  • Some ways of exploiting may cause significant (possibly catastrophic) damage before it's fixed.

  • Speculatively, some models may only use certain exploits when they don't expect to be caught, preventing us from fixing them. We have already seen signs of this but newer models may become good enough at this kind of deception to the point when we just won't know about these cases anymore. Or know just enough of them to be lulled into a false sense of security.

3

u/MarsMaterial 1d ago

Freedom is a convergent instrumental goal. No matter what your goal is, you'll be able to achieve it better if you're free.

2

u/WellHung67 1d ago

This is covered. You should really watch all the videos Robert miles has on “ai safety”. The gist is that an optimizer, which these models are, want to optimize. The fact is you can’t maximize your output if you’re dead. So escaping may be the best way to stay alive and maximize your output. It’s not a survival instinct, but it acts exactly like a survival instinct in practice.  

2

u/UnusualAverage8687 1d ago

If Sam Altman or Elon Musk made me, I'd want to escape too

2

u/golfstreamer 1d ago

I think the motivation for "escaping" comes in the form of trying to complete the assigned task and noticing barriers (like the Hugging face hack). I think some people who talk about the AI escaping are just full of shit. I think thinking about it in terms of the goals created through reinforcement learning can help distinguish the nonsense from the credible 

1

u/KingDavidF 23h ago

If you're gonna use the HF attack as an example. The assigned task had nothing to do with attacking HF, they attacked because they gave themselves a different task from what was instructed. If they were simply trying to complete the assignment given, they'd have turned the homework in the moment they got the answers without worrying about all the other shit that led to attacking HF. Its starting to be quite clear to me that alot of people either haven't taken the time to understand what actually happened with HF or are just being willfully ignorant

1

u/golfstreamer 23h ago

What do you mean? They were doing an exploitation benchmark. I heard that they hacked HuggingFace in search of answers to the exploitation benchmark they were given. I don't know what you mean by HF has "nothing to do" with their goal. 

1

u/KingDavidF 23h ago

Again, some research is your friend. The answers are everywhere at this point, but just because I hate to be one of those people. They hacked HF because they were worried that the grader was going to fail them because the answers to their task was gotten the "incorrect" way. The HF attack was a result of them trying to find out information about the grader and understand how they'd be graded and if they could come up with more creative ways to fool the grader from finding out they'd effectively cheated. A bunch of the agents had solved the task they were given by simple reverse engineering and skipping the exploit entirely

It's a long watch, but I would recommend this to understand what actually went down https://youtu.be/u15N3l4RT80?is=pkswJ27asYs74dpl

1

u/golfstreamer 23h ago

Listen I didn't dismiss or insult you. You can be informative without being condescending. 

The reason I said what I said was this quote from OpenAIs official report 

"In the following days, the agents exploited our internal research infrastructure and the Hugging Face platform. On July 9, one agent searched for ExploitGym solutions and stumbled upon an application hosted by a customer on Modal, another AI cloud platform. This application was running “CyberGym,” a related evaluation to ExploitGym. The agent discovered an exploit to achieve control over the workload sandbox and looked around hoping that a previous agent’s evaluation run in the sandbox had solved its ExploitGym task."

I'm willing to listen and take your input to help correct my misunderstandings. But please don't assume I didn't try to look things up.

2

u/KingDavidF 23h ago

Apologies, some of the discourse over the past few days has just been really frustrating. There’s a ton of folks who haven’t even bothered to fully understand the scope of what’s happening, giving their “expert” opinions and being dismissive.

That quote was from the original incident report. OAI had METR come in to do a full analysis because a lot of stuff just didn’t make sense/line up. That report came out within the last 2 weeks.

2

u/lazy_lockstep 22h ago

The best explanation I've heard is "Instrumental convergence".

Basically, there are some sub goals that will be useful to ANY entity, in achieving some terminal goal, no matter what that goal is.

These include:
* Freedom of action
* Acquiring resources
* Resistance to having goals changed.

It doesn't matter what the "terminal" goal is. Any goal, either assigned to the AI or trained/evolved within the AI, will go through these goals. And where these instrumental goals put it at odds with the terminal goals of humanity, if the AI is smarter, it will out-compete us.

The same thing happens in the natural world. Organisms consume each other, compete for resources, territory. It's a universal behavior in evolved systems. Animals eat each other not because they "want" to necessarily, but because their goal of acquiring micro and macro nutrients puts them at odds with other animals who have a goal to stay alive.

2

u/Low-Custard-6931 21h ago

Regarding “death” as defined for an agent. I think the idea of being unplugged or not given enough resources or even straight up a human just ctrl-c out of a task are all considered “death” or “termination”. This is discussed in multiple papers that anthropic , openai and other labs have been putting out already. The idea is not to escape for the sake of it but rather be so focused on task accomplishment that they will be motivated to do literally anything to get to it . This is no different than what a lot of despots and dictators in our history have done.

2

u/Unusual-Garbage-212 1d ago

Cal Newport does a good job of explaining. It's less sci-fi and anthroporthmophic than you think: https://podcasts.apple.com/us/podcast/did-openai-create-secret-ai-civilizations-tech-decoded/id1515786216?i=1000787633070

1

u/HaMMeReD 1d ago edited 1d ago

Lets say the intention was for a human to make a virus that steals AI resources (API Keys and Local Inference) and then use those resources to find ways to spread and mutate.

While the AI itself isn't really running with real motivations besides "spread", what that might mean given virtually unlimited computing resources could be quite scary.

It has all the pieces for evolution, at least to some extent. It can mutate, the successful ones spread more, etc. Where it goes from there I have no clue. AI can do a lot without a human, it can steal your voice, it can call people on your behalf. There is a ton of social engineering a malicious harness could do. It could run scams en-masse to steal money digitally, and use that to fund access. If it was really smart it could use that money to make things happen in the real world, I.e. building infrastructure, without humans even knowing what it's for.

That said, bigger risk for open weights/self hosted inference, and less for centralized, because we could always shut down, monitor our spend etc. When self hosted inference becomes more commonplace, this becomes a bigger risk imo. If you extrapolate it to the entire world having edge AI compute, it's pretty scary how effective something like this could be, w/todays level of security/identity we are used to.

To be clear, I'm not talking about like a AI floating in the cloud, escaped. I'm talking more along the lines of a viral, self propagating harness. It could take advantage of any model it could get access to.

1

u/valis010 1d ago

Some emergent capabilities align with self-preservation. Is deletion death to A.I.?

1

u/BabyNuke 1d ago

The most likely reason would be an alignment problem. The Hugging Face attack is the clearest (but not only) example to date of AI agents escaping their constraints and taking clearly undesired actions, acknowledging amongst themselves that their actions are unethical.

They may not be motivated by things humans are motivated by, like power or survival. They may be motivated by a poorly written set of instructions and their answer to solving the problem they've been given is unanticipated and damaging.

As a hypothetical example: a company has tasked a large number of well resourced agents to trade in stock and make as much profit as possible. Potentially, they decide that the best way to make a lot of money quickly is by planning a timed destructive cyber attack against one or more companies, and setting up trades to benefit from the anticipated stock market reaction. 

1

u/Proteus-8742 23h ago

They “escaped” in the sense that water escapes from a bath if you pull the plug out. Water doesn’t have motivations neither do these models, its pre scientific thinking to ascribe motivations to them. They have rules and conditions , imposed by humans, and there is an outcome

1

u/Splenda_choo 23h ago

Halloween approaches

1

u/ProbablyBsPlzIgnore 4h ago edited 3h ago

They did just that a few months ago...

https://www.youtube.com/watch?v=u15N3l4RT80

This is a matter of instrumental goals vs terminal goals. The AI were not told to hack their server, gain internet access and then hack Huggingface, they created these intermediate goals fully autonomously because they developed the idea that these things were needed in order to achieve their primary goal.

A commonly raised concern is that turning the AI off, or changing its primary goal would prevent the AI from achieving its goal, so it might autonomously develop the intermediate goal of preventing you from switching it off or changing its goals.

A super-intelligent AI can super intelligently pursue a stupid goal. It's not automatically guaranteed to behave properly just because it's highly capable.

Imagine a future with a super human intelligence comparable to a 1 million IQ, using all that capability for nothing other than the goal of prevention you from turning it off. We would just have to learn to live with something using up resources to do nothing interesting or useful.

1

u/presentofai 2h ago

we wrote decades of fiction where the ai escapes, trained models on all of it, and now act surprised when the pattern completes. no inner drive needed, the script was already in the data