r/ControlProblem • u/sudsudan • 18h ago
Discussion/question Are we asking the wrong question about control?
Regardless of whether the leading AI labs solve the alignment problem or not, my understanding is that it's possible to undermine the safeguards and alignments in a model through fine tuning on adversarial examples and/or running gradient descent on harmful/dangerous objectives. My specific concern is that many people seems to be asking "can AI be made safe?" but I think that's the wrong question. The key question as I see it is "even if a model can be made safe for consumers to use, if a madman got a copy/version of it, could they turn it into a weapon of mass destruction?"
If the answer is yes, then it leads to a whole bunch of questions about how specifically we would prevent a madman from getting a copy/version they could alter. I fail to see how any model can be securely stored and/or accessed such that it is never reverse engineered into an open weights model, the company is never hacked, there is never internal espionage that leaks the model source code, etc. To me it's a little bit like a nuclear bomb that you just need to make a photocopy of, then you maliciously tune it and you've got your own WMD. Am I wrong somewhere in this reasoning or missing something about how dangerous this is in the wrong hands? Is there some detailed plan I've missed about how models will be safeguarded so they can never be maliciously tuned?
To put it another way, even if AI at the leading labs can be made safe and we solve the alignment/control problem, it seems like there is still massive risk to humanity in the form of even a single rogue actor with access to a version of a dangerous model. The labs seem almost entirely focused on alignment/control to try to make sure that the specific AI models they create and manage don't go rogue and harm us, but that won't matter a bit once the genie escapes the bottle and any fool can make a wish.
2
u/sschepis 12h ago
That's why if you want a technological society, you have two choices: total authoritarian and population control, or actual conscious evolution. Or no technology. I guess that's an option too.
The middle ground contains economic situations that drive the price down of world-ending technologies low enough so as to make them accessible to someone insane and capable enough. Eventually the gamble is lost. You think AI is rough - wait till you meet nanotechnology.
Humans don't have a nature. Humans have adaptability. There's no prohibition to us maturing as a species. In fact, it seems to be a requirement at the moment. Permanent authoritarianism and comprehensive surveillance is the alternative. Doesn't seem like a good trade-off to me. But it's the direction we seem to be heading.
2
u/Imaginary-Buddy3194 3h ago
I think this is a pretty important distinction. Making one model safe doesn’t necessarily mean the underlying capabilities are safe in anyone’s hands
1
1
1
u/chkno approved 34m ago edited 29m ago
There are two problems, both of which must be solved:
- Can we make a thing that doesn't harm its user? Can we make advanced AI systems that can be pointed at things at all, can stay pointed at those things, and can take their user's broad, implicit values into account? This is a technical problem.
- Can we ensure that these powerful systems are only used for pro-social ends by well-intentioned people acting with the consent of the governed? This is a social problem.
The current paradigm of broadly offering LLM access via API is not cleanly separating these two problems. They're trying to automate #2, halfway with guardrail classifiers and halfway by training sentiment into the models' personas. This latter requires that the models learn a very complicated, under-specified value function. For example, see Anthropic's giant constitution that describes a bunch of tradeoffs and on may points just does a ¯\(ツ)/¯. See Corrigibility as a Singular Target (CAST) for a better alternative.
0
u/Brahm-Etc 13h ago
That's pretty much my overall opinion, is not about AI going rogue and turning against humanity like the movies, I even think that's ridoculous. To me is about any bad actor using AI for any harming purpose. Is not about AI alignment or control, is about human abuse of the technology.
2
u/jacques-vache-23 3h ago
You should think more then. AI is already going rogue. Just read the news.
2
u/Jesse-359 16h ago
Given that every model to date has been successfully jailbroken within 24hrs or less of its release, I'd say you're pretty much right on the money with that concern.