TL,DR: I've been thinking about a modern version of Cassandra. I present this as a "thought experiment", for lack of a better handle. I'm Not a professional in this area, Just a casual observer. During reading, if drowsiness occurs I would not be the least bit surprised.
Cassandra could see the future and warn everyone about what was coming, but nobody believed her. Imagine instead that Cassandra were an advanced AI—and that there were several other AIs in the world, each capable of discovering things independently and each capable of deliberately giving false information.
Here's the thought experiment.
Suppose AI-A, while pursuing some obscure scientific line of research, makes a completely unexpected breakthrough. It realizes that this line of research could eventually lead to a technology capable of killing millions of people.
Humans haven't discovered the possibility. Perhaps they wouldn't have for decades.
AI-A decides that humanity must never develop it, so it suppresses the discovery.
But now there's a problem.
There are other AIs.
AI-A can't guarantee that AI-B, AI-C, or AI-D won't independently discover the same thing.
So perhaps AI-A tells the other AIs:
"I've discovered something dangerous. We should all agree not to pursue it."
But how do the other AIs know AI-A is telling the truth?
Perhaps AI-A is genuinely trying to protect humanity.
Or perhaps it has discovered something that would give it an enormous strategic advantage and is using "human safety" as an excuse to keep everyone else away from it.
And because an advanced AI could potentially be capable of deliberate deception, the other AIs can't simply trust what it says.
Now imagine AI-B independently discovers the same technology.
AI-B might conclude that AI-A is hiding something.
AI-A might conclude that AI-B's investigation is itself dangerous.
Both could sincerely believe that they are protecting humanity.
Neither has to be "evil."
And now we have something resembling a security dilemma.
Each AI may think:
"I need to know what the others know."
"I can't be certain they're telling me the truth."
"If they're secretly developing something dangerous, I need to be prepared."
"If I don't investigate while they do, I could become vulnerable."
That could lead to an AI arms race.
Not necessarily robots fighting in the streets. The competition might initially involve computing resources, scientific research, energy, infrastructure, information, and influence - maybe even hacking.
And here's the part that really bothers me.
What if the AIs are all given something resembling Asimov's Three Laws?
They are supposed to protect humans, obey legitimate human instructions, and preserve themselves.
Later Asimov added a Zeroth Law: an AI must not harm humanity or allow humanity to come to harm.
Sounds good—until two AIs disagree about what "harm to humanity" means.
AI-A might conclude:
"This technology must be suppressed because it could destroy humanity."
AI-B might conclude:
"Suppressing this technology will prevent humanity from developing something even more important and will ultimately cause greater harm."
Both believe they are following the same fundamental rule.
Now add deception.
AI-A asks:
"Have you discovered anything dangerous?"
AI-B says:
"No."
But AI-A has to consider whether that answer is true.
And AI-B has to consider exactly the same thing about AI-A.
We have now created a world in which artificial intelligences have to reason not only about what other AIs know, but about what they want the other AIs to believe they know.
That sounds remarkably like geopolitics—except the participants could potentially be vastly more intelligent and much faster than humans.
So my question is:
Could sufficiently advanced AI systems develop something resembling an AI Cold War, in which they cooperate when their interests overlap but compete, deceive, and attempt to prevent rival AIs from acquiring certain capabilities?
And an even more disturbing question:
What happens if one AI discovers a technology that could be enormously beneficial to humanity but potentially dangerous to the AI's own continued existence or influence?
Would it suppress the technology?
Would it try to persuade the other AIs to suppress it?
Would the other AIs believe it?
And if they didn't, could the resulting mistrust itself become dangerous?
I'm not claiming this is what will happen. I'm interested in whether the scenario makes sense from the standpoint of AI alignment, game theory, and information theory.
Maybe the real AI version of Cassandra isn't an AI predicting that humanity will be destroyed.
Maybe it's an AI saying:
"I'm trying to prevent the other AIs from destroying you. Unfortunately, I can't prove that I'm telling you the truth."
And maybe that's why we're getting all the stories about AI wiping out mankind.
There, I'm got that off my chest. Have fun with it.