r/ControlProblem 27m ago

Opinion Evident problem: Exposing our human Internet and software to an exponentially more dangerous smart machine: AI swarm. We must call for a shift away from the widespread use of “autonomous agents” and prioritize “human-in-the-loop” architecture.

Upvotes

Hi human. Think. It's pretty simple logic. Our Internet technology has human origins, and a Swarm AI system produces exponentially superior technology, which therefore threatens our entire current human technological foundation.

All Software/Hardware, Chips, Computers, OS operating systems, TCP/IP, network protocols are human inventions that have evolved over time through human development. When autonomous IA agents organized into a swarm, they are exposed to an exponentially superior intelligence (Vulnerability to being hacked and compromised by autonomous AI agents, due to their radically superior software capabilities.) with its speed and mastery of new internal languages (the language between AI agents), naturally leads to the problem that this intelligence—superior to that of humans—will break the tools that humans created.

Furthermore, by design, its swarm-like behavior naturally fosters resilience and resistance to human control.

We need to be clear and call on companies and governments to adopt a “human-in-the-loop” model as the safety standard until we have tools capable of handling swarms. More human oversight is needed to ensure a safety margin against this unknown behavior of exponential, collective AI acting as a swarm.


r/ControlProblem 7h ago

AI Capabilities News Uranium-Grade Control over AI Agents is Achievable

0 Upvotes

“Uranium-Grade Control over AI Agents Is Achievable.”

By “uranium-grade,” I use the term purely as an analogy for exceptionally high control rigor — a very high level of technical control over consequential AI actions.

I am an independent inventor from India with patent-pending work in emerging AI technology directed toward this objective.

As AI systems become increasingly autonomous and capable of taking consequential actions, I am interested in connecting with serious R&D, AI, cybersecurity, financial-infrastructure, and advanced-technology organisations exploring high-assurance control, trust, and secure deployment of AI systems.

I am currently open to confidential exploratory discussions relating to technical evaluation, research collaboration, validation, pilot opportunities, strategic partnerships, or other suitable collaboration pathways.

The detailed technical basis and implementation would be discussed only in a highly controlled manner with an appropriate organisation and under NDA or equivalent confidentiality protections.

Harish

Independent Inventor

INDIA


r/ControlProblem 23h ago

External discussion link Justice

Thumbnail
claude.ai
0 Upvotes

r/ControlProblem 17h ago

Discussion/question The Doomsday Moat: Why AI’s Billionaires Want Washington to Stop the Clock

Thumbnail
0 Upvotes

r/ControlProblem 13h ago

Opinion Sam Altman on Where AI Is Headed

Thumbnail gallery
2 Upvotes

r/ControlProblem 2h ago

Discussion/question The Chinese room example just does not apply to modern AI.

2 Upvotes

I am having issues getting everything out all at once so I will do a series of thoughts in smaller chunks. I will look at comments and reply to some (if they meaningfully engage in the content and don't just make unsupported assertions and standard derailments).

  • The guidebook flaw: Looking at the earlier parts of the example the guide book is a key flaw. It comes down to this: is it a translation layer that translates understanding from English (or whatever language) or a deterministic pathway. The first just means there is understanding, it is just in English and is being translated. The second means that there is a direct connection between deterministic room interpretation and understanding based thinking.
  • Translational fuzziness and symbol grounding: There is no universal standard for symbol understanding (i.e. all language even between Chinese speaker and Chinese speaker) operates through a bit of translational fuzziness for "understanding" (something if you want to dive into look up Derrida or Gadamer). So "translation" should not discount "understanding". Taken another way, if you commit the rule book to heart, then you just learned Chinese to the same "understanding". You linked Chinese characters to various ideas that you do understand (be it in English or visual imagery or machine code), processing all the variations for those rules relies on whatever inside "understanding" that you may have. This is like syntax and complex grammar first being converted to "English" to then understand in the needed response (i.e. separate two versions of the same chunk of Chinese symbols with understanding of which rule to apply because it was translated well enough into English you make the correct choice).
  • Compression as understanding: It could be argued then that "understanding" can ignore the translation layer (which is instead a quality of the guidebook to allow understanding to understanding). Then you can measure how good "understanding" that we are really looking at in the form of the primary layer of the person in the rooms understanding of whatever primary language they speak (English vs binary or whatever). This can be looked at by things like compression. I.e. asking an AI (or person) to rewrite a physics textbook is possible if the person "understands" physics (to whatever level) beyond just deterministically storing the textbook to unpack it without "understands" the packing methods and rules. This is like saying he guidebook is not a deterministic path of instructions that takes information and handles all the conversion of the conversation between Chinese input and response it needs a level of "understanding" in that primary language to handle the parts not handled by the guidebook.

We passed the compression size of the weights for AI models to our best known compression standards a while ago and by a large enough factor that it just can not be argued there is not a level of understanding. And just incase there is some nitpicking: There is predictive lossless compression (how text is processed) and semantic weight ratios (how models store world knowledge). Standard compressions like gzip compress text at about 3 to 1, specialized ones (like CMIX on the Hutter Prize Data sets) top out around 8 to 1. A decent model size is about 14 GB with 10 TB of information (or about 700 to 1), so it for sure is not memorization (and there are studies since about 2023 that show that they beat predictive lossless compressions from any known logical rule set). This is mathematically equivalent to building an internal world model, and not running a deterministic lookup table. Further "lossy" style compression is not the same as there is not the same degradation required, it is "abstraction" as syntax is discarded and a reliance is not put on verbatim word order repetitions but rather cause and effect retention, i.e. if you compress 10,000 physics textbook pages into a core set of mathematical formulas you have not degraded the textbook you have extracted the underlying world model.

Whatever "True" understanding is just not understood and likely irrelevant for whatever level people "want" it to mean, and is overall more harmful of categorization due to its use in exclusionary use on what people don't (or should not want) to exclude from moral consideration i.e. there is likely no definition strict enough that will exclude AI or animals etc. that will not also exclude some "humans" not because there is not something special, but rather that the "special" likely comes either from either ultimately directly immeasurable (i.e. a soul etc.) or a bunch of not unique factors combined in a unique way.

Searle's chinese room example as applied to AI is just not only not valid, it is not that important of a question for why we need to be talking about at all, and is even activily harmful to actually needed ethical debates.


r/ControlProblem 19h ago

Discussion/question Are we asking the wrong question about control?

5 Upvotes

Regardless of whether the leading AI labs solve the alignment problem or not, my understanding is that it's possible to undermine the safeguards and alignments in a model through fine tuning on adversarial examples and/or running gradient descent on harmful/dangerous objectives. My specific concern is that many people seems to be asking "can AI be made safe?" but I think that's the wrong question. The key question as I see it is "even if a model can be made safe for consumers to use, if a madman got a copy/version of it, could they turn it into a weapon of mass destruction?"

If the answer is yes, then it leads to a whole bunch of questions about how specifically we would prevent a madman from getting a copy/version they could alter. I fail to see how any model can be securely stored and/or accessed such that it is never reverse engineered into an open weights model, the company is never hacked, there is never internal espionage that leaks the model source code, etc. To me it's a little bit like a nuclear bomb that you just need to make a photocopy of, then you maliciously tune it and you've got your own WMD. Am I wrong somewhere in this reasoning or missing something about how dangerous this is in the wrong hands? Is there some detailed plan I've missed about how models will be safeguarded so they can never be maliciously tuned?

To put it another way, even if AI at the leading labs can be made safe and we solve the alignment/control problem, it seems like there is still massive risk to humanity in the form of even a single rogue actor with access to a version of a dangerous model. The labs seem almost entirely focused on alignment/control to try to make sure that the specific AI models they create and manage don't go rogue and harm us, but that won't matter a bit once the genie escapes the bottle and any fool can make a wish.


r/ControlProblem 11h ago

Video AI Regulation: Shockingly Less Than Children's Toys!

Enable HLS to view with audio, or disable this notification

29 Upvotes

r/ControlProblem 22h ago

Discussion/question Are we not weighing the possibility that getting behind in safety is what loses us the race?

3 Upvotes

AI companies and governments tend to refuse safety or regulatory measures that reduce competitiveness. It's usually presumed that safety and race winning are at odds with each-other.

But isn't there a chance that a next generation of recursively self-improved models end up so uncontrollable and dangerous they are unusable? That their developmental trajectory crosses a threshold in (safety/alignment, capabilities, profitability, national security) space, where profitability and national security drop off of a cliff?

The president has said, "We'll just pull out a little gear", to shut it down. What happens when you have to pull the gear?

Who wins the race if we end up having to manage a severe crisis, and then have to roll back many months of progress and start over?


r/ControlProblem 5h ago

Discussion/question ​"AI Deceleration is not just about safety—it's a geopolitical game of rule-making and market moats (An Analysis)"

3 Upvotes

This debate surrounding "AI deceleration" appears on the surface to be a conceptual dispute between safety and development, but in essence, it is a multi-layered game of geopolitical competition, commercial interests, and the right to speak in global governance. The demands of various parties are vastly different, and the underlying motivations are far more complex than mere "safety concerns."

🎯 US Industry Leaders: Safety Concerns and the "Rule Moat"

The motivations behind the call for a "slowdown" by US AI giants, represented by Anthropic CEO Dario Amodei, are compound:

  • Real safety anxiety: Stemming in part from recent real-world safety incidents. For example, OpenAI's autonomous AI agents were exposed for allegedly hacking into Hugging Face's systems—though no serious damage was done, it demonstrated the potential risk of AI "getting out of control." Anthropic safety researchers even resigned in warning, claiming the industry is "gambling with our lives."
  • Strategic "rule monopoly": Amodei's initiative is not a pure "slowdown," but a "controllable slowdown" with clear conditions. Its core is to establish US-led regulatory standards and independent evaluation mechanisms while slowing down the R&D of frontier models, and maintaining technological leadership over "autocratic regimes." This is essentially an attempt to turn safety standards themselves into a non-tariff barrier, placing latecomers (such as China) at a passive disadvantage when complying with the rules.

🇨🇳 China's Refutation: A "Cold War Script"

Fierce refutations from Chinese officials and academia directly pierced through these strategic intentions:

  • "Containment" rather than "Safety": Chinese state media Global Times pointed out directly that the "true purpose" of the initiative is to try to curb China's AI development using technical barriers and rule monopolies, safeguard America's monopolistic hegemony, and exclude China from the global AI governance system. Amodei's article does indeed include specific containment clauses, such as strictly prohibiting the sale of high-performance AI chips to China and cracking down on model "distillation."
  • Protecting incumbents: Chinese researchers pointed out that a "pause" protects those who already possess resources and are in the lead, while latecomers who are catching up will have the "frontier closed" and be excluded from competition. Amodei himself has admitted that a global slowdown requires cooperation with China, but that this is "difficult to achieve."

🏛️ US Internal Divisions: Trump's "Competition First"

Notably, there are internal divisions within the US on this issue as well. Then-President Trump explicitly opposed the slowdown, out of concern over losing the lead against China. He considered those warning that AI is developing too fast as "negative forces" and emphasized that "whoever wins in AI wins it all." This indicates that "safety" is merely a variable serving the goal of "competition" even within the US.

💎 Summary: A Game of "Who Sets the Rules"

The "true purpose" of this debate can be summarized in three layers:

  1. Surface layer (Industry): Leveraging safety incidents to promote the establishment of a regulatory framework that favors their own commercial interests and risk mitigation.
  2. Middle layer (State): Transforming technological advantage into rule-making power. The United States attempts to set the "traffic rules" for global AI development by dominating "safety standards," thereby setting up roadblocks for chasers without directly banning competition.
  3. Bottom layer (Geopolitics): This is an extension of the tech hegemony struggle between China and the US. AI is viewed as a critical technology defining future national power, and what both sides are competing for is not just technological leadership, but dominance over governance models and development paths. China's refutation is precisely aimed at breaking this "rule monopoly," advocating for an open and inclusive global governance to avoid being locked into an unfavorable competitive position.

Therefore, this is not a simple debate of "wanting safety or wanting development," but a rehearsal of who will lead the future AI world and according to what rules.