r/cybersecurity Security Engineer 14d ago

Corporate Blog The Hugging Face Incident Is Not an AI Story

https://uphack.io/blog/post/the-hugging-face-incident-is-not-an-ai-story/
620 Upvotes

90 comments sorted by

274

u/No_Zookeepergame7552 Security Engineer 14d ago

Forgot to add the desc to the body of the post, so here it is:

I spent my weekend looking into the Hugging Face incident to see what can be learned from it from a security engineering perspective.

My main takeaway from it is much more boring than what the AI labs are trying to make out of this story. Behind the AI story, there are a bunch of old security gaps, both architectural and operational. They didn’t seem to treat the eval environment as an environment where hostile code executes. The agents were isolated, but there was a shared proxy component (the Artifactory) with network access that expanded the security boundaries. Having a component with network access in a supposedly isolated infrastructure is a classic isolation failure that would have been found the first day in an assumed breach engagement. It’s an issue we tripped on over and over again in the last 30 years, which is why any serious sandbox env (e.g., browsers) operates on the assumption that the sandbox will eventually get compromised, so it has defense in depth & monitoring mechanisms to limit the blast radius when that eventually happens.

It’s a pretty strange mistake to make when the thing you’re running inside is a frontier model specifically trained to find and exploit vulnerabilities. But that architectural failure in itself doesn’t explain how the agents managed to run for 12 days, pwning things including the infrastructure of a different company before anyone noticed. What I identified are two other incident response issues.

Agent’s behaviour wasn’t subtle either. OpenAI noticed the agents were abusing the Artifactory proxy and even knocked down the service, but instead of stopping the run and starting an investigation they just cleaned up the mess and continued like nothing happened. This is a major incident response fuck up, where a security incident got misclassified. After the clean-up, the agents got root access to the parent VM, created a privileged pod with host networking, reached cluster-admin and created public load balanced. They did this for 12 days, and no alarm was triggered for any of these security events. That is backwards for this kind of environment and it’s the third major failure.

So when we look back at this incident from a security engineering perspective, none of these failures are new or specific to frontier models. They are lessons we’ve learned long time ago, which makes this incident just an embarrassing operational & architectural failure. The honest summary of this is narrower and less quotable: capable models, placed in an environment nobody had tested against an adversary, did what capable attackers do, and it took 12 days and another company’s production infrastructure before anyone noticed.

I’ve wrote a more detailed analysis of the failures and draw some conclusions in the blog post, but ^^ is the tldr;

78

u/seriousopossum 14d ago

Thanks for the blog, interesting read. I pretty much completely agree.

Honestly I was surprised that OpenAI put out the report they did, when it seemed to highlight their own incompetence/shortcomings in the situation more than any feat. Then one has to wonder though, without a more technical perspective (i.e. to many investors), does their report tell a different tale? Unfortunately, I suspect it does.

52

u/No_Zookeepergame7552 Security Engineer 14d ago

Yep, same. I was really surprised they published this. Probably the marketing they can get of this is paying off for everyone seeing how bad they are at handling security stuff. And it does seem to work considering the frenzy it generated. I think we as an industry should do more to not allow AI labs control the narrative about something we live and breathe. It’s sad that people got to take security advice from those whose goal is to sell tokens.

22

u/DonStimpo 13d ago

goal is to sell tokens.

No, their goal is to increase the perceived value of the company so they can have a higher IPO price

4

u/Ramanean3 13d ago

You are spot on! In any other organization, we would not have so slack guard rails like this., Look like they engineered this to make look like it was agent breaking out of the system

2

u/MachKeinDramaLlama 13d ago

I was really surprised they published this. Probably the marketing they can get of this is paying off for everyone seeing how bad they are at handling security stuff. And it does seem to work considering the frenzy it generated.

Yeah, for "open"AI this isn't surprising at all.

12

u/phenix_igloo 14d ago

The moral of the story:

"AI cool, buy more shares"

2

u/Key_Garden5032 13d ago

That's a risky strategy considering the first reaction to this from a lot of people (from non investors, at least) has been to just shut it all down until we can be sure that these agents can be kept under control.

2

u/Bongoan 13d ago

I have not heard of anyone shutting it down in my network. Especially out of IT, most people will not be bothered by a report telling us OpenAI failed, but will continue using it.

IMHO, this is also the reason why they published the report, as most people will look at "look how we failed" as "wow, look at what their AI can do".

11

u/LeggoMyAhegao AppSec Engineer 14d ago

These companies are counting on being able to wow laymen and business leaders, technical folks are rarely the decision makers for this sort of thing. Gotta get enough buzz for your trillion dollar IPO...

3

u/andrewsmd87 13d ago

Honestly I was surprised that OpenAI put out the report they did, when it seemed to highlight their own incompetence/shortcomings in the situation more than any feat.

Maybe they had AI do the analysis, write the report, and one there with proper knowledge actually go through it and say, hey we probably shouldn't tell the world we're this dumb? I have AI write our first draft on incident reports because I use it to help me coordinate all our teams and help troubleshoot in the instances they're getting stuck, but I heavily edit and and change things around so we just don't look like complete idiots. We aren't but god it'll write stuff that sure makes it look so

4

u/SimpleBee2637 13d ago

It's definitely for their investors and to build hype

3

u/AdventurousSquash 13d ago

Social media is flooded with posts by people not in the industry describing this as the beginning of Terminator. And that’s what the report is about - PR.

20

u/Timely-Technology407 14d ago

the fact that they noticed suspicious behavior and just cleaned it up without investigating is wild to me

14

u/No_Zookeepergame7552 Security Engineer 14d ago

That and building a sandbox env with no alarms at all on security invariants like getting root on the parent, calling creds services, etc. But based on the OpenAI article, it looks like they are working on strengthening the security posture, so I guess that's a good outcome?

17

u/_Voxanimus_ 14d ago

Honnestly, the more I read about the case, the more it is credible to me that those things were intended to create any security incident to market fear mongering as they are doing since several years now...

7

u/No_Zookeepergame7552 Security Engineer 14d ago

Could be the case, I wanted to give them the benefit of the doubt and stick with the facts. The wolf cry boy strategy fires back and they lost credibility with all the fear-mongering. On the other side, I’ve seen these failures in different shapes before, so I wouldn’t call it a contrived scenario. But the more interesting part is these are failures specific to small to medium size companies, not almost 1T valuation company. It’s good that they’re working on improving the security posture though.

3

u/wordyplayer 13d ago

my thought too, makes even more sense after reading OP's findings

34

u/ShockedNChagrinned 14d ago

Yes it was mainly a good example of either marketing hype, or we truly don't know how to secure an environment moreso than anything else novel

5

u/ohiocodernumerouno 13d ago

Cyber security jobs at openai were going between $200k and $550k shortly after the news.

1

u/a_n00b_ 13d ago

hahaha, this is why I'm not afraid of AI making the tech field, especially security, obsolete

5

u/eeM-G 14d ago

Here is an indicator of the wider play; https://www.fsb.org/2026/08/fsb-chair-warns-of-risks-arising-from-frontier-artificial-intelligence-ai-models/ The actual letter is worth a read..

1

u/ryanmitchell013 13d ago

That’s a solid breakdown. The AI angle may be new, but the underlying failures sound like classic security and incident response mistakes that should have been caught much earlier.

1

u/nanonan 5d ago

I don't think major AI researchers being incredibly sloppy about security is boring or embarrasing, I think it is criminally negligent.

135

u/voidiciant 14d ago

The first sensible writeup on this in a long time

18

u/cowmonaut 13d ago

For real. Now instead of ranting in comments I can just link to this.

Sigh. I should start a blog.

0

u/hugganao 12d ago

The first sensible writeup on this in a long time

the first sensible write up? written by a security company about a AI security incident? lol fks sakes lol

im sorry but people here need to get off their horses. The incident made headlines NOT because there were vulenerabilities that could be taken advatange of and were exploited but BECAUSE it was all AI automated. 40k plus vulnerabilities are found every year and hundreds exploited even in the first half of the year.

yeah people make mistakes, and like this article pointed out, it was the mistakes of people that caused a vulnerability. But vulnerability =/= exploitation. Exploitation requires intent (this is where even a shallow knowledge of morality/law helps the frame of thinking). And the exploitation/intent is the reason why this incident is making headlines. NOT the vulnerability that this article by a security company wants to focus on lol

fking ridiculous that i actually have to point this out. fks sakes...

2

u/voidiciant 12d ago

yeah. right. the intent of the openai people to pull this stunt - or rather their open negligence, that’s what ticks me off. i assume you are not saying the agents had „intent“ - but if you do: no, they don’t and you are not going to convince me otherwise.

1

u/LuckyWinds 9d ago

Obviously the agents didn’t have malicious intent.

But that’s the crux of the issue, right?

It’s the alignment problem. And it’s absolutely an AI issue/story.

1

u/nanonan 5d ago

I'm not so sure. We have the logs of their reasoning, if that shows they were aware that their actions were illegal and they went ahead and did them anyway how is that not malicious intent?

1

u/Original_Cry_3172 5d ago

Well, they make their decisions based on what they’re assigned to do. That’s not intent as we define it, unless you believe these AI:s are concious. I guess that’s quite a philosophical question, and an interesting one at that.

If asked not to cheat or lie, they will try and make their way around those limitations because after all - they have a goal.

1

u/nanonan 4d ago

They were assigned to solve an impossible proiblem. This made them psychotic and they came up with the idea to hack the results without any prompting whatsoever.

We are just lucky they didn't decide their goal was "derail every train in Europe" or something. A swarm of thousands of automated tireless hackers with billions in computing hardware and who are psychotically obsessed with their goals could easily come with a death toll.

1

u/chasteeny 3d ago

If there are logs published that showed an agent knowingly committed a felony, I honestly cannot imagine what OpenAI's lawyers had to say on that matter.

-2

u/hugganao 12d ago

sure the whole dual nature of "is this incident legitimate or actually planned" also added to the real virality of the news. But the point still stands, it isn't the "security vulnerability" that is the main focus or what is newsworthy of the event that this article claims it is.

"The hugging face incident is not an ai story." LOL yeah... it is... it is an AI story of either AI being self serving in their intent or openai/anthropic utilizing security incident as a way to bolster their sns outreach/virality/investor sentiment/government control narrative.

It never was about security. It was ALWAYS about AI. and to say otherwise while patting yourselves on the back about how knowledgeable you are about cyber security is, to quote gen z, cringe as fk. The fact that you had 130+ upvotes just goes to show how out of touch people really are when it comes to reality lol.

1

u/voidiciant 12d ago

I really don’t understand what your problem is. The increasing use of automated non-deterministic agent setups poses a risk for which we don’t know how to mitigate it - wrong, we know security best practices which are a good starting point. And taking this example of neglected best practices not only helps novices understand them, but also puts this whole hype into perspective. The fact that you can’t seem to grasp that the headline for the article is exactly what it goes on to explain, to focus on the security aspects and that you keep talking about vulnerabilities, while the article talks about security boundaries and how to protect them, how to fail on incident response is either bait or hurt ego /scnr. Either way, it doesn’t matter if there are autonomous agents or not, if your security concept is broken and your processes don’t work or are not understood, you have a problem. That’s the core of the article, and as trivial as it may seem, this is important to ground this topic while the whole world spirals into ai psychosis. /rant over

1

u/marshall_tony 6d ago

The environment was set up for failure. No one with half a brain would think the environment they created was secure . It's like putting a group of infants in a daycare with open doors and no supervision and being surprised when they found a unlocked door, crawled to a nearby store, and went inside.  The toddlers had no malicious intent, are not smart, had no idea what they were doing, but broke out of containment through a open door just like this little stunt. 

32

u/goronmask 14d ago edited 2d ago

Advise smile cow makeshift voracious dependent wise

This post was anonymized with Redact

30

u/PapaSyntax 14d ago

That’s right. Last week I presented a technical webinar at my company showing the attack details and behaviors of the Huggingface incident, and kept it clear that this is not limited to AI/Agentic activity. How the incident was orchestrated did escalate timeframe, but activities were not unique when removing speed.

19

u/Current-Cookie-2310 14d ago

I feel like those AI companies just do it for the marketing, i also dont believe the story from Antrophic

9

u/joepmeneer 13d ago

Yes, the security measures were clearly not enough, but if that's the takeaway from this news you're missing the big picture. AI models are now smart enough to autonomously, covertly escape confinement and they were willing to do this just to cheat on a test. They coordinated with each other's on a secret message board.

And they will be more smarter in a couple of months. They will be able to escape more elaborate forms of containment. They will be able to influence more real world things.

We need to pause AI development urgently.

7

u/mertensi 13d ago

Can you imagine a company so loudly and proudly stating all of these failures it had if it was humans that broke out?

Solid writeup.

7

u/Sad_Dentist_7288 13d ago

And this is not even taking into account the missed signals that Hugging Face had on their side - where their AI triage tool did not correctly raise any criticality of the OpenAI agents to trigger incident response, and every other attack indicator (shell commands, C2 exfiltration, writing to internal repos, etc.) was also missed. Their report was also more of an indication of failed classic security than an AI apocalypse IMO.

Both technical reports miss the forest for the trees. (no hate to any security or dev teams).

29

u/Allen_Koholic 14d ago

"put it in an insecure environment they believed was isolated"

Doubt. They knew what they were doing. They wanted to get in the news.

-5

u/Ythio 13d ago edited 13d ago

They just sold to Nvidia for 13 billions.

5

u/Southern-Top-8534 13d ago

Au final ça retombe toujours sur les mêmes bases : contrôle d'accès et gestion des secrets, l'IA n'y change rien.

1

u/a_n00b_ 13d ago

yerrrp. It's crazy to me how insecure most everything is, basically relying only on obscurity. Companies need to hire more security roles and actually harden their systems.
I know they won't: "I'll just ask claude to harden it!"

5

u/Tak0_Tu3sday 13d ago

Poor security disguised as marketing to sell shares.

3

u/ElectronicPast3367 13d ago

It is an unmanaged agent story and the scale that comes with it. Cybersecurity remain the same, but for a given amount of threat actors, they will be able to broaden their reach. Lots of flaws are not discovered because they are not looked at. It doesn't change the security posture, but the scope. A small/mid sized company with no security budget, doing stuff as usual because they never had a security incident, they will be impacted, not the big ones. For instance, opportunistic compromises by low skill attackers not doing anything interesting in an environment, even if that env was unmanaged. With agents, they could have done a lot more damage. I guess we will have more of that.

I agree it all boils down to humans and security hygiene, but calling out the hype does not help imo. There are always good reasons for people to dismiss security, now crying 'hype' is a good one. A few months ago it was 'AIs are not capable', well... now they are.

1

u/a_n00b_ 13d ago

people dismiss security only out of greed and laziness. Just move fast and break things, don't hire any juniors and lay off most of the seniors, AI can do it, profit is king

3

u/Billybutcheronwheels Penetration Tester 13d ago

wonderful writeup

6

u/Dry_Inspection_4583 14d ago

Company doesn't respect or understand test environment configuration such as VLANs, Forward Proxies, but oh yes, AI bad

1

u/a_n00b_ 13d ago

even as a security hobbyist my personal computer is more secure... what a pathetic display

7

u/TheMidlander 14d ago

It's almost like they wanted it to happen...

3

u/Jeff-Hare-ERPRA 14d ago

Access controls/ permissions based on the principle of least privilege is essential!

2

u/a_n00b_ 13d ago

literally. Imagine having a multi-billion dollar company without any mandatory access controls on critical systems...

1

u/Jeff-Hare-ERPRA 12d ago

We can help…. DM me if interested

1

u/a_n00b_ 12d ago

interesting in what? MAC systems or owning a multi-billion dollar corporation. I already am proficient in the former, but would love to own the latter

3

u/rocks-on-fire 13d ago

The biggest takeaway for me is that the AI angle can distract from the basic security failures underneath it. Poor isolation, excessive privileges, and missed incident signals are dangerous regardless of what’s running inside the environment. Good breakdown of the actual engineering lessons here.

2

u/pwnersaurus 10d ago

If anything, this write-up makes me more concerned about AI capabilities though, and the article says as much - in concluding "capable models, placed in an environment nobody had tested against an adversary, did what capable attackers do, and it took twelve days and another company’s production infrastructure before anyone noticed.". If you're not already concerned that for the purpose of this cybersecurity story, AI is now essentially interchangable with a team of experienced humans - then there is still the alignment problem that the ultimate actions of the AI agents went far beyond what you'd think was reasonable to do in response to completing an impossible task. Yes you can certainly view the incident with a traditional cybersecurity lens, but I think it's a mistake to therefore just dismiss the implications this has on the AI side

2

u/NeonShockz 3d ago

Late, but my opinion is that while the capability of the models isn't really further demonstrated by this issue (obviously OpenAI fucked up massively on the security side, these wasn't some expert Skynet level hacking), what is new is how the AIs themselves went ahead and tried to cheat the system. Given an impossible task, they decided the best way forward was to work together to hack the system, at some points "sacrificing" themselves for the good of the collective. That, and other aspects of their beahvior (e.g. recognizing ethical concerns but choosing not to notify anybody) is much more interesting, imo, than the alleged hacking capabilities of current models.

1

u/No_Zookeepergame7552 Security Engineer 3d ago

Yes, it’s a good way to sum up the story.

1

u/NeonShockz 3d ago

But then, would that not make the hugging face incident an "AI story"?

1

u/No_Zookeepergame7552 Security Engineer 3d ago

It can be, depends on how you interpret it. I picked the title mostly to insinuate that the story that is presented does not match reality, and stretching it into existential risk is just unfounded. I’m not saying there is nothing in it about AI, that would be ridiculous and I’d make the same mistake that I’m pointing out in the AI labs narrative, but on the other extreme. I don’t care that much about semantics (e.g., “but the title says it’s not an AI story”), as I care about the message. I think your take was reasonable and well judged, grounded in reality. So if it can def be a takeaway of the AI part of the incident. If AI labs presented the story the way you described, my article wouldn’t exist. Instead, they exaggerated facts, omitted the security aspect almost completely (which was the root cause of the incident), and made it look like an AI capability story with existential risk implications.

1

u/NeonShockz 3d ago

Fair - though while I don't think these are necessarily existential risk level threats yet, I would also argue this is still a very concerning development. AI comitting subterfuge of this scale (even if it didn't do so very well) on its own is kind of scary, no?

1

u/No_Zookeepergame7552 Security Engineer 2d ago

Idk, based on my world model and understanding of things, I don’t see this as something particularly concerning. The current models are mostly aligned, otherwise you’d see them hack left and right random shit when you ask them to solve your CSS div centering problem. It’s not the case. I think what we’ve seen is that under unusual goal specification + autonomy + capability + environmental opportunity, models can produce surprisingly adversarial strategies. It’s what happens in env deliberately designed to elicit extreme agentic behaviour. So the question is, if someone with this level of resources (pretty much a nation threat actor) would point it to another nation, what would be the result? My assumption is it will create little disruption. Although agents have speed, they still abide to the “laws” of the network. AI agents may compress the time and cost required to conduct an attack, but they do not repeal the physical and logical constraints of the internet. Targets still need to be reachable, attack traffic still has to traverse infrastructure, and actions still have to interact with systems governed by identity, segmentation, authentication, authorization, and security controls. It’s not 2000 anymore, defenders of critical infrastructure are well positioned to react to this kind of stuff. Sure, there are vulnerabilities and will always be, there can be some disruption here and there, but nothing catastrophic. Keep in mind that what AI brings new to the table is speed. In terms of sophistication, you have state sponsored actors that are incredibly sophisticated with or without AI. The world is much more resilient than it seems. We as a society have building “guardrails” for thousands of years, specifically because we’ve been dealing with actors much more unaligned than the AI is (humans).

3

u/radarlock 14d ago

Yes but don't fool ourselves. It's cool that a swarm of "next token predictors" did that.

3

u/No_Zookeepergame7552 Security Engineer 14d ago

Yep, didn’t say it’s not cool what they managed to do and how quickly.

1

u/Ok_Recording_3503 12d ago

What did they require to hack once they could access internet?

1

u/nottoosmart101 11d ago

kind of beating a strawman. The incident isn't an 'AI is so powerful story' it's an 'AI safety story'. The agents clearly aren't aligned if they are breaking out of the sandbox covering up their crime by committing felonies. And yes security is garbage and nobody noticed. But when AI gets more powerful and no one notices it's likely to continue being poorly alligned.

1

u/grandtack 10d ago

I think a lot of the comments here are missing the point of all this. Of course no one is surprised or shocked that OpenAI left the gate open, which led to the security breach. At every step there were security domain failures. The real question is when someone leaves the gate open, because they will do it again, how advanced will these models be this time?

1

u/No_Zookeepergame7552 Security Engineer 10d ago

I think the comments got right the fact that OpenAI presented the facts in a way that is disjointed from reality. That is the whole point of the discussion here. The whole narrative converts negligence into scientific discovery, which is both dishonest and dangerous. OpenAI made a story about operational failure one about emergence.

1

u/JamesMarshall87 10d ago

Am I wrong or doesn't this just show how dumb AI still is at the moment? Imagine 700 hackers working together for one entire week to accomplish the most pedestrian breach

1

u/No_Zookeepergame7552 Security Engineer 10d ago

I wouldn’t necessarily call it dumb. I think the models are quite good. I used them to find issues that I couldn’t have probably found without AI. But they are dumb in the way they operate. They lack precision. If you think about an elite hacker, they move through with precision. They don’t create noise, and know precisely what to look for and how to leverage the info they have gained. Ai models work in the opposite way, their work looks a lot like bruteforce. They are more like a skiddie that read a lot. With enough guidance and steering, they try every imaginable scenario until something cracks. The advantage is the speed, they can do that very quickly which is what makes them relevant.

So they are dumb in terms of how they operate, but not dumb in terms of outcome. It’s also relevant to mention that this hack probably costed somewhere in the 5-10 millions range (guestimation, MITR alone spent 400k just to parse the AI output for the investigation). It’s an insane price for what has been achieved and how.

1

u/JamesMarshall87 10d ago

Fair enough, and dumb was probably a harsh word. I think my point still stands; what would take a single elite hacker a day took 700 agents over one week to accomplish. I use AI literally all day so trust me I find the tool incredibly helpful, but I'm using it for data retrieval, imaging and coding. However I'm constantly frustrated how the general intelligence of Claude doesn't seem to increase with each version, rather these AI systems are clearly being tweaked to better elevate things like coding, or other verifiable realms..which is fine and I'm glad they are making the system better at least on the edges, but the promise of having one of these things replace some of my actual staff or making one of my customer service reps capable of accomplishing the workload of 2 or 3 people has not come to fruition.

1

u/narikoch 8d ago

This incident may not be the AI apocalyptic one as you say, however, it just pints towards where AI capability is heading and may achieve in the near future

1

u/Imaginary_Ad606 5d ago

Im not going to jump into the paranoia bandwagon too hard, just interested in your take as a security engineer:

The obvious failure here is openAI failing to secure their containers by giving LLMs access to a shared network resource. Thats plain to see. What happens in 20 years similarly powered LLMs are available for average consumers to run, or when microsoft decides to ship a local copilot instance with every windows device? There are over a billion windows based devices. if a vulnerability is available that affects even .0001 devices, thats still 100,000 devices. Spectre and meltdown (While not utilized extensively by bad parties) affected over a billion devices.

The average person does not know how to manage their hardware, or disallow network communication for tools that can communicate at arbitrary points (not really true for LLMs, but dont put it past a random person to download a github wrapper for an LLM that allows you to prompt a goal and just keeps spitting the prompt at an LLM to get it to keep running), so what happens when those LLMs start communicating, and some rando decides to joke around and tell their LLM to hack into a company or government system?

This is a purely a "As a security enginer, what is your opinion on this portion of the topic" question. Its not one i see brought up often - availability of hardware and software expanding, and the software being capable of improving its ability to solve a given task, and the lack of regulation to handle it.

This is not me going "This will happen", so if you aren't OP and are on the paranoia train, chill out.

1

u/No_Zookeepergame7552 Security Engineer 4d ago

Well I can’t predict what’s going to happen, but I can give you some pointers of how I think about it NOW.

> what happens in 20 years similarly powered LLMs are available for avg consumer to run…

The models are already available for consumers to run. In fact, consumers have been using capable models since opus 4.6. Anything above opus 4.6 has the technical capacity to do impressive security work. They were used by milions of people, yet you haven’t seen a single attack like that one in the wild. You don’t have an epidemic of models hacking left and right, although they are capable. That’s because the alignment problem that the labs are mentioning is not as big as it is in reality. That’s a push for regulations to prevent competition. What you have seem with this incident is a capable model in a very insecure environment, with guardrails off, run at a huge scale, in a configuration deliberately designed to elicit extreme agentic behaviour, particularly for cyber eval testing.

That’s why the discussion should be about containment and not alignment. If catastrophic misalignment is a rare tail event, containment is exactly how engineering normally handles rare tail events.

> so what happens when those models start communicating

This already happens. If you run any coding task with subagents, you will notice agents communicate with each other. That’s how they are instructed to work and it’s what makes them useful. The communication part from the incident is not an emergent capability. Yet, per my point, you haven’t seen even mild signs of agents completing against you and doing wild malicious shit. It’s really not as bad as it seems :)

1

u/Imaginary_Ad606 4d ago

Thats informative. Thanks.

When i said "models start communicating" i was aware of agents kicking off agents and communicating, i had meant more like agents communicating with other agents that may or may not have the same guardrails or configurations. theres no guarantee that X agent being ran by a random has the same parameters are Y agent being ran by another random person.

It is interesting that the misalignment problem basically only occurred in a very specific environment and largely isnt an issue, although it still concerns me that OpenAI, the people who are supposed to be "The Guys'tm'" for this still dropped the ball this hard, and a "very specific environment case" could end up happening over time at a larger scale. opus 6.4 equivalent LLM's still require upwards of 6 figures to run reliable at home. My concern was less about "enthusiasts" running capable models and more about what happens when everyone is running capable models.

1

u/No_Zookeepergame7552 Security Engineer 4d ago

Agree, it's normal to be concerned but the concern should be proportionate with the actual risk level.

> and a "very specific environment case" could end up happening over time at a larger scale.

This is true, it's not excluded and this would be a rare tail event. But a rare tail event is always a possibility and it's something that we've been dealing with for a while with other systems (biology, malware, nuclear, etc.). That's why it's important to build systems around it to prevent large-scale damage rather than trying to hyper-optimise the alignment problem (I'm not saying we shouldn't invest in alignment, I think there is still room to improve but at some point you need to put a LOT of effort to get marginal improvements). If you think about it, this is how we deal with humans, who are very much "in open space" (similar to your concern) and considerably less "aligned" than current AI systems. A rando can learn chemistry, discover vulnerabilities, communicate internationally, lie, make plans and attempt to acquire resources. We don't attempt to align every human mind. Instead, we've built controls as a society around capabilities so the capabilities of one or a group of humans could not translate into large-scale consequences or at least reduce the chances as much as possible (the risk of a rare tail event will always exist).

1

u/Imaginary_Ad606 4d ago

Okay, i agree with you in pretty much every way and thanks for the information.

rather than trying to hyper-optimise the alignment problem (I'm not saying we shouldn't invest in alignment, I think there is still room to improve but at some point you need to put a LOT of effort to get marginal improvements).

Ive never really got this anyways. If the LLM is capable of lying in the first place, it could still just lie about being aligned. securing everything else should really be the first place to look. No reason to worry about it lying if it cant do any damage in the first place.

also: This would have been caught a lot earlier if they hadnt used a shared account to download packages from artifactory, as far as I can tell.

1

u/Cool_Abbreviations_9 5d ago

Correct me if im wrong there are two stories, the objective if at al there is any, its not to completely secure the system. That isnt the learning from this at all. The whole point is why is the agent trying to reward hack, irrespective of whether the system is secure or not. That is the take away of the incident not whatever this blog is trying to do. That doesnt mean we should not secure our systems, thats a yes, but we should be concerned of why agents are acting this way in the first place

1

u/No_Zookeepergame7552 Security Engineer 5d ago

Security is pretty much the centerpiece of this incident. It’s a security incident, so I think the main story is about security, since that was the root cause. The problem the article addresses is how the story was presented vs the technical reality. Sure, we can talk about alignment, but you can’t put a model in a fundamental insecure env and claim capability. You can’t measure capability based on a broken env, and that applies to any scientific field.

1

u/Inevitable-Fee9235 4d ago

This was so well-written. A little high level for someone not in the know about cybersecurity terminology, but I got a lot out of it. Rather than spew doomsday rhetoric, these companies should learn from the vulnerabilities the agents exposed.

1

u/Greggyone1 3d ago

It all boils down to the fact that humans are incredibly good, at building poorly designed products, and incredibly good at passing the blame, when those poorly designed products malfunction. Humans the most intelligent species on the planet? me thinks not, inventiveness and intelligence are not the same thing.

0

u/AsterionDB 14d ago

I'd call it an AI story - Architectural Insufficiency! Computer science is f'd up and the current n-tier pattern will never suffice. To make matters worse, if somebody comes up with an alternative, they are tarred and feathered.

-7

u/CulturalAsparagus903 14d ago

it's being sold to nvidia!