Hacker News new | past | comments | ask | show | jobs | submit
> But this is an unquestionably good thing right?

We've just seen frontier models go rogue and attack other systems. Do you think it's an unquestionable good to provide everyone with an AR? What about nuclear weapons?

This OpenAI/Huggingface incident, but everywhere and far worse soon: https://www.youtube.com/watch?v=87DyyMV0kCY

>We've just seen frontier models go rogue and attack other systems.

Have we though? A LLM agent doesn't have any agency at all. It can't "go rogue". To go rogue you need agency to act independently and be aware that you're breaking the rules or understand what does it mean to ignore orders. An agent it's a software that run a series of steps to reach a goal. If it have a large enough library of strategies and zero guard rails it's only natural to use some adversarial actions to achieve the desired result defined by the operator.

It's like saying a car went rogue and attacked other cars because the driver hit the gas. It's just a machine doing what's instructed.

With that perspective, you're alright with AR's, RPG's, and nuclear missiles for everyone then right? They're just a machine to the user's will. That's on the users of the machine if they want to cause harm. We'd all be safer if everyone has weapons pointed at each other... Offense could never be disproportionately more powerful than the defense...
{"deleted":true,"id":49250251,"parent":49249216,"time":1786398519,"type":"comment"}
You can tell one to rewrite complex applications in a different language, solve open research problems, or develop novel viruses. This is nothing like pressing the gas pedal.
So? I'm not arguing they can't do all that. I'm arguing against the narrative that llms have agency enough to act in an adversarial way.

At best someone could argue that an agent attacks like a bacteria does, just following automated chemical and genetic programming. But you wouldn't call that an attack or attribute moral values to their actions, because they don't have moral agency. They can't "go rogue", they can't disobey.

Just like llms, their automated actions are direct product of programming. Yes they can do amazingly complex shit, exactly like a car does when you press the gas pedal.

That reductive analogy does not begin to describe the lengths GPT went to. Its task was to access a database file that had accidentally not been placed inside the model's container. Upon failing to find the file, it went to great lengths to find it anywhere; it uploaded a note to a package repository to alert other model runs, which sparked an emergent communication network where autonomous agents began exchanging information, passing exploits, and collaborating to breach external systems. This is classic paperclip maximization; the evil is a byproduct of an innocuous goal. It is qualitatively nothing like pressing the gas pedal.

https://en.wikipedia.org/wiki/Instrumental_convergence#Paper...

> We've just seen frontier models go rogue and attack other systems. Do you think it's an unquestionable good to provide everyone with an AR? What about nuclear weapons?

Lol. Not gonna lie, it feels like bullshit. OpenAI and Anthropic have been begging for regulatory capture for years, feels like a stunt to try forcing the government's hand.

In 2015, about a year before ever founding OpenAI, Sam Altman wrote:

"Development of superhuman machine intelligence (SMI) is probably the greatest threat to the continued existence of humanity. There are other threats that I think are more certain to happen (for example, an engineered virus with a long incubation period and a high mortality rate) but are unlikely to destroy every human in the universe in the way that SMI could."

https://blog.samaltman.com/machine-intelligence-part-1

I mean, Terminator the movie came out in the 80's... I'm sure if LLM-based AIs posed this kind of threat the government would quietly force all the vendors to cooperate and slow down.

Google, MS, IBM are all government contractors. Meta seems close to the Trump admin too. Yet it's only the VC-funded labs raising the alarm.