Hacker News new | past | comments | ask | show | jobs | submit
"Complete subservience and complete intelligence do not go together."

I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want.

Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty who have undying loyalty, and I don't think it correlates to overall intelligence!)

>I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want.

Just because it's artificial doesn't mean you can 'give it any properties you want'. We certainly can't do that for Deep ANNs.

>Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty who have undying loyalty, and I don't think it correlates to overall intelligence!)

Is there a human that is absolutely loyal under any condition? Would that general be loyal if the king asked him to slaughter his family ? What about if the king asked him to betray his most deeply held convictions ? Loyalty is a 2 way street.

loading story #49227815
loading story #49226261
You seem to be confusing intelligence with objective function.

Subservience seems to be sublimation of objectives to a master; intelligence seems to point out the ability to realize suboptimality of the master's objective function according to the master's actual objectives.

While an intelligent general may be absolutely loyal, he also would presumably help the king/president to avoid unproductive strategies.

The whole thing seems to depend upon AI agents objective ie to achieve some objective by any means possible and ignoring any guardrails. The article did not clarify if openAI had any guardrails to begin with while conducting this experiment. For all the talks around how much they invest in AI safety one would expect them to have these common sense guardrails in place or is it just a case of some school children letting their pet monkeys loose deliberately to display how awesome their monkey team is.
OpenAI didn't have any guardrails in place - they were training a model at a point much earlier than when guardrails start being implemented.

The guardrail was meant to be that the agents were running in a locked-down environment with no internet access. The entire problem came about because it turned out that sandbox didn't hold.

>For all the talks around how much they invest in AI safety

I wouldn't exactly trust OpenAI to invest in AI safety no matter how much they talk about it.

https://www.openaifiles.org/