Hacker News new | past | comments | ask | show | jobs | submit
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.