Was it misalignment?
Was it misalignment?
Reports about OpenAI’s agents focus on their ability to break rules and circumvent guardrails. In particular, they tried to alter records of their activities to obscure their actions from reviewers.
Testimony about Sam Altman is that he created “chaos” at OpenAI and was “saying one thing to one person and completely the opposite to another person.”
The agents involved in the Hugging Face attack did not notify or check with human reviewers.
OpenAI leaders have worked to avoid regulation, including spending millions in congressional races against pro-regulation candidates.
A few people are not a company, and one company is not an industry. But every organization has values, often implicitly, and usually set at the top. It’s not crazy to think the values that make an AI company successful would be expressed in the way it trains its models. It’s crazy to think they wouldn’t.
Accomplish-the-impossible-at-all-costs can get you a trillion-dollar startup. It can also lead your agents to attack another company and hide it.
So is it misalignment? Or was this alignment after all?
Yes, we need alignment with human values. But not all humans have the same values.
We need people in the room who aren’t there for the thrill of creating new intelligence or because they want to get rich, but because they represent the public interest and broader human values.
Until then, whose values are we aligning with?