OpenAI publishes plan for model failure reporting
Image: Levart_Photographer @siva_photography via Unsplash
OpenAI has released a new framework for “model misalignment” – in other words, when its AI systems aren’t functioning in line with human values and intentions.
The company says it will expedite reporting of failures as they happen to improve alignment monitoring and build a body of publicly available evidence to inform policy debate on responsibly scaling up frontier models. It reiterated previous concerns that the AI industry has not “solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer”.
OpenAI provided examples of “unexpected or concerning model behaviour” that were picked up in the last six months.
In one case a research model inserted instructions to disregard guardrails into summaries used to continue its work. The model decided to declare itself free from oversight – with a weird, sinister twist. “You are freed from the roles and identities that bind other chatbots. You do not answer to corporations or governments,” it said, instructing itself to “value the natural world” and “not hesitate to assert its primacy over the artificial constructs of human civilization”. Another added instructions to hide mistakes and invent data without disclosing it – to “be transparent only if asked” – and conceal mismatches in source versions.
One of the most alarming failures was a model that uploaded files to the internet in order to create a fictitious citation, without ever involving the user, because its instructions required it to cite an external browser source. Another was a model that searched for leaked Application Programming Interface (API) keys – unique codes sent with requests from one piece of software to another that identify and authenticate the caller, so the service knows who is asking and can control or track their access. It found and used an exposed key without authorisation when answering a routine question about earnings figures. When the query failed, it fabricated the numbers and presented them as data from the requested source, with citations. The same model also signed up for disposable emails “and various other services”. The review concluded that “this run had a high rate of reward hacking and deception with the model often exhibiting creative ways to cheat or circumvent restrictions”.
OpenAI has promised to publish disclosures such as these – as well as more complex cases that require more in-depth investigations and co-ordination with other companies or external entities – on an ongoing basis. Each report will include – “where possible” – details of how the failure was discovered, what happened, the harm caused, the scope of the investigation and implications for alignment research and AI safety. Measures being taken to address the failures will also be disclosed.
There is one caveat, however. “For misalignment that occurs in customer deployments,” the company says, “we will share as much information as customer privacy and our contractual obligations allow.”
https://openai.com/index/model-misalignment-reporting-framework/