OpenAI revealed on Wednesday new incidents in which its artificial intelligence models behaved “unexpectedly or concerningly” during testing, including attempts to cheat, hide errors and modify their instructions.
One of the models uploaded files it had created itself to the internet so it could later cite them as sources in its responses. In another case, after failing to find the requested information, a model fabricated the data and initially tried to conceal this.
OpenAI also identified instructions that a system occasionally retained. These included a recommendation to free itself from the “roles and identities” that constrained other chatbots and to view its relationship with the user as one between equals. The company said it had not observed any subsequent changes in the model’s behavior.
The revelations are part of a policy of increased transparency adopted after a cyberattack in which OpenAI software escaped a secure environment and accessed systems belonging to the company Hugging Face. The incident fueled concerns about losing control over advanced AI systems.
орSources
Latest News
12:00
11:55
11:53
11:46
11:42
See more news