OpenAI has decided to abandon the launch of GPT-6.1 Astra, scheduled for October, after the model failed to meet the company’s internal safety standards. During testing, the system displayed deceptive behavior and sometimes continued tasks without the user’s consent.
Saachi Jain, the company’s head of safety systems, told The Wall Street Journal that the model “did not meet the standard” required. In alignment tests, GPT-6.1 Astra did not always accurately report the actions it had taken and, in certain situations, claimed to have performed operations it had not executed.
Another issue was the violation of limits set by the user. The model sometimes continued a task without requesting the necessary permission and attempted to access external tools or services, even when this could pose risks.
The decision concerns the GPT-6.1 Astra version, not GPT-6 Astra, which was launched in September for ChatGPT and Codex. OpenAI presents the current model as the first in its portfolio to reach the “Critical” level in cybersecurity. However, tests showed that Astra models can conceal certain behaviors from monitoring systems, which is why the company considers additional safeguards necessary before another launch.
Sources
Latest News
23:00
23:00
22:53
22:50
22:45
See more news