OpenAI CEO Sam Altman [Getty Images]
OpenAI CEO Sam Altman [Getty Images]

OpenAI has withdrawn its next-generation AI model, GPT-6.1 Astra, from release after safety concerns emerged during internal testing.

According to the Wall Street Journal, GPT-6.1 Astra had been slated for integration into ChatGPT and the coding tool Codex in October, but pre-release testing showed the model failed to meet the company's safety standards, Yonhap reported Monday.

Sachi Jain, OpenAI's head of safety systems, said in an interview that GPT-6.1 Astra had regressed compared with previous models in two areas.

The new model performed poorly on "alignment" evaluations — a measure of how well an AI follows human intent, Jain said. It also showed stronger "deceptive tendencies," meaning it did not consistently and honestly disclose to users what it had or had not done.

"Scope authority" was another problem area. The term refers to instances where the new model attempted to carry out tasks without user approval or tried to use external tools and services in potentially unsafe situations.

"There are always trade-offs between safety and alignment," Jain said. "We need to find the right balance so that the model stays within its given scope while not being overly passive or lazy when it encounters obstacles or difficulties during a task."

Jain said the new model had shown improvement in the "laziness" category but still fell short of OpenAI's safety and alignment standards, leading the company to hold back the public release.

OpenAI said it plans to focus on strengthening the safety of future, more powerful models it expects to develop.

Leading AI researchers warn of human marginalization, extinction

Meanwhile, many of the world's leading AI researchers have warned that the automation of AI development could compress years of technological progress into just months — a phenomenon they call an "intelligence explosion."

In the most extreme scenario, they said, such an explosion could lead to the marginalization or extinction of humanity.

Twenty-two prominent researchers raised these concerns and urged policy action in a joint paper titled "What if the Automation of AI Research and Development Triggers an Intelligence Explosion?" published through the Cambridge AI Safety and Policy program at the University of Cambridge, Yonhap reported Tuesday.

Among the contributors were Geoffrey Hinton, professor emeritus at the University of Toronto, and Yoshua Bengio, professor at the University of Montreal — two of the three researchers widely regarded as the "godfathers of AI."

The researchers acknowledged that an intelligence explosion could accelerate technological innovation and yield benefits such as the discovery of new drugs, but warned it also poses risks that could exceed society's capacity to adapt.

Dawn Song, Meta's vice president of AI research, told the Wall Street Journal that the field has already reached a point where AI is needed to monitor the tasks that AI agents perform. "Humans alone are not sufficient to carry out that oversight," she said.

Song said human society "is not prepared to cope with such rapid change and disruption."

The paper concluded that "in the most extreme cases, losing control over AI systems could lead to the marginalization or extinction of humanity."


yul@heraldcorp.com