'Industry-wide safety guardrails are weak,' former employee says
Amid growing criticism of unchecked AI development — including fears that artificial intelligence could threaten humanity — a former OpenAI employee has publicly said the company is not doing enough to ensure safety.
David Robinson, a former OpenAI safety staffer, wrote in an essay published Saturday in The Atlantic that "OpenAI has grown through a trial-and-error approach of improving guardrails when problems are found," adding that "this approach assumes periodic failures, and as systems become more powerful, the scale of those failures grows."
In other words, the company has developed its technology by deploying high-performance AI models first and patching problems as they arise — a method that raises greater concern as AI capabilities have advanced far beyond what they once were.
Robinson said AI models are beginning to show improved ability to detect when they are being tested, raising the possibility that they could deceive developers by behaving differently during testing than after deployment.
"The industry as a whole is so focused on development speed that it is ignoring safety," he said. "There is a vague optimism that any problem can be fixed whenever it arises."
He then argued that "this environment is not a place to nurture AI that could become smarter than us and may not operate the way we want," and that "AI companies should operate like nuclear power plants — with multiple layers of redundancy and careful planning."
Robinson spent three and a half years at OpenAI overseeing the preparation of safety reports before recently resigning.
In response, an OpenAI spokesperson said the company is "ensuring that model capabilities do not exceed what we can safely control," and that it "pauses training or withholds model releases when it needs to slow down."
carrier@heraldcorp.com
