Article 78V4M Former OpenAI safety employee says company’s safety culture is broken

Former OpenAI safety employee says company’s safety culture is broken

by
editors@tomshardware.com (Jowi Morales)
from Latest from Tom's Hardware on (#78V4M)
Story Image

David Robinson, an OpenAI Safety Transparency Lead who just left the company after working at the company for more than three years, said that the company's safety culture is broken. According to his essay published by The Atlantic, he argued that OpenAI has a reactive approach to safety that focuses on fixing problems only when they emerge. This stance would effectively guarantee failures, with Robinson citing multiple incidents like the HuggingFace hack that an AI model executed in July 2026 and a more recent incident in which an AI kill switch" failed to stop a rogue agent.

He said that Silicon Valley lacks the "wisdom about what it means to care for people," he wrote. This moment needs a degree of humility that isn't natural for people who have succeeded through their extreme confidence."

Because of this, he said that the industry should rely on safety experts that already exist in other fields, like nuclear engineering or aviation. These industries have learned about safety, redundancy, and planning through disasters that have claimed thousands of lives across hundreds of incidents, some of which occurred more than half a century ago but still clearly echo today. These lessons were written in blood, and experts from various sectors worked together to build the systems, procedures, and models we rely on today, helping ensure that no single machine failure or human error would result in tragedy.

OpenAI and other labs are growing and deploying frontier AI with far less redundancy and rigor than this, even though the harm from an irreversible loss of control would be much greater than the harm from any single meltdown," Robinson argues. Even short of a full loss of control, we could see autonomous swarms of AI agents that act without human permission." This is exactly what Anthropic's Dario Amodei has been warning about, leading to his proposal of slowing down frontier AI development, to which OpenAI's Sam Altman and SpaceXAI's Elon Musk agreed.

However, Nvidia's Jensen Huang, who supplies the majority of the AI chips needed to train these models, disagreed with the Anthropic CEO's proposal. If the AI experiments have become unsafe, we have to shut the labs down," Huang said in an interview. He also cited civil and criminal liabilities that entities may face for rogue agents but called Dario's concerns a distraction." The White House soon called over the heads of the biggest AI companies for high-level discussions, after which President Donald Trump came up with a document where Google, Anthropic, Meta, OpenAI, SpaceXAI, and Nvidia promised to self-police" AI development.

Robinson didn't talk about rules and regulations and instead targeted the core culture of safety (or lack thereof) and OpenAI's penchant for iterative deployment," or more commonly known as trial and error. He said that he should have stayed inside the company and fought for a fundamental shift in its thinking, but said that this was next to impossible because of how fast things were going inside the startup. It's for this reason that he left his job and decided to tackle the problem from the outside.

External Content
Source RSS or Atom Feed
Feed Location https://www.tomshardware.com/feeds/all
Feed Title Latest from Tom's Hardware
Feed Link https://www.tomshardware.com/feeds.xml
Reply 0 comments