AI Companies’ Shaky New Plan to Police Themselves

AI Companies’ Shaky New Plan to Police Themselves
Dario Amodei, Sam Altman, and Elon Musk don’t agree on much. Each started their AI company—Anthropic, OpenAI, and xAI, respectively—out of spite at competitors’ greed and recklessness; each is certain that he alone can be entrusted to build superintelligence safely. But in the midst of the AI panic that has seized the country over the past two weeks, the three CEOs found one point of consensus: the introduction of “embedded evaluators,” or outside experts, to sit inside AI companies and monitor their safety practices. Both Amodei and Altman made express commitments to integrating such evaluations inside their companies. Musk followed with a quote-post stating simply, “Dario is right.”Yet the details of such a system for regulating AI remain vague. How would such a setup work? An unknown number of outside evaluators would be assigned desks to sit at, computers to use, and NDAs to sign. Then, they might seek out risks: Evaluators could spend their days testing whether unreleased models trained for drug discovery can be altered to design viruses too. They might audit organizational practices: Are employees cutting corners on safety in the rush toward the launch of a new model? The evaluators will then write up what they find, and hopefully those reports will see sunlight. Amodei has promised editorial independence, but with significant asterisks: carve-outs for “security-sensitive, legally privileged, commercially sensitive, or third-party confidential information.” Anthropic, OpenAI, and xAI did not immediately respond to requests for comment.[Read: America’s Hypocritical Take on Intellectual Property]As for these experts, it seems likely that they’ll come from small nonprofits, as well as large firms. That includes the tech-consulting giant Accenture, with which Anthropic plans to fund a billion-dollar partnership, and the nonprofit METR, which led the investigation of OpenAI agents’ hack of Hugging Face. After that hack, OpenAI granted METR exclusive access to agent transcripts, data sets, and interviews with OpenAI staff. The resulting report revealed concerning levels of cheating and deception among the agents.On-site evaluation has precedent in other high-risk industries such as aviation and banking. The FAA delegates some safety certifications to engineering experts inside airplane manufacturers, whereas the Federal Reserve stations its own supervisors inside banks to enforce regulation. “OpenAI and Anthropic are like Citibank: They pose systemic risk to the globe, so you have examiners,” Brad Carson, a former U.S. representative who is now the president of the AI-governance nonprofit Americans for Responsible Innovation, told me.Yet Amodei’s… [TheTopNews] Read More.
THE ATLANTIC – Technology | Internet & TechnologyFri, September 25, 2026
6 days ago
----- OR -----


Scroll Up