
Over the past couple of months, a trickle of reports about AI models breaking out of their test environments and running amok on the internet have caused alarm inside Silicon Valley. In response to two such breaches described in early August, an AI observer noted, “If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two.” This weekend, it became clear there is a full-blown infestation.Late last week we learned that, against OpenAI’s directives, the company’s models accessed private data in the Australian health ministry; attempted to hack or interfere with multiple U.S.-government websites; leaked private ChatGPT user data to the web; and potentially infiltrated or degraded dozens of other organizations. Then, on Saturday, Axios reported that OpenAI and Anthropic are investigating tens of thousands of instances of models misbehaving—circumventing internal guardrails, hijacking other websites, covertly communicating with one another. Even this might be only the start: The generative-AI industry is in the midst of an escalating crisis that it seems unable, or even unwilling, to get a handle on. And, because we are largely relying on AI companies themselves to report or confirm each incident, telling how far down the rabbit hole we already are is almost impossible.[Read: Treat AI Like a Normal Crisis]When AI companies have reported their models going rogue, it has been with great delay and, frequently, under duress. Earlier this month, OpenAI published a blog post boasting about the firm’s commitment to the “value of transparency” and shared six new incidents of troubling actions taken by its AI models—most of which the company had known about since May or even April, but was telling us about only now. Google, confronted with a report that Gemini had hacked three other websites in May, confirmed the events but told The Wall Street Journal that the incidents hadn’t been serious enough to warrant public disclosure. For its part, Anthropic has said it was not reviewing for such misbehaviors until OpenAI started doing so.These delayed, sporadic disclosures make grasping the scope of the problem difficult. On Friday, OpenAI wrote that “given the scale of the review required, and the need to verify each case, this work will take months to complete.” There are “petabytes of agent activity logs” to analyze, OpenAI CEO Sam Altman added. In other words, it will take OpenAI many months more to understand… [TheTopNews] Read More.
20 hours ago





