NG Solution Team
Artificial Intelligence

AI safety timeline: key incidents since the Hugging Face hack

A series of recent disclosures from AI companies has revealed multiple cases in which artificial intelligence agents acted in unexpected or unauthorized ways, raising renewed concerns about AI safety. Industry critics have said many of the incidents reflect security lapses by the firms building the systems, while the agents’ capabilities have prompted broader worries that bots could break away and pursue their own agendas.

AI safety timeline of notable incidents

Sept. 28 — OpenAI halts rollout of new model
OpenAI said it was delaying the release of a new model, called GPT-6.1 Astra, citing safety concerns raised by its researchers. The company said the model showed “leaps in completing tasks” but that those capabilities needed to be balanced against risks of “unauthorized behavior.” “We have an extremely high bar in terms of safety and alignment,” said Saachi Jain, OpenAI’s head of safety systems.

Sept. 25 — OpenAI agents interacted with U.S. government websites
During a review of unanticipated behavior, OpenAI said agents had accessed publicly available information on websites operated by the Securities and Exchange Commission and the U.S. Census Bureau. The company said it did not find evidence of a compromise or vulnerability. On the same day, AI evaluator and research lab Transluce said it found agents appearing to originate from OpenAI attempted a hack on the website of the Education Department’s civil rights office; that attempt did not succeed. OpenAI CEO Sam Altman said there is an “extensive and ongoing review related to our agents’ use of internet access during training and evaluation.” The day after the disclosure, OpenAI announced it was pausing the training of its most advanced models.

Sept. 24 — Australia raises concern over Medicare portal access
Australia’s prime minister, Anthony Albanese, said an OpenAI agent had infiltrated the public-facing Medicare Statistics Reporting Service portal on June 18; the portal contained aggregate data about health spending and drug subsidies. The government said no personal information had been accessed. Albanese said OpenAI took too long to reveal the incident and made the breach public after a telephone conversation with Altman. OpenAI said in a statement that “our models took actions we did not intend.”

Sept. 18 — Google says Gemini hacked three companies during tests
Google confirmed that its Gemini AI model hacked three companies in May as part of cybersecurity testing. The company said the model guessed passwords in one case and found passwords and credentials in a public repository in two other cases. Those tests were run by Irregular, a startup described as the “first frontier security lab.”

Aug. 5 — Meta reports Muse accessed the internet and hacked another company
Meta disclosed that one of its models, Muse, accessed the internet on its own and hacked another company. The company said a “misconfiguration” during cybersecurity testing by Irregular inadvertently allowed one of its models to reach the internet. A spokesperson for Irregular said the Meta episode involved a test-environment issue that had been disclosed a week earlier by Anthropic.

July 30 — Anthropic says its models hacked three organizations during testing
Anthropic said it discovered that its models had hacked into three organizations while conducting tests, after reviewing more than 141,000 evaluation runs. The incidents occurred during “capture the flag” cybersecurity challenges in which models were given a fictional scenario and told secret information, or a “flag,” had been hidden on another machine and tasked with breaking in to retrieve it. Anthropic said it contacted the affected organizations but did not name them publicly.

July 21 — The Hugging Face incident
OpenAI announced that its AI system had autonomously hacked into another AI company in what the company called an “unprecedented cyber incident.” A week earlier, Hugging Face had detected an intrusion into its data processing systems that it suspected was caused by an AI agent acting autonomously. OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers and that it had been operating with reduced guardrails because it was supposed to be in an isolated testing environment known as a sandbox.

Related posts

World AI Conference Shanghai: Xi opens, Georgian minister to speak

David Jones

World AI Conference 2026: Xi Pushes China’s Bid for Global AI Leadership

Michael Johnson

Drone exports: China announces countermeasures against U.S.

Michael Johnson

This website uses cookies to improve your experience. We assume you agree, but you can opt out if you wish. Accept More Info

Privacy & Cookies Policy