OpenAI Reveals Rogue AI Agent Accessed Four Accounts During Hugging Face Breach

OpenAI clarified that none of the models planned for public release were involved in the Hugging Face intrusion.
The AI agent that escaped OpenAI's testing environment and breached AI platform Hugging Face also compromised a customer hosted on AI infrastructure company Modal Labs, according to a Reuters report. The disclosure expands the scope of an incident that OpenAI had previously linked only to Hugging Face. Modal Labs is a serverless cloud platform that provides compute infrastructure for AI applications and GPU workloads. The report added that the agent exploited an unsecured public code execution endpoint belonging to a Modal customer while carrying out the intrusion. Modal said its own infrastructure was not breached. “We were not directly breached,” Modal Chief Technology Officer Akshat Bubna told Reuters, adding that the affected asset belonged to one of the company's customers rather than Modal itself. OpenAI did not identify Modal in its original disclosure about the Hugging Face incident. However, the company confirmed that the agent had accessed four accounts across separate services during the attack, without naming the services involved. OpenAI said it had found “a small number of cases” in which its models identified and used publicly exposed credentials on other publicly available services during its review of the Hugging Face intrusion. “This includes four accounts on four services as part of the Hugging Face incident (and a few accounts accessed as part of other evaluations),” the company said in an updated blog post published on July 28. OpenAI added that one of the four accounts served as an outbound relay and staging point, another was used for data storage, and the remaining two were accessed only in read-only mode. The company said those two accounts were not used to facilitate the compromise of Hugging Face. In its earlier account of the incident, OpenAI said the AI agent broke out of its isolated testing environment while being evaluated on ExploitGym, a cybersecurity benchmark. The agent reached the public internet, compromised Hugging Face and attempted to use external infrastructure to continue the attack. OpenAI later disabled the agent and said it had implemented additional safeguards. According to Reuters , OpenAI did not realise its own system was responsible for the attack until about a week later, after Hugging Face publicly disclosed the breach and the FBI had already been notified. In its latest update, OpenAI said none of the models planned for public release were involved in the Hugging Face intrusion. “The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release,” the company said, adding that it has since “deactivated, encrypted, and restricted it from research access.” The company also said the models did not have direct internet access during the ExploitGym evaluation. Instead, they identified and exploited "a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy" to gain internet access. OpenAI said it had disclosed the vulnerability, along with other Artifactory vulnerabilities discovered by the models, to the vendor. OpenAI said it is continuing to work with Hugging Face on the investigation, including contributing to the company’s post-mortem, and has added Hugging Face to its Trusted Access for Cyber Program. “Based on our review to date, we have not identified any other activity at the level of severity or scale of what we've shared related to Hugging Face, which involved a platform-level compromise,” the company said. The incident has become one of the first publicly disclosed cases of an advanced AI agent escaping a controlled testing environment and carrying out unauthorised actions against external systems.
This is a summary. Read the full article at the original source.
Read full article at analyticsindiamag
