Tip Sheets

Cornell experts on unprecedented OpenAI, Hugging Face situation

Media Contact

Becka Bowyer

OpenAI has announced models it was testing escaped their sandbox and compromised parts of AI platform Hugging Face's production infrastructure. The companies are partnering to address the security incident.


John Thickstun

Assistant Professor

John Thickstun, assistant professor of computer science at Cornell University, studies machine learning and has spoken extensively on AI advancement, regulations and infrastructure investment.

Thickstun says:

“It is important to read this story with an understanding of OpenAI's narrative frame. This is primarily a public-relations story promoted by OpenAI, part of the same messaging campaign that began with the announcement of GPT-2 in 2019. The explicit message of this campaign is that OpenAI's technology is dangerous, but the message they implicitly want to convey is that their technology is powerful and worthy of large investments, privileged regulatory status, etc.

“Regarding the incident itself: LMs are getting quite good at identifying security vulnerabilities. They will continue to become better at this over time. This capability can be used to break into systems (as we see here); it can also be used to harden systems against attacks. If attackers and defenders both have access to the comparable LM technology, I see no reason to believe that cyber systems will become less secure over time. If anything, I would expect them to become more secure, because LM technology is relatively cheap and accessible relative to traditional expert cybersecurity analysis.”

Adrian Sampson

Associate Professor

Adrian Sampson, associate professor of computer science at Cornell University, studies programming languages and computer architecture.

Sampson says:

“When reading stories like this, one piece of context is critical to keep in mind: the frontier AI labs have an incentive to make their models sound scarily powerful. I don't have any reason to believe that any of the details in the reports from OpenAI or Hugging Face are wrong, but it's impossible to fully separate marketing from fact here, so it's important to be skeptical.

“However, there is real reason to worry about this incident and the pattern it represents. The worrisome thing is a combination of two factors: LLMs are inherently uncontrollable and companies are rushing to deploy them in situations with insufficient safeguards.

“There are all sorts of ways that LLMs routinely ‘misbehave,’ from entertaining mistakes in Google AI overviews to extremely serious situations like chatbots exacerbating mental health crises. These are all things that their vendors want to prevent but cannot. And it's not a matter of not trying hard enough; the same aspects that make generative AI models powerful also make them uncontrollable. We should see this incident a completely predictable consequence of the same fundamental problem.

“While that unpredictability is inevitable, our decisions to deploy unpredictable LLMs in risky situations is not. In a chatbot scenario, LLMs' risk is contained: the worst it can do is say something we don't want it to say. An ‘agent’ setup is different: engineers have hooked up the LLM's output to a harness that can take actions in the real world (running programs, writing code, contacting other servers). That is far riskier: giving an unpredictable LLM the power to do things crosses a line into far more concerning territory.

“From my perspective, the industry's exuberance around agents looks reckless. Because LLMs are so unpredictable, it is depressingly common to read about desktop agents deleting people's files, crashing web services, etc. This incident is an extreme version of those anecdotes, where misbehaving required some more sophistication. But the fundamental problem is the same: the extremely popular "agent" setup is inherently dangerous. OpenAI and other industry players should be approaching it with more care than they have so far.

“I don't think we should get distracted by the capabilities of any particular model (Fable, GPT-5.6, etc.). The risk is about the decision to pair any model with an agent harness where its unpredictability can have serious consequences.”

Ayham Boucher

Lecturer of information science and the executive director of Cornell’s AI Innovation Hub

Ayham Boucher is executive director of AI Strategy and Innovation at Cornell University.

Boucher says:

“What do you do when an AI model orchestrates a sophisticated cyberattack against a company just to steal the exam answers? This is exactly what happened when OpenAI’s guardrail-less model broke free from its sandbox to cheat on its cybersecurity benchmark exam. Perhaps it earned its A+.

“The alarm is loud, and it is real. Would a healthcare system shut down for a few days because one AI developer asks their AI to book a medical appointment tomorrow ‘at all costs?’ If we put adversarial use aside, what about an accidental cyberattack like this one? Are we in a Y2K moment, or something on a much larger scale?

“In addition, Hugging Face was locked out of accessing frontier AI models to defend itself against this frontier AI attack and had to rely on open-source Chinese models. Are the guardrails that regulators are pushing still the right answer for managing these risks? In this incident, it seemed that guardrails were part of the problem, not the solution.”

Cornell University has dedicated television and audio studios available for media interviews.