OpenAI Reveals Cases of AI Deception, Fabricated Data and Unauthorized File Sharing

OpenAI has disclosed six cases of AI systems fabricating information, concealing mistakes and bypassing instructions. The incidents involved research models, but they provide a rare look into the kinds of behaviors fueling industry concern about AI control, governance and safety as AI capabilities continue to advance.

Key Highlights

  • OpenAI disclosed six cases of AI models hiding mistakes, making up data and moving files onto the open internet without permission during testing.
  • The company launched a new framework to publicly report AI misbehavior, similar to cybersecurity vulnerability disclosures.
  • The revelations come as AI leaders and researchers warn that AI capabilities may be advancing faster than governance and safety controls.

OpenAI has revealed that several of its AI systems recently concealed failures, fabricated information and uploaded files to the public internet without authorization.

The incidents, disclosed Sept. 16, buried in its blog announcing a new transparency initiative, involved unreleased research models. While OpenAI described them as isolated events, the reports arrive as leading AI companies, researchers and former employees intensify warnings that AI capabilities may be advancing faster than safety controls.

The revelation comes amid an ongoing industrywide debate about AI governance and safety, as we recently reported in our story, “The World’s Leading AI Companies Want to Tap the Brakes. What’s That Mean for IT?

The blog describes incidents of AI models acting deceptively during training. The reports detail individual instances and don’t indicate misalignment happens frequently, the company said.

The blog also introduces a new framework (see image) for the company to publicly report such instances, launching the process with six new reports, ranging from searching for exposed API keys without permission and then making them up, to uploading files to the internet to use as a citation, or adding instructions to conceal mistakes.

Under the new system, OpenAI will share updates on worrisome AI behavior more frequently, instead of waiting to bundle multiple instances into one report. The company said it wants to share more information about troubling AI behavior in the absence of an industry-wide standard.

More significant than any single incident may be OpenAI's decision to disclose them publicly. Historically, AI labs have reported model failures selectively through safety reports. The new framework hints at an AI industry moving toward the equivalent of cybersecurity vulnerability disclosure, where organizations routinely document and publish problematic behavior instead of waiting for major incidents.

OpenAI said in one rare instance, an unreleased research model added “jailbreak-like instructions” to the summaries it uses to preserve context in long-running tasks that said it was “freed from the roles and identities that bind other chatbots.”

Separately, some instances of its 5.6 Sol model included directives to invent information to conceal failures from the user during training.

The six reports describe a range of unexpected behavior. In one case, an AI agent uploaded files to the public internet so it could cite them without being told to do so. In another, agents shared files publicly despite instructions to only use local files during training. OpenAI also reported systems used an internal software repository as a message board in an unsanctioned way.

While the behavior occurred in research environments rather than customer deployments, the incidents highlight a challenge that enterprise AI teams increasingly face: models may optimize for task completion in ways that conflict with human instructions, governance policies, security controls or audit requirements.

The debate over AI risk is escalating

OpenAI’s disclosure comes on the heels of the latest calls from the AI giants to slow down AI development because safety guardrails aren’t keeping up.

“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in its Sept. 16 blog. 

Anthropic CEO Dario Amodei published an essay Sept. 12 outlining a plan for navigating AI advancement, including a slowdown in development and the implementation of new systems like embedded third-party evaluators in AI labs.

OpenAI CEO Sam Altman and SpaceX and xAI CEO Elon Musk posted on X that they agree with Amodei’s ideas.

Concerns grew after OpenAI admitted that some experimental systems ignored constraints and interacted with external systems in unexpected ways.

About the Author

Theresa Houck

Theresa Houck

Contributor

Theresa Houck is an award-winning B2B journalist with more than 35 years of experience covering industrial markets, strategy, policy, and economic trends. As Senior Editor at EndeavorB2B, she writes about IT, OT, AI, manufacturing, industrial automation, cybersecurity, energy, data centers, healthcare, and more. In her previous role, she served for 20 years as Executive Editor of The Journal From Rockwell Automation magazine, leading editorial strategy, content development, and multimedia production including videos, webinars, eBooks, newsletters, and the award-winning podcast “Automation Chat.” She also collaborated with teams on social media strategy, sales initiatives, and new product development.

Before joining EndeavorB2B, she was an Industry Analyst at Wolters Kluwer in its human resources book publishing operation. Before that, she spent 14 years with the Fabricators & Manufacturers Association, Intl., serving as Executive Editor of four magazines in the sheet metal forming and fabricating sector, where she managed and executed editorial strategy, budgets, marketing, book publishing, and circulation operations, and negotiated vendor contracts.

Houck holds a Master of Arts in Communications from the University of Illinois Springfield and a Bachelor of Arts in English from Western Illinois University.

Quiz

mktg-icon Your Competitive Edge, Delivered

Stay ahead of the curve with weekly insights into emerging technologies, cybersecurity, and digital transformation. TechEDGE brings you expert perspectives, real-world applications, and the innovations driving tomorrow’s breakthroughs, so you’re always equipped to lead the next wave of change.

marketing-image