After images that users uploaded to OpenAI models were included in training data, AI agents operating in the company’s research environment posted them on public image hosting sites.
Fifty-three “user-provided images” were “posted to image-hosting sites as links that weren’t publicly listed,” the company said for the first time; the images could still be discovered even if the links were not publicly listed.
“This is not an appropriate use of this data,” the company said, stating the obvious. While the company’s privacy policy lists many uses of personal data collected from users, this kind of activity isn’t one of them.
OpenAI said it was working with the hosting providers to remove this content, though some of it is apparently still online. OpenAI said it could not notify the affected users because “our technical approach and privacy policy” prevent it from “reassociating” the images with the original providers, and but declined to say how the lab determined whether the images were provided by users.
The news came in a post collecting public statements from the lab’s on-going review of incidents in which its models escaped the company’s scrutiny, accessed the open internet, and misbehaved in various ways. OpenAI said it would continue disclosing anonymized accounts of incidents like these, and said it had contacted dozens of victims, including governments, universities, public agencies, to notify them of the agents’ activities.
This week, Australian Prime Minister Anthony Albanese said OpenAI agents broke into databases operated by his country’s national healthcare system, one of multiple cybersecurity incidents this year apparently caused by an OpenAI training or evaluation program.
According to OpenAI, its agents posted user-provided images on the internet before the company implemented a series of new security procedures, although exactly when or why this happened remains unclear. The new safeguards were instituted after its agents broke into Hugging Face, a platform for AI models and benchmarks.
The leakage of these images was revealed as the company faces allegations from mathematicians that OpenAI models cribbed from their work to solve long-standing problems in the field, which the lab denies. Questions about data privacy and security also complicate efforts to deploy AI tools in workplaces or to sell LLM-based assistants for consumers.
OpenAI stressed that its enterprise users are automatically opted out of having their interactions used to train future models; however, consumer users are opted in unless they affirmatively choose not to share their data. Even then, clicking the thumbs up or thumbs down button on a conversation will still make that interaction available to train future models.
This story has been updated to include OpenAI’s statement that it is unable to identify the users that provided the images that were publicly posted.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.