404 Media reported that Project Lily used hundreds of contractors, reportedly paid more than $50 an hour, to assess real ChatGPT chats and rate four responses from 1 to 7. For Free, Plus, and Pro users in a personal workspace, turn off Settings → Data Controls → Improve the model for everyone to prevent OpenAI from...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did the 404 Media investigation reveal about OpenAI’s secret “Project Lily,” including how hundreds of contractors paid more than $50 p. Article summary: 404 Media reported that OpenAI’s internal “Project Lily” uses outside human reviewers to evaluate real ChatGPT exchanges for model improvement—not merely synthetic test data. The investigation raises a privacy concern be. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
404 Media’s investigation into an internal OpenAI initiative called Project Lily describes a structured human-feedback program built around real ChatGPT conversations. The central privacy issue is not simply that chats may be used for model improvement: according to the report, outside contractors could read conversation material that had usernames removed but could still include highly personal context. 2
According to leaked materials reviewed by 404 Media, OpenAI recruited hundreds of contractors through staffing firms to review real user prompts and ChatGPT responses. The work reportedly paid more than $50 per hour. Reviewers summarized the user’s intent, compared four candidate responses, and scored them on a 1-to-7 scale. 2
The reported objective was to improve response quality—particularly to make ChatGPT sound less robotic, reduce overly human-like claims, and curb excessive agreement or flattery. This is a familiar use of human feedback in AI development, but the reported use of real consumer conversations makes the privacy implications more immediate. 2
404 Media reported that reviewers could receive full exchanges and, in some cases, a user memories summary. The report said usernames were not shown, but the available context could still reveal approximate location, profession, past interests, and personal circumstances. 2
That distinction matters. Removing a name is not the same as making a conversation harmless or anonymous in practice. A combination of contextual details—such as work, family situation, health concerns, or a distinctive event—can be sensitive even if it does not directly identify an account holder.
OpenAI says its Privacy Filter is a model designed to detect and redact personally identifiable information in text. 46
47
In the Project Lily reporting, OpenAI said it removed usernames and used an automated filter before material reached contractors. But 404 Media reported that sensitive details could remain in reviewed conversations despite those measures. 2
This illustrates the limit of automated redaction: a system may recognize obvious identifiers such as a name, email address, or phone number, while subtler identifying context can be much harder to classify reliably.
The reporting describes guidelines aimed at more than basic correctness. Contractors were reportedly asked to identify issues including:
The emphasis on sycophancy is significant. A model that is too eager to validate a user can produce responses that feel supportive but fail to challenge incorrect, risky, or harmful assumptions. Project Lily was reportedly intended in part to help tune that behavior. 2
Writing a request for secrecy inside a prompt does not change how an account’s data controls apply. It is content within the chat, not a technical instruction that overrides OpenAI’s data-use settings.
For personal ChatGPT workspaces, OpenAI’s documented control is Improve the model for everyone. When it is enabled, chats may be used to improve models; when it is turned off, OpenAI says it will not train on new conversations. 21
30
For signed-in users, OpenAI’s instructions are:
This setting applies to the account rather than to one specific device. Your existing chat history can remain available after switching it off, but OpenAI says new conversations will not be used to train its models. 19
30
OpenAI also offers a Do not train on my content option through its privacy portal; the company says either that option or the in-product Data Controls setting is sufficient for ChatGPT and Codex tasks. 21
OpenAI says personal workspaces on Free, Plus, and Pro have data sharing for model training enabled by default, though users can opt out. For ChatGPT Business, Enterprise, Edu, and the API Platform, OpenAI says it does not use provided inputs and outputs to train its models by default. 23
That does not make every business conversation automatically risk-free for every purpose. Organizations should still review their workspace settings, retention practices, sharing controls, and internal policies before placing confidential material into any AI tool.
Project Lily, as reported by 404 Media, is a reminder to treat consumer AI chats as a service with configurable data-use settings—not as an inherently confidential channel. If a prompt contains highly sensitive personal, legal, medical, financial, or proprietary information, the safest immediate step is to disable Improve the model for everyone before starting a new conversation. 2
21
Opting out controls OpenAI’s use of new conversations for model improvement. It is not a request embedded in a prompt, and OpenAI’s documentation does not describe it as a way to retract material already submitted. 21
30
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
404 Media reported that Project Lily used hundreds of contractors, reportedly paid more than $50 an hour, to assess real ChatGPT chats and rate four responses from 1 to 7.
404 Media reported that Project Lily used hundreds of contractors, reportedly paid more than $50 an hour, to assess real ChatGPT chats and rate four responses from 1 to 7. For Free, Plus, and Pro users in a personal workspace, turn off Settings → Data Controls → Improve the model for everyone to prevent OpenAI from using new conversations to improve its models.
Business, Enterprise, Edu, and API inputs and outputs are not used for model training by default, according to OpenAI.