404 Media reported on September 15, 2026, that OpenAI’s “Project Lily” uses hundreds of outsourced contractors to assess real consumer ChatGPT chats, including potentially sensitive context. Reviewers reportedly summarize user intent, compare four candidate replies, and score them from 1 to 7 to help make ChatGPT le...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did the 404 Media investigation published on September 15, 2026, reveal about OpenAI’s secret “Project Lily,” including how hundreds of. Article summary: 404 Media reported that OpenAI’s internal “Project Lily” uses a large outsourced human-feedback operation to improve ChatGPT using real consumer conversations—not merely public web data or synthetic tests. The report is . Topic tags: general, general web, academic, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
404 Media’s investigation into OpenAI’s internal Project Lily describes a human-feedback operation built around real consumer ChatGPT conversations. According to leaked materials, examples of prompts, and statements reviewed by the outlet, outsourced contractors may read chat context to evaluate model responses—even when conversations contain sensitive personal details. 1
For people using ChatGPT for personal, health, relationship, legal, or work-related questions, the central lesson is straightforward: a prompt asking the chatbot to keep something private is not an access-control setting. The relevant control is whether the account permits new conversations to be used for model improvement. 1
404 Media reported that hundreds of contractors, hired through staffing firms and paid more than $50 an hour, review a large flow of real ChatGPT conversations. Their assignments reportedly include:
The reported evaluation criteria focus on improving the feel and quality of ChatGPT’s answers. Reviewers are instructed to flag AI-sounding phrasing, incorrect factual claims, over-validating or sycophantic language, safety concerns, and responses that falsely imply the model has personal experiences or emotions. 1
That feedback loop is designed to make responses sound more natural and less mechanical. But it also means quality improvement can involve human access to the substance of user conversations. 1
OpenAI told 404 Media that usernames are removed and that a Privacy Filter is applied before material is sent to reviewers. The investigation nevertheless found examples in which review materials could include full conversation context and model-generated memory summaries. Those summaries reportedly contained details such as approximate location, profession, and highly personal circumstances. 1
The limitation is contextual identification. A name or account identifier can be removed while a combination of life events, job details, geography, or unusual references still makes a conversation recognizable to someone who knows the person—or simply highly sensitive to expose. 404 Media reported that OpenAI acknowledged filters can have difficulty with uncommon identifiers and ambiguous, context-dependent private references. 1
Telling ChatGPT to keep a discussion confidential is an instruction within the conversation, not a privacy preference for the account. It does not override the service’s data-use settings or prevent a chat from being selected for an improvement workflow described by 404 Media. 1
Users should therefore avoid treating a consumer chatbot as a confidential channel merely because the chat feels personal. If information would be harmful to disclose to an authorized reviewer, it is safer not to enter it in a personal workspace with model improvement enabled.
OpenAI’s Data Controls documentation says personal-workspace users can turn off model improvement through:
With that setting off, new conversations remain in chat history but are not used to train or improve ChatGPT’s models. The setting applies to the account rather than to one specific device.
The timing matters: OpenAI says opting out applies to new conversations. It does not retroactively remove chats that were already submitted while model improvement was enabled.
OpenAI states that it does not use inputs and outputs from ChatGPT Business, Enterprise, Edu, or its API Platform to train models by default. In contrast, model-improvement sharing is enabled by default for ChatGPT Free, Plus, and Pro users in personal workspaces, though they can opt out through Data Controls.
That distinction does not eliminate every organizational data-governance obligation. Teams should still follow their employer’s policies and review the specific workspace settings that apply to them. But it is an important difference from the default treatment of consumer chats.
Human review is not unique to OpenAI. Anthropic’s privacy policy says it may train and improve its models using inputs and outputs supplied by users or crowd workers unless users opt out. An academic review of major AI providers’ privacy practices also found that both Google and OpenAI discuss human review of user chats for model training or improvement.
The broader issue is not whether human feedback exists—it is whether users understand when their conversations can enter that process, what de-identification can and cannot protect, and how to change the applicable settings.
Project Lily, as reported by 404 Media, is a reminder that consumer AI chat is not automatically a confidential one-to-one exchange. De-identification measures may reduce exposure, but they cannot necessarily remove every revealing detail from a conversation rich in personal context. 1
Before sharing sensitive information in ChatGPT, check Settings → Data Controls → Improve the model for everyone. If model improvement is enabled, treat new chats as potentially eligible for use in improving the service; if it is disabled, OpenAI says new chats will not be used for that purpose.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
404 Media reported on September 15, 2026, that OpenAI’s “Project Lily” uses hundreds of outsourced contractors to assess real consumer ChatGPT chats, including potentially sensitive context.
404 Media reported on September 15, 2026, that OpenAI’s “Project Lily” uses hundreds of outsourced contractors to assess real consumer ChatGPT chats, including potentially sensitive context. Reviewers reportedly summarize user intent, compare four candidate replies, and score them from 1 to 7 to help make ChatGPT less robotic, less sycophantic, and more reliable.
OpenAI reportedly removes usernames and applies a privacy filter before review, but 404 Media found that contextual details and memory summaries can still expose sensitive personal information.