OpenAI reportedly removed multiple contractors for using AI to perform ChatGPT evaluation work because Project Lily is meant to supply independent human judgments, not model generated feedback. Project Lily reportedly uses hundreds of contractors to read and assess real ChatGPT conversations, while reporting describ...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: Why did OpenAI fire contractors working on Project Lily for using AI tools to evaluate anonymized ChatGPT conversations, how do its more tha. Article summary: OpenAI reportedly removed contractors who used generative AI to produce the judgments because Project Lily’s purpose is to collect independent human preferences. AI-written evaluations would turn a human-feedback pipelin. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Project Lily, as reported by 404 Media, is a human-review operation in which contractors assess ChatGPT responses to help improve future behavior. That purpose explains the reported removals of contractors who used generative AI to do the evaluation: the value of the data is supposed to come from independent human judgment. If a model supplies the judgment instead, the pipeline no longer cleanly measures what people prefer. 19
20
The reported workflow gives reviewers real ChatGPT prompts and, in some cases, full conversations. They summarize the user’s intent, compare candidate responses, and assign scores on a 1–7 scale. Reporting describes up to four candidate answers per task, with emphasis on behavioral qualities such as tone, excessive flattery or sycophancy, and overly human-like claims—not simply factual accuracy. 8
19
Those ratings can become preference data: a structured record of which response a person considers better for a particular context. Such data is useful for tuning a model toward desired behavior only when the judgments are sufficiently consistent, informed, and independent.
Project Lily reportedly involves hundreds of outside reviewers. They are described as being recruited through Crossing Hurdles and paid through Mercor, with one reviewer reporting compensation above $50 an hour. Separately, reporting on OpenAI’s broader evaluation operation refers to more than 10,000 contractors across its rating pipelines; that figure should not be read as the size of Project Lily alone. 19
6
An AI evaluator may be fast and sometimes useful for internal testing, but it answers a different question from a human preference review. A model tends to reproduce patterns learned from its training and may favor responses with familiar phrasing, style, or assumptions.
If output from a model, or a closely related model, is repeatedly used to reward future model outputs, a feedback loop can emerge:
That does not prove every use of AI evaluation makes a model worse. Automated evaluation can be valuable when it is validated against human judgment and used with clear limits. But substituting it for a task commissioned specifically as human preference data weakens the independence of the training signal—and that is the apparent reason the reported policy prohibited it.
The reporting says supervisors were instructed not to rely on AI-writing detectors such as GPTZero. Instead, they were to look at work patterns and the written output through human review, including signals such as unusually uniform phrasing, punctuation or formatting patterns, and implausibly rapid completion. 6
That approach has an obvious limitation: none of these signals alone proves that a worker used AI. A fast writer may be experienced; similar language may reflect a shared rubric. The credible use of such signals therefore depends on review processes that distinguish a prompt for investigation from evidence strong enough to remove someone from a project.
Organizations generally avoid publishing detailed anti-gaming criteria because a precise checklist can become an instruction manual for evasion. Contractors could alter wording, add delays, or change formatting while still outsourcing judgments to a model.
However, secrecy creates its own risk. When workers do not know how a quality decision was reached, mistakes are harder to challenge. A durable integrity program needs confidential detection methods alongside clear rules, documented decisions, and a meaningful appeal or correction process.
Mercor reportedly said that using AI to perform this evaluation work violates policy and that confirmed cases result in immediate removal from the relevant project. Reporting also says the rules barred AI assistance more broadly, including tools used to write feedback or comments. 6
40
The reported case of deliberate poor ratings makes the larger point: “human feedback” is not automatically reliable simply because a human submitted it. A reviewer may rush, misunderstand the rubric, optimize for throughput, act in bad faith, or try to game opaque quality systems.
This is an incentive problem as much as a technical one. The model developer needs calibrated, good-faith judgments that predict better real-world behavior. A contractor may instead be rewarded mainly for speed, task completion, or avoiding rejection. When those incentives diverge, the resulting data can be noisy or systematically distorted.
Human review remains important for judgments that are contextual, subjective, or difficult to verify automatically. But it works best as a system of checks rather than a simple stream of individual scores. Practical safeguards include:
The central lesson is straightforward: preference training is only as trustworthy as the process that creates its preferences. Human reviewers are not a magic guarantee of quality, and AI-generated judgments cannot simply stand in for independent human feedback when the goal is to learn what people actually value.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI reportedly removed multiple contractors for using AI to perform ChatGPT evaluation work because Project Lily is meant to supply independent human judgments, not model generated feedback.
OpenAI reportedly removed multiple contractors for using AI to perform ChatGPT evaluation work because Project Lily is meant to supply independent human judgments, not model generated feedback. Project Lily reportedly uses hundreds of contractors to read and assess real ChatGPT conversations, while reporting describes more than 10,000 contractors across OpenAI related rating pipelines.
The episode highlights a basic alignment challenge: a rating entered by a person is not necessarily careful or independent, so training pipelines need quality checks as well as human reviewers.