OpenAI reportedly removed contractors who used AI to help evaluate its models. The issue was not simply that an AI company’s workers used AI: these assignments were intended to produce independent human judgments about ChatGPT’s responses. Reporting connects the removals to OpenAI’s broader contractor evaluation work, while identifying Project Lily as one program built around human review; it does not establish that every removed worker was assigned to Lily.
16
13
How Project Lily’s human ratings work
Project Lily reportedly uses hundreds of outside contractors to review real ChatGPT prompts and, in some cases, the surrounding conversations. Reports describe reviewers recruited through Crossing Hurdles and paid through Mercor. One worker reported earning more than $50 an hour; that is not a verified pay rate for every reviewer.
2
20
A reviewer summarizes what the user was trying to accomplish, then critiques and scores candidate ChatGPT replies on a 1–7 scale. Accounts describe tasks with up to four candidate replies and an emphasis on qualities such as tone, robotic phrasing and excessive agreement. The scores and written critiques are intended to help improve future model behavior.
4
11
3
1
Why AI-assisted ratings defeat the purpose
A human reviewer can bring an independent perspective to whether an answer addresses the user’s request or sounds too eager to agree. If the reviewer instead lets another AI write the critique or choose the score, that independence is weakened: the feedback may reflect model-generated preferences rather than a person’s assessment. That is a risk to the value of the feedback, not proof that any particular AI-assisted rating harmed a model.
4
16
Reported instructions prohibit reviewers from using AI to review work or write feedback, including Grammarly and AI translation features. They also bar reviewers who audit other contractors from using AI-detection tools, which the instructions describe as unreliable.
13
9
How supervisors spot suspected AI use—and what they cannot know
Reports describe supervisors looking for patterns such as repetitive, AI-like phrasing or unusually fast completion instead of relying on detector scores. Those clues can justify a closer look, but they are not proof by themselves. The available reporting does not show how often supervisors correctly identify AI use or miss it.
9
33
The same distinction matters for ratings made without AI. A person could still choose a poor score deliberately or do careless work. The provided reporting does not verify the specific alleged admission of deliberately choosing poor ratings, so its circumstances and prevalence cannot be established here. The broader lesson is narrower but important: a human-rated label describes who supplied the feedback, not whether each judgment is sound. Reliable evaluation also depends on checking the quality of the judgments themselves.
1
4
Project Lily raises a separate concern for users: reviewers may see real conversation text, and OpenAI’s reported privacy filter may fail to remove some sensitive details. That makes the integrity of both the review process and its handling of user conversations consequential.
8
12