Sampura Research launched publicly on August 25, 2026—not August 24—and is developing hybrid human–AI systems to evaluate increasingly capable models. Its central argument is that AI only oversight can miss novel failures, be manipulated by the systems it evaluates, or fail in ways that overlap with human mistakes.
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Sampura Research, the London-based nonprofit formally launched on August 24, 2026 by former Google DeepMind researchers Rishub Jain,. Article summary: Sampura Research is a London-based AI-safety nonprofit developing “human–AI complementarity for scalable oversight”: systems in which people and AI jointly evaluate increasingly capable models rather than delegating over. Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
Sampura Research is a London-based AI-safety nonprofit focused on a practical question: who should judge whether an increasingly capable AI system is behaving correctly and safely? Its answer is not to remove people from the loop, but to design systems in which humans and AI share oversight responsibilities. The organization says it has received an initial $11 million grant from Coefficient Giving, consisting of $7 million for its first year and a further $4 million pledged. 1
The public launch announcement is dated August 25, 2026, rather than August 24. It identifies Rishub Jain and Josh Jacob as the founders. Earlier social-media posts described work involving Alex Adams, but the launch announcement does not establish that all three formally founded the nonprofit. 112
In Sampura’s terminology, a judge can be a person, an AI model, or a hybrid system that assesses whether another AI system’s behavior is correct and aligned. The behavior being assessed might be a single answer, a conversation, or a multi-step agent trajectory. 1
That distinction matters because oversight is often treated as a separate problem from model capability. A model may produce fluent answers or complete complicated tasks while still misleading its evaluator, exploiting gaps in a reward function, or taking an unsafe action during a longer sequence of steps. Sampura’s research agenda is aimed at improving the systems that detect those failures.
The organization describes this as human–AI complementarity for scalable oversight: using automation where it is reliable while preserving human involvement where judgment, interpretation, or value-sensitive decisions are needed. 1
AI systems are attractive as overseers because they are faster and cheaper to run at scale. But Sampura argues that speed alone does not make an evaluator trustworthy. Its launch statement highlights several risks:
These concerns do not prove that human judges are always better. They explain why Sampura wants to test which combination of human and machine evaluation works best under pressure, rather than assuming that either humans or AI should handle every case alone. 1
Sampura’s initial six-month programme centers on a diverse judge leaderboard containing more than 20 subdatasets. The planned areas include deception detection, cultural bias, and unsafe actions by AI agents. 1
The proposed evaluations go beyond static question-and-answer tests. Sampura also plans scenarios in which the supervised model directly optimizes against the judge during reinforcement learning. That setup is intended to reveal whether a judge continues to work when the system being evaluated has an incentive to find and exploit weaknesses. 1
The research will compare three broad approaches:
The goal is not simply to achieve a high score on one narrow benchmark. Sampura says it wants judges that remain robust across a broad set of tests and continue to work as the models they supervise become better at exploiting evaluation weaknesses. 1
The organization has outlined several possible designs for combining human and AI judgment:
This approach treats human involvement as something to optimize, not as an all-or-nothing requirement. Some cases may be cheap and reliable for automated review; others may justify slower human attention. The relevant trade-off is therefore not “human versus AI,” but the combination of accuracy, cost, speed, and resilience to deliberate gaming. 1
After developing stronger judges, Sampura says it plans to test whether they can reduce reward hacking during model training. It also wants to study how much human involvement is needed, how oversight systems should adapt when base models learn to exploit them, and whether judges can be used for more than classifying outputs. 1
That broader work could include drafting or assessing the specifications, instructions, rubrics, and test cases used to train and evaluate AI systems. Sampura says it intends to publish its work through papers, code, datasets, and leaderboards. 1
The available launch material supports this general dataset, benchmark, routing, and human-assistance programme. It does not establish every more specific project label associated with the organization, including a named “forward-looking harm” dataset or a confirmed open-source human-rating platform. Those descriptions should therefore be treated as unverified rather than as announced products.
Sampura says it is hiring founding technical staff in London. Its programme requires expertise spanning machine learning, evaluation design, human–computer interaction, data operations, and deployment. 1 Earlier public posts referred to recruiting at least six additional researchers and quoted salaries of £100,000 to £290,000, but those exact figures are not confirmed by the nonprofit’s launch announcement. 1213
The broader AI-evaluation field is also competing for a limited pool of specialists. METR’s own listing for a Member of Technical Staff, Evaluation Execution role gives a base salary range of $285,548 to $503,116 and says it may consider paying more for exceptionally experienced candidates. 18
That contrast illustrates the challenge Sampura faces: building reliable oversight requires both technical research and careful real-world evaluation, while the organizations working on those problems are competing for many of the same people.
Sampura’s launch arrives amid broader concern that AI development could outpace the systems intended to monitor it. In July 2026, more than 1,100 employees from major AI companies signed an open letter asking the U.S. government to support international technical and governance tools that could deliberately pace frontier AI development if necessary. The letter did not call for an immediate halt to AI development. 40
There has also been significant movement among senior AI researchers. Reuters reported leadership changes at Google’s AI division and departures of prominent engineers and researchers to other companies and startups. 49 Those developments provide context for Sampura’s formation, but they do not demonstrate that the nonprofit has solved the underlying safety problem.
The strongest evidence supports Sampura as an early-stage research organization building measurement and decision infrastructure for AI oversight. Its benchmarks and hybrid evaluation methods could eventually be useful to AI labs, independent evaluators, or other organizations testing model behavior. 1
That is different from proving that Sampura can contain frontier systems. Its immediate contribution is to investigate how oversight should be designed and tested when the model being supervised is capable, adaptive, and potentially incentivized to pass its evaluations. The central bet is that scalable safety will require automation—but not automation that removes humans from the most consequential judgments.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Sampura Research launched publicly on August 25, 2026—not August 24—and is developing hybrid human–AI systems to evaluate increasingly capable models.
Sampura Research launched publicly on August 25, 2026—not August 24—and is developing hybrid human–AI systems to evaluate increasingly capable models. Its central argument is that AI only oversight can miss novel failures, be manipulated by the systems it evaluates, or fail in ways that overlap with human mistakes.
The nonprofit’s first six months will focus on a judge leaderboard with more than 20 subdatasets, comparing human, AI only, and hybrid evaluators.