Inherent claims Faraday, built on a 27 billion parameter Qwen 3.6 model, outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT 5.5 at independently reproducing published research results. Faraday’s distinctive role is research direction: reinforcement learning was used to teach judgment about which experiments t...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Inherent, a London-based AI startup founded by Google DeepMind alumni with about a dozen employees and a recent $50 million seed ro. Article summary: Inherent claimed that Faraday outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at independently reproducing results from published scientific papers when the agent was not given the original answer in advanc. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Inherent, a London AI lab founded by former Google DeepMind researchers, says its newly released Faraday agent beat Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at a focused task: independently reproducing findings from published scientific papers without being given the answer in advance. The result is notable because Faraday is built around Qwen 3.6, a comparatively small 27-billion-parameter model, rather than a frontier-scale system. 25
The important qualification is scope. This is Inherent’s reported result on a paper-replication evaluation—not an independently established claim that Faraday is a generally more capable model than Claude or GPT-5.5.
The comparison concerned scientific replication: an agent receives a published study and must reconstruct enough of the work to reproduce its findings, rather than simply summarize the paper or predict a known answer. Inherent says Faraday performed better than Claude Opus 4.8 and GPT-5.5 in that setting. 5
A technical account of the evaluation describes Replica as a 310-task benchmark focused on reproducing scientific figures. 8 That framing matters because success in a tightly defined research workflow can reveal useful capabilities without establishing performance across general reasoning, coding, writing, or scientific discovery.
Faraday runs on Qwen 3.6 with 27 billion parameters. Inherent presented that as a substantial model-size contrast with the larger systems used for comparison. 15
The lesson Inherent is advancing is not that Qwen 3.6 alone has replaced the leading general-purpose models. Faraday is an agent system: its outcome depends on the underlying model, post-training, research workflow, tool use, and the way responsibilities are divided between planning and execution. A smaller foundation model can therefore be competitive on a specialized task when the surrounding system is optimized for that task.
That distinction is central to the company’s strategy. Instead of trying to train the largest general-purpose model, Inherent is betting that a compact model with better scientific direction can coordinate stronger specialist tools and make more useful experimental decisions.
Inherent describes the target capability as “research taste”: the judgment to decide which experiments are worth running, how to design them, when results are inconclusive, and when to change direction. 4
The company says it used reinforcement learning to develop this behavior. Rather than prescribing every step, the training approach rewards the quality of the eventual research process and outcome. That is a different objective from optimizing only for a short-answer benchmark, where the correct response is already defined.
Faraday also does not need to perform every implementation step itself. It can use coding agents—including OpenAI’s GPT-5.5 Codex—as tools. In this arrangement, Faraday acts more like a research director: it supplies scientific planning and judgment, while a specialist coding system helps turn that plan into executable experiments. 6
This architecture makes the comparison more nuanced. Faraday’s reported advantage belongs to the complete agent workflow, not necessarily to its 27B base model in isolation.
Replication gives an AI researcher a concrete way to practice science. The agent must interpret a paper, form hypotheses about the underlying method, choose experiments, deal with failed attempts, and compare its results with the published findings.
Inherent’s premise is that the visible paper captures only a small part of the work. The trial and error, discarded approaches, and practical choices that led to the final result are mostly absent. Reconstructing that hidden process can therefore serve as a controlled training ground for scientific reasoning. 5
It is also easier to evaluate than truly original discovery. A replication task has an external target: the agent can be assessed on whether it successfully reproduces the reported result and whether its experimental decisions are scientifically defensible. That makes it a useful intermediate step before asking an AI system to investigate questions whose answers are not yet known.
The company does not describe paper replication as the final destination. It frames Faraday as an early step toward AI scientists that can work across scientific fields and eventually help produce genuinely new knowledge. 13
That is a much larger claim than winning a benchmark. A system that can reproduce a published result demonstrates useful research behavior, but original discovery also requires identifying valuable unanswered questions, generating novel hypotheses, designing reliable tests, and producing results that survive scrutiny. The Faraday evaluation is best understood as evidence for one part of that trajectory, not proof that the broader goal has already been reached.
Inherent emerged from stealth with a reported $50 million seed round and is building in London, close to the talent ecosystem that produced its founders. 5 The company’s reported plan to grow from roughly a dozen people to about 20–25 suggests a small, research-intensive approach rather than a race to match the headcount or model-training budgets of the largest AI labs.
Co-founder and chief scientist Edward Hughes has criticized the U.K.’s garden-leave restrictions, which can limit how quickly researchers move between employers or join startups. Inherent views that as a competitive issue when recruiting experienced AI researchers, including talent from DeepMind, against better-funded companies. 6
The resulting strategy is concentrated rather than scale-driven: use a lean team, invest in reinforcement learning and scientific judgment, and combine a smaller base model with increasingly capable external tools. If that approach works beyond the initial evaluation, it could offer startups another way to compete in AI research without building the biggest model in the market.
Inherent says Faraday outperformed GPT-5.5 and Claude Opus 4.8 at independently replicating published scientific findings, despite using a 27B Qwen 3.6 model. The more consequential idea is not simply that a smaller model won a comparison; it is that research planning, long-horizon reinforcement learning, and tool coordination may matter as much as raw model scale for scientific work.
For now, the evidence supports a narrower conclusion: Faraday is a promising research-agent system on the task Inherent evaluated. It does not yet establish broad superiority over Anthropic’s or OpenAI’s models, nor does replication by itself demonstrate autonomous scientific discovery.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Inherent claims Faraday, built on a 27 billion parameter Qwen 3.6 model, outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT 5.5 at independently reproducing published research results.
Inherent claims Faraday, built on a 27 billion parameter Qwen 3.6 model, outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT 5.5 at independently reproducing published research results. Faraday’s distinctive role is research direction: reinforcement learning was used to teach judgment about which experiments to run and how to design them, while coding agents such as GPT 5.5 Codex can handle implement...
Inherent sees replication as training for AI scientists, with a longer term goal of helping generate genuinely new scientific knowledge across fields.