Inherent says its 27 billion parameter Faraday agent outperformed Claude Opus 4.8 and GPT 5.5 on research replication tasks, but the result does not yet show that Faraday can independently make new scientific discover... Faraday was trained with long horizon reinforcement learning on Replica, a benchmark that asks a...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Inherent, a London AI lab founded by Google DeepMind alumni, announce about its 27-billion-parameter Faraday AI agent’s ability to. Article summary: Inherent says its 27-billion-parameter Faraday agent can independently reproduce findings from published scientific papers without being given the answer beforehand, and that it outperformed Anthropic’s Claude Opus 4.8 a. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Inherent, a London AI lab founded by Google DeepMind alumni, says its 27-billion-parameter Faraday agent can reproduce findings from published scientific papers without being given the answers in advance. The company also says Faraday outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 on that specific research-replication task. 259
That is a notable result if it holds up under broader, independent testing—but it is narrower than the headline may suggest. Faraday has been evaluated on reproducing existing results, not on independently discovering new scientific knowledge.
Inherent created Replica, a benchmark built from 100 machine-learning and AI-for-science papers and converted into 310 replication tasks. Each task asks an agent to recreate a particular figure from a paper while working within limited time and computing resources and without access to the original plot. 34
The challenge is therefore more demanding than simply asking a model to describe a paper or copy a published chart. The agent must interpret the research, implement the necessary experiments, decide how to use its available resources, and produce a result that matches the reported finding closely enough to count as a replication.
Inherent says Faraday surpassed Claude Opus 4.8 and GPT-5.5 on this benchmark despite using a much smaller 27-billion-parameter model. 158 The comparison should be read as a benchmark-specific claim, not as a general ranking of the systems across all scientific or coding work.
Inherent’s argument is that successful scientific work depends on more than obtaining a final result. A capable researcher must decide which experiments are worth running, design them sensibly, interpret ambiguous outcomes, and change direction when an approach fails. The company describes this collection of judgments as “research taste.” 47
That distinction matters because a model can potentially produce a plausible-looking output without following a reliable scientific process. Inherent says Faraday is trained to develop the longer-horizon behavior behind a good investigation: forming hypotheses, testing them, learning from failures, and selecting productive next steps. 4
The company sees replication as a useful training ground for this behavior. Reproducing a published result requires recovering much of the practical trial and error that a paper may not fully document. Inherent’s broader thesis is that learning to navigate this process could help agents progress from verifying known findings toward more open-ended research. 49
Faraday is not positioned as a standalone replacement for a powerful coding agent. Instead, it acts as a supervisory scientific layer that directs OpenAI’s GPT-5.5 Codex during replication tasks. Inherent compares this arrangement with a human scientist using a coding assistant: Faraday supplies planning and scientific judgment, while Codex handles much of the implementation work. 39
This architecture shifts the focus away from model size alone. A smaller model can potentially add value by deciding what the larger tool should do, when to change tactics, and how to evaluate the results. Faraday’s reported advantage is therefore partly an orchestration advantage—the combination of scientific planning and frontier coding capability—rather than evidence that its 27-billion-parameter model is universally more capable than larger systems.
Inherent says it wants to build AI agents that collaborate with researchers as teammates. The intended system would help plan, execute, critique, and iterate on experiments while people retain direction and oversight. 45
Faraday is best understood as an early version of that idea: a model trained to provide scientific intuition and research management on top of general-purpose coding tools. The company’s stated long-term ambition is to develop agents that can contribute to discovering new knowledge, but the current result demonstrates replication rather than discovery. 34
That boundary is important. Reproducing a known result is a valuable test of scientific reliability, yet it does not by itself establish that an agent can formulate important original questions, generate genuinely novel hypotheses, or validate discoveries in the real world.
Inherent emerged from stealth with a reported $50 million seed round and is presenting itself as a focused London research lab founded by Google DeepMind alumni. 25 Its approach is to concentrate on the parts of AI-assisted science that general-purpose models may not learn automatically: experimental judgment, long-horizon decision-making, evaluation, and coordination between humans and software tools.
That strategy also explains why the company highlights Faraday’s use of an external coding system. Rather than building every capability in one model, Inherent is attempting to train a specialized layer that makes existing frontier tools more effective for scientific work. 39
The immediate takeaway is measured but significant: Inherent claims that a relatively small, reinforcement-learning-trained supervisory model can outperform larger frontier systems on a carefully defined replication benchmark. The larger question—whether that approach can produce dependable, original scientific contributions—remains open.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Inherent says its 27 billion parameter Faraday agent outperformed Claude Opus 4.8 and GPT 5.5 on research replication tasks, but the result does not yet show that Faraday can independently make new scientific discover...
Inherent says its 27 billion parameter Faraday agent outperformed Claude Opus 4.8 and GPT 5.5 on research replication tasks, but the result does not yet show that Faraday can independently make new scientific discover... Faraday was trained with long horizon reinforcement learning on Replica, a benchmark that asks agents to recreate figures from published machine learning and AI for science papers without seeing the original plots.
Its design pairs a smaller model that supplies scientific planning and judgment with GPT 5.5 Codex for coding intensive execution.