The aggregate laboratory result was similar: 354 of 1,320 designs bound their intended targets, an overall hit rate of 26.8%. Only one evaluated target produced no binders.
Those numbers should not be treated as a universal success rate. Performance varied substantially by target, with reported per-target results ranging from zero to 90%. A focused run in which Mythos worked on individual targets for 24 hours produced the highest reported rate, 35.1%, while simultaneous multi-target campaigns produced lower rates.
The experiment was not a language model inventing an entirely new field of protein science. A human protein-design specialist first provided an approximately 30,000-token prompt describing the workflow. Claude then operated with minimal further human steering.
Within that framework, Claude:
This distinction matters. The achievement was the ability to manage a long, tool-using design process and make the individual candidate-selection decisions, rather than the invention of every underlying computational method.
The target set also included established benchmark targets and newer competition targets. That makes the result more informative than a purely synthetic computational exercise, although it does not remove the possibility that benchmark familiarity influenced performance.
A computational model can predict that a protein should bind without producing a protein that actually expresses, folds or binds in an assay. Anthropic therefore sent its designs to external testing partners.
Adaptyv Bio received the designs without knowing which model had produced them. Its automated process moved from digital protein sequence to DNA synthesis, protein expression, binding measurement and quality control. The company measured binding using surface plasmon resonance and reported 354 binders among 1,320 designs.
Twist Bioscience also independently produced and tested designs. Anthropic says the results included high-affinity binders for at least six targets and binders that matched or exceeded the best previously reported affinity for at least four.
This is stronger evidence than a benchmark based only on predicted structures or model-written explanations. But the validation still answers a narrow question: does the designed protein bind the selected target under the test conditions? It does not establish that the protein is a useful therapeutic.
One of Anthropic’s most notable comparisons involved RBX1. Anthropic reported an approximately 40% hit rate for Claude’s designs, compared with 3.7% among earlier competition entrants, and said its system outperformed the previous winning design on that target.
Anthropic also reported unusually strong affinity results for some candidates, including designs that bound several times more tightly than the best previously published results. These are promising findings, but they are still company-reported results from a single campaign rather than proof of broad superiority across protein-design tasks.
A protein binder is a molecule engineered to attach to a biological target. That can be useful for therapeutics, diagnostics and research reagents, but binding is only an early milestone.
A viable drug candidate may also need the right:
Anthropic describes the work as an early step toward drug-like molecules, not as an end-to-end drug-discovery result. The campaign did not demonstrate clinical efficacy, safety in people or successful progression through drug development.
The target selection is an important limitation. Protein binders are generally more practical against accessible, extracellular targets than against proteins operating inside cells. Martin Shkreli criticized the reported affinities as too weak to support the broader pharmaceutical implications and pointed to the lack of difficult intracellular targets.
Those criticisms do not invalidate the laboratory finding that many designs bound their targets. They do, however, narrow the appropriate conclusion. The experiment showed that Claude can produce a substantial number of experimentally confirmed binders under a defined workflow. It did not show that the system can reliably create intracellular therapeutics or outperform established drug-development methods across clinically important targets.
The evidence also remains partly company-reported. The physical synthesis and assays were conducted by external organizations, which improves the credibility of the result, but independent replication across broader target sets would provide a stronger test of generality.
Anthropic presents protein-binder design as one component of a larger ambition: an end-to-end, all-modality system that can help accelerate drug development and analyze chemical and biological data. Its broader Claude Science effort combines research workflows, computational tools and scientific data analysis, while the protein campaign demonstrates how an agent might connect those capabilities to experimental work.
Anthropic says its most capable life-science capabilities are not currently available for unrestricted use. The company has identified a scientist-access program as a priority, and says Opus 5 is currently its most capable generally available model for life-science research.
The practical promise is not that an AI system eliminates the laboratory. It is that autonomous systems could reduce the time spent researching targets, running software, triaging candidates and deciding which relatively small set of designs deserves experimental attention.
The workflow has a dual-use dimension. The same ability to research biological systems, design proteins and coordinate specialist tools could support beneficial research—or lower the expertise and time required for harmful biological work.
Anthropic says protein design and other dual-use biological capabilities remain unavailable for general access on its most capable systems, with controlled or trusted access intended for qualified scientists. The company’s stated challenge is to support legitimate therapeutic research while limiting assistance that could contribute to pathogen engineering, toxin-related work or other dangerous biological applications.
That policy is not separate from the technical result. The more autonomously a model can move from a biological question to candidate molecular designs, the more important it becomes to evaluate not only whether the designs work, but also who can access the workflow, what safeguards surround it and how its use is monitored.
Anthropic’s demonstration showed that Claude Mythos Preview and Claude Opus 4.8 can autonomously coordinate existing protein-design tools and produce experimentally confirmed binders at a reported 22.6% to 35.1% hit rate. Across 1,320 designs, 354 bound their targets, and 14 of 15 evaluated targets produced at least one binder.
The result is a significant proof point for AI-assisted early protein design. But it is not an autonomous drug-discovery system, a solution to intracellular targeting or a substitute for the long process of optimization, toxicology, clinical testing and regulatory review. The next test is whether these results replicate across harder and less familiar targets—and whether the capabilities can be made broadly useful without making dangerous biological assistance broadly available.