Hero image for Anthropic's Claude Just Beat Chemists at Drug Design
By AI Tool Briefing Team

Anthropic's Claude Just Beat Chemists at Drug Design


Anthropic published results on August 18 showing that Claude designed working protein binders (the molecules that latch onto a disease-relevant protein and either block it or flag it for the immune system) for 14 of 15 tested targets, with every design physically synthesized and measured in an outside lab, not simulated. That’s the detail that separates this from a benchmark score. Adaptyv Bio and Twist Bioscience built the proteins and ran the binding assays themselves, blind to which sequences came from Claude.

Every Anthropic story we’ve covered this year has been about something adjacent to the model: a valuation, a security incident, an export-control order. This is the first one where the news is a model doing science and having it check out.

Quick Summary: What Happened

DetailInfo
PublishedAug 18, 2026, via Anthropic’s research blog
Models usedClaude Mythos Preview and Claude Opus 4.8 for protein design; Claude Opus 5 for a separate analytical-chemistry test
Targets tested15 disease-relevant proteins, including PD-L1, TREM2, TNFα, and EGFR
SuccessWorking binders produced for 14 of 15 targets; zero binders against one target (maltose-binding protein)
Hit rate, combined session26.7% (Mythos Preview), 22.6% (Opus 4.8) across all 15 targets in one 48-hour run
Hit rate, single-target sessions35.1% (Mythos Preview), one target per 24-hour run
Industry baseline10–15% typical hit rate for de novo protein design campaigns
ValidationIndependent wet-lab synthesis and testing by Adaptyv Bio and Twist Bioscience

Bottom line: Claude didn’t just predict which proteins might bind. It designed them, they got built by someone else, and in a blind lab test they worked more often than the field’s own baseline.

What Actually Happened

Anthropic runs a research effort called Claude Science, which it launched at the end of June to point Claude at problems with a physical, checkable answer instead of a chat window. Protein binder design is a good test case for that, because “does it work” isn’t a matter of opinion. You synthesize the sequence, expose it to the target, and measure whether it sticks.

Across 15 targets pulled from prior public design competitions and benchmark sets (so there was already a human baseline to compare against), Claude Mythos Preview and Claude Opus 4.8 produced 1,320 total protein designs. Of those, 354 bound their targets when tested, for a blended hit rate of roughly 26.8%. Split by model: 26.7% for Mythos Preview and 22.6% for Opus 4.8 when each worked all 15 targets in a single 48-hour session. Give Mythos Preview 24 hours per target instead of splitting attention across all 15, and the rate climbs to 35.1%. Either number roughly doubles the 10–15% hit rate Anthropic cites as typical for this kind of campaign.

One target produced nothing. Every design Claude generated against maltose-binding protein failed to bind. Anthropic published that alongside the wins, which is the right call on a site that’s skeptical of hit-rate stats by default. A 14-for-15 record with the loss disclosed is a lot more credible than a 15-for-15 record would have been.

What Did Claude Actually Achieve, Target by Target?

  1. RBX1 — Mythos Preview hit 40% against this target in an Adaptyv Bio design competition, versus 3.7% for the human entrants in that same competition. Claude’s best design bound at roughly 3.9 nM, about ten times tighter than the human-designed competition winner’s 25.7 nM.
  2. 15-PGDH — Claude’s top binder improved on the previously best-published affinity by more than 50x, from 1.7 µM down to 33.4 nM.
  3. TREM2 — an 80% hit rate, more than double the 38.3% recorded in the prior competition for the same target.
  4. At least six targets overall — Claude produced at least one high-affinity binder (dissociation constant under 10 nM).
  5. At least four targets — Claude’s best design matched or beat the strongest previously published affinity, not just a “good enough to count” threshold.
  6. Maltose-binding protein — zero confirmed binders, the one target where nothing worked.

That’s a legitimate answer to “did AI beat expert human designers here,” and the honest version is: on some targets, by a wide margin; on one target, not at all.

Anthropic ran a second, smaller experiment alongside the protein work: pointing Claude Opus 5 — reviewed on this site for its Fast Mode pricing cut as the prior-generation model — at raw NMR and LC-MS output, the kind of spectral data a chemist normally interprets by hand after a molecule comes out of synthesis. Claude processed the NMR read in 23 minutes and the LC-MS read in 19, landing within 0.08 of the lab’s hydrogen count and calling purity at 96.4% against the lab’s own 96.33% figure. It’s a narrower claim than the protein work, but it’s the same theme: matching a specific, checkable number a trained chemist would otherwise produce by hand.

Why This Matters

Absent from either experiment: Claude Fable 5, Anthropic’s most capable model, and the one this site covered getting pulled worldwide under a Commerce Department export control order back in June. Anthropic kept Fable 5 out of the protein work on purpose, over dual-use biosecurity concerns — the same capability that speeds up finding a binder for a cancer target could, differently applied, speed up finding something nobody wants found faster. It’s the same Mythos lineage this site flagged in April as too dangerous to release without restriction for its cyber capabilities, now doing the thing biosecurity researchers worry about most, in a different domain — with Anthropic using the preview checkpoint rather than the frontier one.

The other thing worth sitting with: the bottleneck here isn’t the AI anymore. Generating 1,320 candidate sequences is cheap for a model — Anthropic disclosed the 48-hour sessions ran on up to 12,500 Nvidia H100 hours, a real compute bill but a trivial one next to what synthesizing and testing those sequences in a wet lab costs. How fast this actually accelerates drug discovery depends on how fast Adaptyv Bio, Twist Bioscience, and labs like them can physically build and test what the model proposes, not on how fast the model can propose it.

Anthropic itself doesn’t oversell this, and the caution is worth quoting directly: “protein binders are not a standard therapeutic modality,” and “designing a high-affinity binder is just the first step” toward an actual drug-like molecule. A binder that sticks to a target in a plate isn’t a treatment. It’s the first gate in a pipeline that still runs through toxicity, delivery, manufacturing, and years of trials most candidates never survive.

Not everyone’s convinced this deserves the attention it’s getting. Martin Shkreli — the former pharma executive whose own credibility on drug development claims is its own separate conversation — called the results unimpressive, arguing hitting a binding target is a long way from anything resembling a drug. He’s not wrong about the distance. He also doesn’t dispute the actual hit-rate numbers, which came from two labs with no stake in Anthropic’s narrative.

What Are Your Options Now

If you work in biotech or pharma R&D, the practical move isn’t “go try Claude Science” — it’s not a self-serve product, it’s a research collaboration Anthropic ran with two named partners. Watch whether Adaptyv Bio and Twist Bioscience turn this into something buyable. Twist in particular is a public company with a manufacturing pipeline already built for this kind of work, which is why its own investor coverage picked this up as material news rather than a research curiosity.

If you’re evaluating AI vendors for scientific or technical work broadly, the signal worth noting isn’t the protein result specifically — it’s that Anthropic keeps choosing to publish results a third party can independently falsify, the same posture we flagged covering OpenAI’s own capability disclosures around Astra this month. A vendor that lets an outside lab blind-test its output and publishes the miss with the hits is giving you a better read on its self-reporting than one that only shows the wins.

If you’re just tracking what Claude is actually good at, file this next to the coding and agent-orchestration story, not instead of it. Nothing here changes how you’d use Claude for a spreadsheet or a codebase today.

The Bigger Picture

This lands in the same month Anthropic has been racing toward a potential $2 trillion IPO on a revenue curve that’s mostly enterprise API and Claude Code usage. A verified science result doesn’t move that number. What it does is give Anthropic a second pitch besides “our coding agent is good,” aimed at a buyer that’s been much slower than software to adopt frontier models — largely because “the AI said so” has never been good enough in a field where verification is expensive and mandatory.

That’s also the honest limit of how far this generalizes. AI has moved fastest in math and code because an answer can be checked in seconds by running it. Protein design sits at the opposite end — checking an answer means weeks and real dollars in a wet lab. This result matters because Anthropic paid that cost, twice, through independent partners, and still came out ahead of the human baseline. It doesn’t mean the next expensive-to-verify domain falls the same way. It means this one did, once, with disclosed limitations.

Our Take

We think the methodology is the actual story, more than the headline hit-rate numbers. Anthropic didn’t grade its own homework — it handed sequences to Adaptyv Bio and Twist Bioscience blind, let them synthesize and test independently, and published the one target where Claude produced nothing alongside the targets where it beat human entrants by 10x on affinity. That’s “trust us” made verifiable, a different posture than a benchmark score Anthropic controls end to end.

Where we’d push back on the framing implied by headlines like this one (ours included): “beat chemists at drug design” is a hook, not a finish line. Claude beat human designers at producing binders that stick to a target in a set of competitions where the target was already well-characterized. It did not design a drug — Anthropic said as much itself, plainly, in the same announcement. The gap between “binds in a plate” and “works in a patient” is where the actual drug-development industry lives, and it’s a gap AI hasn’t closed, this result or any other we’ve seen.

The Fable 5 exclusion is the detail we keep coming back to. Anthropic had its most capable model available and chose not to point it at this problem, on the record, over biosecurity risk — a company managing a capability it’s not fully comfortable with, in public, the same posture we’ve tracked around Mythos since April. Worth remembering next time a “Claude just did X” headline shows up without that caveat attached.

Frequently Asked Questions

Q: Did Claude design an actual new drug? A: No. Claude designed protein binders — molecules that attach to a disease-relevant target — and 14 of 15 tested targets produced at least one working binder in independent lab tests. Anthropic itself said a working binder is “just the first step” toward a drug-like molecule, not a finished treatment.

Q: How was this actually verified, and by whom? A: Adaptyv Bio and Twist Bioscience physically synthesized every design and ran binding assays in their own labs, independent of Anthropic. Adaptyv’s assay used Surface Plasmon Resonance across five target concentrations with duplicate measurements per design.

Q: Which Claude models were used? A: Claude Mythos Preview and Claude Opus 4.8 handled the protein design campaigns. A separate test used Claude Opus 5 to interpret NMR and LC-MS analytical chemistry data. Claude Fable 5, Anthropic’s most capable model, was deliberately excluded from the protein work over biosecurity concerns.

Q: How does a 22–35% hit rate compare to normal protein design work? A: Anthropic cites 10–15% as the typical hit rate for de novo protein design campaigns industry-wide. Claude’s combined-session rates (22.6–26.7%) roughly doubled that baseline, and its single-target rate (35.1%) more than doubled it.

Q: Is this a product I can use right now? A: Not as a self-serve tool. This was a research collaboration between Anthropic, Adaptyv Bio, and Twist Bioscience, run through Anthropic’s Claude Science research effort. There’s no announced commercial protein-design product tied to this result yet.

Q: Did every target succeed? A: No — 14 of 15. Claude produced zero confirmed binders against maltose-binding protein, and Anthropic published that result alongside the successes rather than omitting it.

Q: Has anyone pushed back on these results? A: Yes. Biotech commentator Martin Shkreli publicly dismissed the results as unimpressive, though his critique targets how far this is from a marketable drug rather than disputing the reported hit-rate or affinity figures themselves, which came from Anthropic’s independent lab partners.


Last updated: August 24, 2026. Sources: Anthropic — How Claude is accelerating protein design and analytical chemistry · Adaptyv Bio — Benchmarking Claude’s protein designs in the wet lab · Forbes — Claude Designed Proteins That Worked Against 14 Of 15 Disease Targets · pharmaphorum — Claude Science outperforms experts in protein binder task · TheNextWeb — Anthropic says Claude designed working protein binders, and beat human experts on some · CNBC — Anthropic launches AI drug discovery program, Claude Science · Stocktwits — Martin Shkreli slams Anthropic’s Claude drug-discovery claims.

Related reading: Claude Opus 4.8 Review: Fast Mode Got 3x Cheaper · Fable 5 Pulled: What Buyers Need to Know · Anthropic’s Claude Mythos: Too Dangerous to Release · Anthropic’s IPO Could Top SpaceX · OpenAI’s AI Hacked Hugging Face — Then It Paused Astra