Ubi Titer Issue #12
AI Protein Design: Generation Works, Ranking Does Not
Five papers published in the same week, using entirely different methods and targets, land on the same limitation: current models generate viable candidates at useful rates but cannot reliably rank them, leaving experimental measurement as the arbiter of which designs actually work.
- computational antibody design
- de novo binder design
- prospective benchmarking
- AlphaFold 3
- affinity prediction
- targeted protein degradation
- T-cell engagers
- epitope targeting
The field
This Week in Biologics
Paper 1 · Nature Biotechnology
A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability
511 AI-designed antibodies from 29 organizations tested blind: affinity maturation reached 95 pM, but AI ranking within sequence clusters beat the control less often than random clone picking.
Core finding
Across three blinded, prospective tasks against SARS-CoV-2 RBD, 511 antibody sequences from 29 organizations were expressed as full IgGs and measured under one uniform protocol combining high-throughput SPR, orthogonal KinExA, and a five-assay developability panel. Performance split sharply by task. In silico affinity maturation succeeded, with the winning entry reaching 95 pM, a 2,000-fold improvement over the parental antibody and statistically indistinguishable from the best experimentally derived antibody at 113 pM. Affinity ranking within HCDR3 clusters failed: only 9.8 to 13.8 percent of AI submissions identified antibodies better than the cluster control, against roughly 39 percent for randomly picked clones, and only one participant exceeded that random baseline.
What is novel
This is the first CASP-style prospective, blinded, wet-lab-anchored benchmark for antibody discovery, in which every submission was synthesized and characterized under identical conditions by independent laboratories rather than evaluated retrospectively by its own authors. The study also establishes cross-platform assay concordance for rank ordering, with SPR and single-point KinExA correlating at Spearman rho 0.94.
Limitations
The challenge used a single, extensively characterized antigen with unusually deep sequencing and affinity data, so the results represent an upper bound on current capability rather than general performance across arbitrary antigens. The organizing consortium also reported the results and relied on informatic rather than structural safeguards for blinding, and several prominent groups with public design claims did not participate. The top-ranked out-of-library design failed to elute from the hydrophobic interaction column, a liability the composite developability score allowed to pass because failure in one assay could be offset by strength elsewhere.
Why it matters in context
Prospective blinded evaluation did for structure prediction what retrospective self-assessment could not, and this challenge imports that discipline into antibody design. The finding that most AI ranking strategies underperformed random clone picking sharpens an open question about whether current models learn transferable affinity determinants or exploit dataset-specific structure, and a purely statistical consensus method with no machine learning placing third overall sets a baseline that learned methods have yet to clear. Analysis by method class showed no single approach dominating across tasks, supporting a portfolio view rather than a search for one general model.
Paper 2 · Anthropic (technical report)
Autonomous de novo protein binder design with Claude
An AI agent ran 24 to 48 hour binder design campaigns against 16 targets with no human input into any design decision; 354 of 1,320 designs bound, a 27 percent hit rate.
Core finding
Two Claude models ran de novo protein binder campaigns against 16 targets from a single frozen protocol prompt, autonomously selecting target constructs, epitopes, backbone and sequence design tools, filtering criteria, and final rankings, with no human input into any individual design decision. Across 1,320 designs on 15 evaluable targets, synthesized exactly as delivered and measured blind by two independent contract research organizations, 354 bound, a hit rate of 26.8 percent, with binders against 14 of 15 targets and 49 percent among top-ranked designs. Twelve binders were obtained against the TNF-alpha trimer, a target against which several prior de novo efforts reported none.
What is novel
Earlier work coupling language models to protein design tools kept humans in the loop for target construct selection, epitope choice, tool configuration, or final ranking. Here every decision from target research through final ranking was made by the agent under one frozen protocol, every delivered design was synthesized without modification, and results from campaigns that failed are reported alongside those that succeeded, with per-design provenance recording which tool produced each tested candidate.
Limitations
The evidence is binding, not structure or function. No design was solved structurally or assayed for activity, so every reported pose is a prediction and affinities against the five oligomeric targets are apparent values inflated by avidity. No matched human-expert campaign was run, so the work does not establish that the agent outperforms an expert given the same tools and budget. Each combination of model, format, and target ran once, confounding model differences with format and run-to-run variation. Most targets chosen are extensively characterized in the literature. The report is not peer reviewed, the corresponding author is an Anthropic employee, and both contract research organizations were paid to perform the validation.
Why it matters in context
The comparison against open design competitions measured in the same laboratory with the same assay provides a rare quasi-benchmark, with hit rates exceeding the assembled field on four of six shared targets and matching it on the other two. More consequential than the headline rate is what the scoring data show: co-folding confidence separated binders from non-binders within a target at mean average precision 0.52 against 0.31 expected by chance, yet predicted affinity poorly among binders and gave almost no warning of whole-target failure, with designs against maltose-binding protein scoring nearly as well as those against targets that yielded dozens of binders. That pattern matches independent findings from prospective antibody benchmarking and de novo scaffold design in the same period, placing ranking and filtering rather than generation at the field's current limit.
Paper 3 · mAbs
AlphaFold 3 prediction of TCR mimic Fv-pMHC complex structures and their validation and optimization
AlphaFold 3 modeled TCR mimic antibody-pMHC complexes accurately and passed mutagenesis validation, but all three AF3-designed affinity mutants failed experimentally.
Core finding
AlphaFold 3 modeled TCR mimic Fv-pMHC quaternary structures reliably when confidence was high, recovering the canonical diagonal binding geometry across the peptide-binding groove for three NY-ESO-1/HLA-A*02:01-specific CAR clones, with Fv-pHLA contact areas of 1065 to 1204 square angstroms against 1331 for the natural T-cell receptor. Predicted epitopes and paratopes were confirmed by alanine scanning on both the peptide antigen and the CDRs. Affinity optimization then failed completely: of 21 modeled mutants, the three selected for larger contact surfaces and additional hydrogen bonds were all predicted at high confidence, yet two lost tetramer binding outright and the third bound more weakly than the parent.
What is novel
The study couples structure prediction to systematic experimental validation on both sides of the interface and then tests the prediction-to-design step directly rather than stopping at modeling accuracy. It also quantifies the training-cutoff effect for this complex class, with 57 percent accuracy on structures released before the cutoff against 40 percent after, and shows that degraded confidence scores correctly flagged four of six experimentally confirmed loss-of-function paratope mutants.
Limitations
Correct prediction was achieved for only 6 of 12 tested TCR mimic clones, with failures concentrated in clones whose binding geometry departs from canonical T-cell receptor engagement. The design search path was restricted, beginning at one CDR loop and then a second, so beneficial mutations elsewhere or in combination may have been missed, and only one peptide-MHC system was examined. Underlying data are available from the corresponding author on request rather than deposited.
Why it matters in context
The overall 50 percent correct-prediction rate is on par with reported performance for antibody-antigen and nanobody-antigen complexes, so the modeling result is consistent with the wider benchmarking literature. The design failure has a concrete physical explanation: the model removes explicit water molecules during processing and does not predict their positions, yet buried waters mediate hydrogen bonds and contribute much of the shape complementarity at this interface. More generally, the model produces no thermodynamic output, so buried surface area and hydrogen bond counts are geometric proxies rather than free energies, and they prove nearly uninformative when ranking point mutants within a single complex. The authors point to force field methods and newer specialized predictors as the route forward.
Paper 4 · mAbs
Mimic antibodies: leveraging ligand mimicry for epitope-targeted antibody discovery
Ranking a 20,000-sequence repertoire by how closely each antibody reproduces a cognate ligand's contact residues gave 11 binders from 31 picks, eight of them sub-nanomolar.
Core finding
A systematic search of the Protein Data Bank identified 1,834 antibodies sharing five or more epitope residues with a target's natural ligand, of which 892 reproduce four or more residues of matching physicochemical character, establishing that antibody mimicry of natural ligands is widespread and mostly scattered across CDRs rather than copying a ligand loop. Turning that observation into a selection prior, the authors co-folded a 20,000-sequence immunized rabbit repertoire against the IL-18 receptor alpha D3 domain, ranked candidates by how many IL-18 contact residues each antibody reproduced in spatially equivalent positions, and obtained 11 binders from 31 tested, eight of them sub-nanomolar and stronger than the natural IL-18 ligand at 18 to 69 nanomolar, with no experimental affinity maturation.
What is novel
The work identifies scattered mimicry, where key interaction residues are conserved in spatial position without the backbone topology being copied, as the dominant and previously unexploited mimicry mechanism, and converts it into a prospective screening prior. A head-to-head comparison against surface fingerprinting, which returned no binders from 11 picks on the same repertoire, target, and predicted structures, isolates residue-level correspondence as the discriminating signal over surface shape complementarity.
Limitations
The approach needs a structurally rigid target epitope and a solved cognate ligand complex; the authors chose this particular interface because the alternative binding site on the same receptor engages a conformationally heterogeneous region. Binding was confirmed to the isolated domain but there is no direct evidence the validated antibodies engage the specific predicted mimicry epitope, and cross-reactivity with other natural partners that bind through overlapping regions cannot be excluded. Within the pre-filtered candidate pool, neither confidence score nor total mimicry count alone separated binders from non-binders. The method selects mimics already present in a repertoire rather than generating them, and its performance is bounded by the accuracy of the underlying co-folding model.
Why it matters in context
This offers a practical counterweight to de novo generation by exploiting somatically matured diversity already present in an immune repertoire together with a structural prior. It addresses a longstanding gap in repertoire mining, where standard analysis ranks candidates by clonal abundance on the assumption that expanded clones make the best leads, despite frequency correlating poorly with affinity and sequence-based mining carrying no information about binding mode or epitope. The finding that mimicry concentrates at evolutionarily conserved, structurally central interface positions connects the result to established binding hotspot theory, and a Protein Data Bank-wide analysis showing residues near the interface centroid are mimicked at nearly twice the background rate suggests the prior generalizes beyond the single system tested.
Paper 5 · Journal of the American Chemical Society
De Novo-Designed Bifunctional Proteins for Targeted Protein Degradation
A designed 67-residue protein scaffold carrying a BCL-xL binder and an E3-recruiting motif degrades BCL-xL in cells and triggers apoptosis, approaching a clinical PROTAC.
Core finding
A de novo 67-residue helix-turn-helix scaffold, confirmed crystallographically at 1.5 angstrom resolution with 0.36 angstrom backbone deviation from its design model, was engineered to present a BCL-xL binding site on one helix and a motif recruiting the E3 ligase adapter KLHL20 within its loop. Hot-spot grafting followed by computational sequence design and structure-prediction filtering gave a 75 percent success rate for submicromolar binders, with the best reaching 3 nanomolar and co-crystal structures confirming the designed poses. The bifunctional construct degraded 54 percent of endogenous BCL-xL in lung cancer cells and induced 56 percent apoptosis, against 78 and 73 percent for the clinical PROTAC DT2216, while neither half alone did either.
What is novel
The degrader is built entirely from designed protein rather than small-molecule chemistry, sidestepping both linker optimization and the narrow set of chemically ligandable E3 ligases. Grafting the recruiting motif into a structurally constrained loop, rather than presenting it as a disordered peptide, promotes the compact conformation the ligase adapter recognizes while leaving the scaffold's fold and stability intact, so one small protein displays two epitopes with defined relative geometry. Crystal structures were obtained for the bare scaffold, both target complexes, and the motif-bearing loop designs.
Limitations
Only one target and E3 ligase pairing was demonstrated in cells, leaving the generality of the approach across other targets, ligases, and orientations untested. Degradation and apoptosis were achieved by transfecting the encoding plasmid rather than dosing a defined amount of protein, so intracellular concentrations are uncertain and comparison with the dosed small-molecule PROTAC is indirect. Delivery of designed proteins across cell membranes remains unsolved for therapeutic use, and few designs were tested experimentally.
Why it matters in context
Targeted protein degradation faces two acknowledged bottlenecks that this work addresses together: degradation efficiency depends on ternary complex formation governed by binding constants, linker length, and the proximity of suitable lysines, and most degraders recruit just two E3 ligases out of a large repertoire. Peptide-based approaches promised access to more ligases through natural recruitment motifs, but most such motifs are intrinsically disordered and therefore hard to present in precise, stable complexes, which is exactly what a rigid designed loop provides. The authors also report that none of the confidence predictors or physical scoring metrics they applied correlated well with measured binding affinities, an observation that echoes contemporaneous large-scale binder benchmarks and reinforces a stated field-level need for better predictors of binding potential.
Paper 6 · mAbs
Design and preclinical characterization of an affinity-tuned BCMA×CD3 bispecific T-cell engager for broad B-cell depletion
Deliberately weakening the CD3 arm gave a greater than 50-fold bias toward BCMA and near-identical potency against cells differing 23-fold in target density.
Core finding
Gamgertamig is a humanized IgG4-based bispecific T-cell engager built by deliberately selecting a reduced-affinity CD3 arm, giving 0.65 nanomolar affinity for BCMA against 34 nanomolar for CD3, a greater than 50-fold bias toward target compared with roughly 15-fold for a teclistamab analog. Potency proved essentially independent of antigen density, with half-maximal killing at 0.10 nanomolar against a line carrying 25,300 BCMA receptors per cell and 0.16 nanomolar against one carrying 1,100. The molecule depleted naive, transitional, unswitched memory and switched memory B cells across a 16-donor cohort spanning healthy donors and five disease states, and produced sustained circulating B-cell and bone-marrow plasma-cell depletion in rhesus monkeys with a terminal half-life exceeding five days.
What is novel
The design applies CD3 affinity de-optimization specifically to reach the low-density BCMA found on early B-cell subsets, rather than restricting activity to high-expressing plasma cells as prior BCMA-directed engagers developed for myeloma have done. It also demonstrates that the resulting pan-B depletion profile holds across autoimmune indications with widely differing baseline effector-to-target ratios, and that depletion proceeds through BCMA engagement even where surface BCMA sits below conventional flow cytometry detection.
Limitations
The work is entirely preclinical, with no human clinical data. Reductions in inflammatory cytokines relative to the comparator were numerical trends that did not reach statistical significance in the donor cohorts tested. Plasmablasts and plasma cells were too scarce in peripheral blood samples to permit reliable measurement of killing, so the plasma-cell effect rests on the primate bone-marrow data. Xenograft efficacy used reconstituted immunodeficient mice in which graft-versus-host disease caused deaths in control groups, and per-indication sample sizes of two to three donors limit inference about disease-specific behavior.
Why it matters in context
The molecule sits within a broader shift of BCMA-directed engagers from relapsed myeloma toward autoimmune immune reset, where CD19- and CD20-directed therapies cannot eliminate terminally differentiated plasma cells that lack those markers and whose persistence is thought to drive relapse after CAR-T therapy. Engagers offer practical advantages over cell therapy in this setting through scalable manufacturing, no need for lymphodepleting preconditioning, and step-up dosing to mitigate cytokine release. The work also adds to accumulating evidence that maximizing binding affinity is the wrong optimization target in biologics engineering, joining cases where affinity was deliberately reduced to widen a therapeutic window or tune pharmacokinetics.
Primary papers
- [1] A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability (Nature Biotechnology)
- [2] Autonomous de novo protein binder design with Claude (Anthropic (technical report))
- [3] AlphaFold 3 prediction of TCR mimic Fv-pMHC complex structures and their validation and optimization (mAbs)
- [4] Mimic antibodies: leveraging ligand mimicry for epitope-targeted antibody discovery (mAbs)
- [5] De Novo-Designed Bifunctional Proteins for Targeted Protein Degradation (Journal of the American Chemical Society)
- [6] Design and preclinical characterization of an affinity-tuned BCMA×CD3 bispecific T-cell engager for broad B-cell depletion (mAbs)