← All UbiTiter issuesUbiTiter Issue #17
Validated Tooling: Developability, Cyclic Peptides, ADC Payload Distribution
Four platforms this week each report the control that could have sunk them: CDR H3-clustered splits, first-round hits without optimisation, spatial payload gradients, and repeat runs of one analysis.
- antibody developability
- protein language models
- cyclic peptides
- non-canonical amino acids
- antibody-drug conjugates
- quantitative systems pharmacology
- binding-site barrier
- pharmacometrics
The field
This Week in Biologics
Paper 1 · bioRxiv
An interpretable open platform for sequence-based antibody developability prediction
DELPHI predicts polyreactivity and SEC behaviour at AUC 0.959 and 0.933 under CDR H3-clustered cross-validation, then transfers to a 246,293-antibody public library at 0.950 without training on it.
Core finding
DELPHI is open software for training antibody developability predictors from labelled assay data, shipped with ready-to-run retrainable models. It compares 25 language-model and classifier combinations under CDR H3-cluster cross-validation, reaching mean AUC 0.959 on polyreactivity and 0.933 on size-exclusion behaviour using in-house data. Trained on those data alone it transfers to a 246,293-antibody public library at AUC 0.950, ranking polyreactivity at a level similar to the best reported Ginkgo competition point estimate without having trained on that competition's data.
What is novel
The evaluation protocol is the contribution as much as the models. Antibody repertoires contain families of closely related sequences, so a random train/test split places near-duplicates on both sides and the resulting accuracy partly measures how well the model recognises sequences it has already seen. Grouping antibodies into CDR H3 clusters and holding out whole clusters removes that overlap, which makes the reported AUC an estimate of performance on a CDR H3 the model has not encountered, the situation that actually applies when a new candidate is screened. Two further outputs make the platform usable rather than only benchmarkable. It characterises how predictive accuracy changes as the number of labelled training examples grows, which lets a laboratory estimate how much of its own assay data it must generate before retraining improves on the shipped model. And it attributes predictions to individual residues, so a flagged antibody comes with candidate positions to mutate rather than only a score, converting a pass/fail filter into a starting point for engineering.
Limitations
The models cover two assay endpoints, polyreactivity and size-exclusion behaviour, and both label sets were generated within one organisation, so the assay protocols, instrumentation and scoring conventions behind the training data are uniform in ways a second laboratory's data would not be. Developability decisions in practice also turn on viscosity at formulation concentration, aggregation propensity, thermal stability and expression titre, none of which are predicted here, so this addresses part of a candidate triage rather than replacing it. The residue-level attributions are explanations of what the model weighted, not measurements of what each residue contributes to the liability, and the two can diverge; they are hypotheses for mutagenesis rather than conclusions to design against. The transfer result on the public library is also a ranking comparison against a reported competition estimate rather than a head-to-head evaluation under one shared protocol.
Why it matters in context
Sequence-based developability prediction sits at the point in discovery where hundreds of candidates must be narrowed before anyone commits expression and purification capacity to them, which is why the accuracy number attached to such a model matters commercially and not only methodologically. Two conditions determine whether another group can act on a published model: whether the software and trained weights are distributed, and whether the reported accuracy came from a split that controls for sequence similarity between training and test antibodies. This release satisfies both, shipping the training code, the trained models, the clustered cross-validation protocol and a transfer evaluation on a public library of roughly a quarter million antibodies. The practical consequence is that a team can reproduce the stated numbers, apply the CDR H3 clustering protocol to its own internal validation splits to check whether its in-house models have been scored optimistically, and retrain on its own assay labels with a published baseline to compare against.
Paper 2 · bioRxiv
De novo design of small cyclic peptide inhibitors from non-canonical amino acids by cofolding-guided search
nCycle-Forge designed 1,472 cyclic peptides against thrombin and MDM2 from target sequence alone, reaching Ki 0.14 uM and EC50 2.4 uM in a single round with hit rates up to 24-fold above larger experimental libraries.
Core finding
nCycle-Forge is a simulated annealing framework that designs cyclic peptide binders by iteratively substituting building blocks and scoring each proposal by cofolding with the target. The authors designed, synthesised and screened 1,472 cyclic peptides against two structurally different sites: the active-site pocket of thrombin and the shallow p53-binding surface of MDM2. Both campaigns started from target sequence alone with no fixed binding motif, and yielded chemically diverse hits in a single round with no experimental optimisation. The best thrombin inhibitor reached Ki 0.14 uM and the best MDM2 binder EC50 2.4 uM.
What is novel
The method needs no building-block-specific parameterisation, no conformer generation and no retraining, so the chemical alphabet extends past canonical amino acids without a per-block setup cost, and the designs start from target sequence alone rather than from a known interaction motif or a library carried over from an earlier campaign. Hit rates reached up to 24-fold higher than in larger experimental libraries. Surface plasmon resonance confirmed direct binding, competition with orthosteric inhibitors at both sites, and weaker binding under conditions expected to linearise the designs, which ties the activity to the intended cyclic conformation.
Limitations
Potency is sub- to low-micromolar, a starting point rather than a therapeutic endpoint. Two targets do not establish general applicability, though a deep pocket and a shallow protein-protein interface are a reasonable span of site geometry. The reported hit rates are benchmarked against experimental library screens rather than against another computational design method, so the comparison shows that the designs beat random library sampling and not that this search outperforms an alternative in silico approach on the same targets.
Why it matters in context
Macrocyclic peptides occupy the space between small molecules and biologics: large enough to cover a protein-protein interface that a small molecule cannot engage, small enough to reach compartments an antibody cannot, and constrained by cyclisation into a limited conformational ensemble that improves affinity and proteolytic stability relative to a linear peptide. Non-canonical amino acids extend that range further, adding side-chain chemistry and backbone geometry unavailable from the canonical twenty, and they are also what makes the class computationally awkward: a docking-based pipeline needs force-field parameters derived for each new building block and a conformer ensemble enumerated for each macrocycle, so every extension of the chemical alphabet carries a setup cost before a single design can be scored. Cofolding models now predict peptide-protein complex structure accurately enough to act as the scoring function inside a design loop, and using them that way is what lets the search move through non-canonical chemical space without per-block preparation. The trade is that the score is a learned structural prediction rather than a physics-based energy, so its reliability across unusual building blocks is established empirically, by synthesising designs and measuring them, which is what the 1,472-peptide screen does.
Paper 3 · European Journal of Pharmaceutical Sciences
A multiscale quantitative systems pharmacology platform for early screening of MMAE-based antibody-drug conjugates
Across three vc-MMAE ADCs differing only in target antigen, a QSP model predicted intratumoral free MMAE within -38% to +21% and showed the highest-expressing target giving the least payload penetration through a binding-site barrier.
Core finding
A multiscale QSP model linking systemic pharmacokinetics, Krogh-cylinder tumour penetration and single-cell disposition predicted intratumoral unconjugated MMAE exposure within -38.13% to 20.52% across three vc-MMAE antibody-drug conjugates that share one linker-payload platform and a drug-to-antibody ratio of 4, differing only in target antigen: RC48 against HER2, RC108 against c-Met and RC118 against Claudin 18.2. Combining predicted payload exposure and spatial distribution with in vitro cytotoxicity reproduced the observed in vivo efficacy ranking of the three.
What is novel
Holding the linker and payload constant and varying only the antibody isolates the contribution of the targeting arm, which is what makes the central result interpretable. RC48 had the strongest in vitro potency, an ADC IC50 of 0.46 nM against a cell line carrying 1.94 million HER2 copies per cell, and the least intratumoral penetration: high antigen density with fast binding kinetics captures drug in perivascular regions and leaves deep tumour compartments underexposed. RC108 and RC118, with antigen expression roughly 20- to 40-fold lower, distributed more homogeneously and outperformed in vivo. The model therefore supplies a mechanistic account of an in vitro to in vivo disconnect rather than only a prediction, and global sensitivity analysis attributes intratumoral exposure principally to target density, ADC permeability, systemic clearance and payload uptake and efflux rates.
Limitations
Tumour growth is not modelled dynamically, so predictions degrade once tumours regress substantially, as seen for RC108 beyond 168 hours post-dose. Prediction error was widest for RC48, which the authors attribute to the same spatial heterogeneity the binding-site barrier creates, compounded by homogenate-based measurements being sensitive to sampling location and by a calibration cell line derived from a trastuzumab-resistant subline that may carry altered payload influx and efflux. All three conjugates share one payload class, so generalisation to DXd, DM1 or pyrrolobenzodiazepine dimers is untested. Validation depends on xenograft data, and the authors note that clinical data are rarely parameterised in a form a model like this can consume.
Why it matters in context
An ADC's activity depends on a chain of steps that a single potency number collapses together: plasma exposure, extravasation into tumour tissue, diffusion through that tissue, antigen binding and internalisation, linker cleavage, and payload retention or efflux. An empirical exposure-response analysis correlates a systemic exposure metric with an outcome and therefore cannot separate how much payload reached the cells from how effective it was once inside them, which is why a molecule can rank first in a cytotoxicity assay on a high-expressing line and lose in vivo. Resolving payload concentration spatially, along the vascular-to-interior gradient represented by a Krogh cylinder, is what lets this model attribute the observed in vivo ranking to distribution rather than intrinsic potency, and it identifies the specific parameters that control the outcome: target density, ADC permeability, systemic clearance and payload uptake and efflux. The design implication runs against a common default. Target selection and affinity maturation are usually pushed toward higher antigen expression and tighter binding, and the binding-site barrier means both can be pushed past the point of benefit, since a conjugate captured by the first antigen-positive cells it meets never reaches the tumour interior. That trade-off has been described in the ADC literature before; what this work adds is a calibrated model that puts an expression window on it for three clinical-stage molecules sharing one payload.
Paper 4 · CPT: Pharmacometrics & Systems Pharmacology
PMxAgent: An Agentic Platform for Pharmacometrics
Exposing validated R pharmacometric tools over MCP gave 98.3% NCA accuracy across 1,820 subjects and bit-for-bit reproducibility across 10 runs, where a frontier agent writing its own code varied between 92.7% and 95.6%.
Core finding
PMxAgent is an open-source platform that exposes R pharmacometric functions as agent-callable tools: an R API server built with Plumber publishes an OpenAPI specification, and a Python MCP server turns that specification into tools any MCP-compatible agent can discover and invoke. Five tools cover non-compartmental analysis, exposure-response modelling, PK simulation, data standardisation and population PK dataset generation. Benchmarked against four frontier AI agents across 182 drugs and 1,820 simulated subjects with PKanalix as reference, PMxAgent reached 98.3% NCA accuracy, against 98.2% for Claude Opus 4.8, 95.6% for GPT 5.5 pro, 94.5% for Claude Sonnet 4.6 and 91.6% for GPT 5.5 thinking.
What is novel
The accuracy ranking spans four percentage points; the reproducibility result is categorical. Across 10 identical runs PMxAgent returned bit-for-bit identical results with full tool-call accuracy, while GPT 5.5 pro varied bimodally between 92.7% and 95.6%, having regenerated its analysis code on each run. The error analysis localises the difference to a single step: C_max and T_max were reproduced exactly by every agent across all 1,820 subjects, and essentially all divergence traced to the choice of terminal timepoints for the elimination-rate estimate, which propagates into half-life, AUC to infinity, clearance and volume. The lower-accuracy agents matched the reference terminal points for 77% to 78% of subjects against 96% for the others. PMxAgent was orchestrated by Claude Sonnet 4.6, the same model that scored 94.5% performing the analysis directly, so a smaller model calling a validated tool reached the accuracy of a much larger one.
Limitations
The benchmark used simulated data that needed no cleaning, so the data-wrangling work that dominates real pharmacometric analysis was not tested. Handling instructions for values below the limit of quantification were inadvertently omitted from the prompt, which the authors say may have improved accuracy had they been included. Authentication rests on the MCP framework's built-in OAuth without role-based access control, persistent token management or validation documentation, so this is a single-user local prototype rather than a deployable regulated system. Terminal-point selection is a judgement call without one correct answer, which complicates treating any reference as ground truth.
Why it matters in context
Pharmacometric analyses supporting a regulatory submission must be reproducible and traceable: the same inputs return the same parameter estimates, and the code path that produced them can be inspected and validated. An agent that writes fresh analysis code on each execution fails both tests, not because the code is necessarily wrong but because it is a new, unvalidated artefact every run, which is what the bimodal 92.7% to 95.6% spread across repeat runs demonstrates concretely. Separating orchestration from execution addresses this directly: the language model decides which analysis to run and with what arguments, while the numerical work happens in established R packages that are versioned and validated once. That division also explains why a smaller orchestrating model matched a much larger one, since the accuracy of the result depends on the tool rather than on the model's arithmetic. The second design choice is data movement. Passing datasets between tools by file reference through host-mounted directories, rather than serialising tables into the model's context, keeps large datasets out of the context window entirely; the authors note that transcription error in agent-mediated transfer is nonzero and grows with dataset size, so this is a correctness argument as much as an efficiency one, and it generalises to any MCP server operating over real scientific data.
Our work
This Week at UniBio Intelligence
Material additions to our data, models, tools, and research platform.
tool
Fc sequence analysis and construct comparison
Identify an Fc sequence and compare explicit constructs side by side, so format choices can be checked against each other before you commit to one.
tool
Immunogenicity prediction across HLA-II, DRB1 and MHC-I
Assess HLA-II presentation, population-level DRB1 risk and MHC-I presentation through HLAIIPred, ImmunoGeNN and MHCFlurry. Useful alongside this week's developability work, which covers polyreactivity and size-exclusion behaviour but not immunogenicity.
tool
Structure-aware binder design and scoring
Binder design and scoring workflows are now public, including BoltzGen, HalluDesign and DISCO. The cyclic peptide work in this issue uses cofolding as a design-loop scoring function, which is the same pattern these workflows support.
platform
PK/PD model fitting and controlled execution in the modeling workspace
The public modeling workspace now runs PKPD workflows, including model fitting, with execution you can trace. Relevant to both pharmacology papers in this issue.
Primary papers
- [1] An interpretable open platform for sequence-based antibody developability prediction (bioRxiv)
- [2] De novo design of small cyclic peptide inhibitors from non-canonical amino acids by cofolding-guided search (bioRxiv)
- [3] A multiscale quantitative systems pharmacology platform for early screening of MMAE-based antibody-drug conjugates (European Journal of Pharmaceutical Sciences)
- [4] PMxAgent: An Agentic Platform for Pharmacometrics (CPT: Pharmacometrics & Systems Pharmacology)