All articles

Research · CASP17

Choosing Protein Models: Lessons from CASP17

· Vivek Kohar · 5 min read

A predicted protein structure is useful when it helps us choose an experiment: which residues to mutate, where to test binding, or how to investigate an assembly. When thousands of models suggest different answers, the challenge is deciding which structural ideas deserve that investment.

In CASP17, UniBio Intelligence used UbiStruct-QA to estimate the quality of 70,697 supplied models across five targets: two antibody–antigen complexes, two peptide assemblies, and a 24-subunit protein cage. CASP tests predictions against experimental structures withheld during prediction. Its assessment results are still pending. [20]

Our contribution was a method for scoring those candidates. It builds on a long-standing challenge in CASP: judging whether the parts of a complex are assembled correctly. [1][2][3][4] The lesson from our analysis is that a useful ranking should make its biological evidence and its sensitivity to assumptions visible.

Where the models came from

We received structures from CASP participants and large prediction pools supplied by the MassiveFold team. MassiveFold makes extensive sampling practical: repeated runs explore alternative structures rather than returning just one prediction. [5] Each antibody and peptide target included 17,085 MassiveFold entries, alongside 403–606 participant models. The cage pool contained 342 participant models only.

The supplied MassiveFold files were labelled AlphaFold-Multimer, the AlphaFold2 extension for complexes; [6][7] AlphaFold 3, which uses diffusion to generate atomic structures; [8] ColabFold, which makes sequence searches and AlphaFold prediction more accessible; [9] and ESMFold2, a protein language-model-based predictor. [10] File names recorded variations in sampling seeds, sequence alignments, and template settings. We could identify the predictor family from these file names, but could not reconstruct every setting used to generate each model. Method information was more limited for the CASP participant models.

Each MassiveFold pool contained 8,040 AlphaFold 3, 5,025 AlphaFold-Multimer, 3,015 ESMFold2, and 1,005 ColabFold entries. These are sample counts, not accuracy measurements.
The same family counts occurred in each of the four antibody and peptide pools. AlphaFold 3 contributed eight times as many entries as ColabFold. The bars show how many candidate structures each method contributed. Comparing their accuracy will require the experimental structures.

We thank the CASP participants and the MassiveFold team for generating the models, and the method developers cited above for the prediction tools. UbiStruct-QA scored the supplied structures using their confidence outputs and our structural checks. We did not train a new quality-estimation model or generate new sequence alignments for the submitted scores.

Count evidence, not repeated predictions

We found byte-for-byte identical peptide structure files under different group IDs. Related entries may have reused the same prediction, or shared a prediction pipeline; the files alone do not establish why. Counting each copy separately would make one prediction look like agreement between independent predictions.

For the antibody targets, we calculated how often each contact appeared within each CASP group’s models and within each MassiveFold predictor family. We averaged these contact-frequency summaries separately for the CASP participant pool and the MassiveFold pool, giving each distinct summary equal weight. If two groups or families within a pool had identical summaries, we counted that summary once. We then combined contact support from the CASP participant pool and the MassiveFold pool. For peptides, we applied the same principle to groups of structurally similar models. This reduced the influence of repeated predictions and unequal numbers of submitted models.

This does not make all remaining predictions independent: methods can share training data or assumptions, and each candidate still contributes to its own agreement score. Peptide scoring had another limitation. CASP participant models carried labels for the requested assembly states; MassiveFold models did not. A labelled participant model could receive support from participant models with the same label and from the MassiveFold pool. An unlabelled MassiveFold model received support only from the MassiveFold pool. The agreement scores therefore used different evidence for the two pools.

Check the biology of the assembly

Agreement alone is not enough. We combined it with predictor confidence, completeness, packing, and clashes. Confidence values were not calibrated across predictors; interface-specific scores such as ipTM and actifpTM supplied additional evidence where available. [11] The structural questions differed by target:

  • Antibody complexes: do the antibody and antigen make plausible contacts? Published NDM-1 and MYDGF structures provided antigen references; [12][13] ImmuneBuilder supplied antibody references. [15] Distance comparisons adapted from lDDT checked component geometry. [14] Well-folded components can still be docked in the wrong orientation.
  • Peptide assemblies: do the chains form a consistent repeating arrangement? One target was annotated as an amyloid, motivating cross-β checks; [16] public metadata for the other motivated helical packing checks. [17] That metadata guided an assumption; no unreleased coordinates were used.
  • The protein cage: do 24 subunits form a complete, well-packed shell? We used structural checks and confidence, with a related experimental cage as a reference, rather than consensus across candidates. [18][19]

Test the shortlist

For peptide target T2464, we ranked all 17,683 candidates by our baseline score. Here, “top 100” means the 100 models with the highest scores for that target. We then kept the models fixed, changed the scoring weights, and counted how many original top-100 model IDs remained in the new top 100.

Giving confidence more weight retained 64 of the original 100: 36 left the shortlist. Yet the rank correlation across the entire pool was 0.993, close to the value of 1 for identical ordering. Removing other evidence produced larger changes; the strongest stress test, using structural checks alone, retained 19.

Original top-100 models retained after rescoring: more confidence, 64 with rank correlation 0.993; no model-agreement term with other weights rebalanced, 47 and 0.994; structure-heavy without model agreement, 49 and 0.993; structural checks alone, 19 and 0.983.

All bars compare the same 17,683 candidates with the baseline shortlist. Correlation describes the ordering of the whole pool; the bars count original top-100 models retained. Neither measures agreement with experimental structures.

Baseline weights were 80% structural checks, 15% confidence, and 5% model agreement. From top to bottom, alternatives used 70/25/5, 82/18/0, 90/10/0, and 100/0/0. Removing model agreement also rebalanced the remaining weights. Chart data.

A model leaving the top 100 could be replaced by a nearly identical structure, so 36 replacements do not necessarily mean 36 different biological interpretations. Even so, the result shows that a ranking can look almost unchanged across thousands of models while its shortlist changes substantially. Comparison with experimental structures is needed to determine which scoring weights select more accurate models.

For computational teams, a useful report should show the top-ranked models and how that selection changes when scoring assumptions change. For wet-lab teams, the next step is to compare the shortlisted structures: which binding contacts or assembly features recur, and where do they disagree? Those differences can guide an experiment—for example, testing binding after mutating a residue predicted to contact the partner in only some models. CASP’s independent assessment will test how well our scores reflect structural accuracy.

Explore the UbiStruct-QA code and methodology.

References and acknowledgments

We thank the CASP organizers and experimental target providers, participating prediction teams, the MassiveFold team, and the developers of the methods below. References 1–19 are the works cited in our methods abstract.

  1. [1] Studer, G., Tauriello, G. & Schwede, T. (2023). Assessment of the assessment—All about complexes. Proteins 91, 1850–1860. doi:10.1002/prot.26612. Source.
  2. [2] Edmunds, N.S., Alharbi, S.M.A., Genc, A.G., Adiyaman, R. & McGuffin, L.J. (2023). Estimation of model accuracy in CASP15 using the ModFOLDdock server. Proteins 91, 1871–1878. doi:10.1002/prot.26532. Source.
  3. [3] Roy, R.S., Liu, J., Giri, N., Guo, Z. & Cheng, J. (2023). Combining pairwise structural similarity and deep learning interface contact prediction to estimate protein complex model accuracy in CASP15. Proteins 91, 1889–1902. doi:10.1002/prot.26542. Source.
  4. [4] Fadini, A., Adiyaman, R., Alhaddad, S.N., et al. (2026). Highlights of Model Quality Assessment in CASP16. Proteins 94, 314–329. doi:10.1002/prot.70035. Source.
  5. [5] Raouraoua, N., Mirabello, C., Véry, T., et al. (2024). MassiveFold: unveiling AlphaFold’s hidden potential with optimized and parallelized massive sampling. Nat. Comput. Sci. 4, 824–828. doi:10.1038/s43588-024-00714-4. Source.
  6. [6] Jumper, J., Evans, R., Pritzel, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. doi:10.1038/s41586-021-03819-2. Source.
  7. [7] Evans, R., O’Neill, M., Pritzel, A., et al. (2021). Protein complex prediction with AlphaFold-Multimer. bioRxiv (preprint). doi:10.1101/2021.10.04.463034. Source.
  8. [8] Abramson, J., Adler, J., Dunger, J., et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500. doi:10.1038/s41586-024-07487-w. Source.
  9. [9] Mirdita, M., Schütze, K., Moriwaki, Y., Heo, L., Ovchinnikov, S. & Steinegger, M. (2022). ColabFold: making protein folding accessible to all. Nat. Methods 19, 679–682. doi:10.1038/s41592-022-01488-1. Source.
  10. [10] Candido, S., Hayes, T., Derry, A., et al. (2026). Language Modeling Materializes a World Model of Protein Biology. bioRxiv (preprint). doi:10.64898/2026.06.03.729735. Source.
  11. [11] Varga, J.K., Ovchinnikov, S. & Schueler-Furman, O. (2025). actifpTM: a refined confidence metric of AlphaFold2 predictions involving flexible regions. Bioinformatics 41, btaf107. doi:10.1093/bioinformatics/btaf107. Source.
  12. [12] Feng, H., Ding, J., Zhu, D., et al. (2014). Structural and Mechanistic Insights into NDM-1 Catalyzed Hydrolysis of Cephalosporins. J. Am. Chem. Soc. 136, 14694–14697. doi:10.1021/ja508388e. Source.
  13. [13] Ebenhoch, R., Akhdar, A., Reboll, M.R., et al. (2019). Crystal structure and receptor-interacting residues of MYDGF — a protein mediating ischemic tissue repair. Nat. Commun. 10, 5379. doi:10.1038/s41467-019-13343-7. Source.
  14. [14] Mariani, V., Biasini, M., Barbato, A. & Schwede, T. (2013). lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests. Bioinformatics 29, 2722–2728. doi:10.1093/bioinformatics/btt473. Source.
  15. [15] Abanades, B., Wong, W.K., Boyles, F., Georges, G., Bujotzek, A. & Deane, C.M. (2023). ImmuneBuilder: Deep-Learning models for predicting the structures of immune proteins. Commun. Biol. 6, 575. doi:10.1038/s42003-023-04927-7. Source.
  16. [16] Bücker, R., Seuring, C., Cazey, C., et al. (2022). The Cryo-EM structures of two amphibian antimicrobial cross-β amyloid fibrils. Nat. Commun. 13, 4356. doi:10.1038/s41467-022-32039-z. Source.
  17. [17] Bloch, Y., Rayan, B., Ragonis-Bachar, P. & Landau, M. (2025). helical Aurein 3.3, mP form. wwPDB / RCSB Protein Data Bank. PDB 9SR1 (public metadata). Accessed 23 September 2026. PDB metadata.
  18. [18] King, N.P., Sheffler, W., Sawaya, M.R., et al. (2012). Computational Design of Self-Assembling Protein Nanomaterials with Atomic Level Accuracy. Science 336, 1171–1174. doi:10.1126/science.1219364. Source.
  19. [19] Edwardson, T.G.W., Mori, T. & Hilvert, D. (2018). Rational Engineering of a Designed Protein Cage for siRNA Delivery. J. Am. Chem. Soc. 140, 10439–10442. doi:10.1021/jacs.8b06442. Source.
  20. [20] Protein Structure Prediction Center. CASP17: Critical Assessment of Structure Prediction (2026). CASP17.
Cite this post
@misc{kohar2026casp17,
  author = {Kohar, Vivek},
  title = {Choosing Protein Models: Lessons from CASP17},
  year = {2026},
  url = {https://unibiointelligence.com/blog/casp17-choosing-protein-models/},
  month = {September}
}

Related reading: Predicting CYP inhibition: lessons from the OpenADMET challenge