RefineNovo
algorithm · Transformer (NAR)
Curriculum learning
| Kind | algorithm |
| Deep learning | yes |
| Acquisition | DDA |
| Family | Transformer (NAR) |
Code
Live stars, open issues and last-push figures are on the Code activity chart.
Checkpoints
| Version | Trained on | Host | Licence | Size | Checked | Backup |
|---|---|---|---|---|---|---|
| RefineNovo-30M | — | Google Drive | MIT | 388 MB | verified 2026-10-02 | copy |
A host marked archival has a DOI and keeps what it is given. The others can move or disappear, which is why they are checked rather than merely listed. verified means the bytes were fetched and hashed on the date shown; live means only that the host answered when last asked. Where a copy is linked, it is a backup of someone else’s weights kept in case the original link goes stale; the original is the link to cite and to prefer.
Reported comparisons (5)
The comparison tables this method’s own papers print, standardised: every value on a 0-1 scale, methods down the side, the measure and then the species across. These are numbers papers report about themselves and their baselines. They are not a leaderboard, and they do not compare across tables: each was produced by a different group, on the dataset named in its corner, with each baseline either retrained, run from released weights or quoted from another paper. Where the paper says which, it follows the method’s name (hover it for the sentence); most papers do not say. Bold is the best value in a column and underline the runner-up, our ranking rather than the paper’s own marks.
Table1
Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing, page 7: Comparison of the performance on the 9-species-V1 benchmark datasets. The models are categorized by their architecture type: DB represents Database, AR stands for Autoregressive Generation, and NAR denotes Non-Autoregressive Generation. The bold font indicates the best performance.
|
Nine-species benchmark original (DeepNovo, 2017) |
Amino acid precision | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mus musculus | Homo sapiens | Saccharomyces cerevisiae | Methanosarcina mazei | Apis mellifera | Solanum lycopersicum | Vigna mungo | Bacillus subtilis | Candidatus Thiodiazotropha endoloripes | Average | |
| PEAKS | 0.600 | 0.639 | 0.748 | 0.673 | 0.633 | 0.728 | 0.644 | 0.719 | 0.586 | 0.663 |
| DeepNovo | 0.623 | 0.610 | 0.750 | 0.694 | 0.630 | 0.731 | 0.679 | 0.742 | 0.602 | 0.673 |
| PointNovo | 0.626 | 0.606 | 0.779 | 0.712 | 0.644 | 0.733 | 0.730 | 0.768 | 0.589 | 0.687 |
| Casanovo · released | 0.689 | 0.586 | 0.684 | 0.679 | 0.629 | 0.721 | 0.668 | 0.749 | 0.603 | 0.667 |
| AdaNovo | 0.646 | 0.618 | 0.793 | 0.728 | 0.650 | 0.740 | 0.719 | 0.739 | 0.642 | 0.697 |
| Casanovo V2 · released | 0.760 | 0.676 | 0.752 | 0.755 | 0.706 | 0.785 | 0.748 | 0.790 | 0.681 | 0.739 |
| π-PrimeNovo · released | 0.784 | 0.729 | 0.802 | 0.801 | 0.763 | 0.815 | 0.822 | 0.846 | 0.734 | 0.788 |
| RefineNovo | 0.800 | 0.730 | 0.818 | 0.819 | 0.780 | 0.825 | 0.835 | 0.854 | 0.742 | 0.800 |
|
Nine-species benchmark original (DeepNovo, 2017) |
Peptide recall | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mus musculus | Homo sapiens | Saccharomyces cerevisiae | Methanosarcina mazei | Apis mellifera | Solanum lycopersicum | Vigna mungo | Bacillus subtilis | Candidatus Thiodiazotropha endoloripes | Average | |
| PEAKS | 0.197 | 0.277 | 0.428 | 0.356 | 0.287 | 0.403 | 0.362 | 0.387 | 0.203 | 0.322 |
| DeepNovo | 0.286 | 0.293 | 0.462 | 0.422 | 0.330 | 0.454 | 0.436 | 0.449 | 0.253 | 0.376 |
| PointNovo | 0.355 | 0.351 | 0.534 | 0.478 | 0.396 | 0.513 | 0.511 | 0.518 | 0.298 | 0.439 |
| Casanovo · released | 0.426 | 0.341 | 0.490 | 0.478 | 0.406 | 0.521 | 0.506 | 0.537 | 0.330 | 0.448 |
| AdaNovo | 0.467 | 0.373 | 0.593 | 0.496 | 0.431 | 0.530 | 0.546 | 0.528 | 0.372 | 0.481 |
| Casanovo V2 · released | 0.483 | 0.446 | 0.599 | 0.557 | 0.493 | 0.618 | 0.589 | 0.622 | 0.446 | 0.539 |
| π-PrimeNovo · released | 0.567 | 0.574 | 0.697 | 0.650 | 0.603 | 0.697 | 0.702 | 0.721 | 0.531 | 0.638 |
| RefineNovo | 0.583 | 0.581 | 0.709 | 0.667 | 0.616 | 0.705 | 0.720 | 0.736 | 0.549 | 0.653 |
Table2
Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing, page 9: Comparison of the performance on the 9-species-V2 benchmark datasets. AT stands for Autoregressive Transformer and NAT stands for non-autoregressive Transformer.
|
Nine-species benchmark revised (main) |
Amino acid precision | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mus musculus | Homo sapiens | Saccharomyces cerevisiae | Methanosarcina mazei | Apis mellifera | Solanum lycopersicum | Vigna mungo | Bacillus subtilis | Candidatus Thiodiazotropha endoloripes | Average | |
| Casanovo V2 · released | 0.813 | 0.872 | 0.915 | 0.877 | 0.823 | 0.891 | 0.891 | 0.888 | 0.791 | 0.862 |
| π-PrimeNovo · released | 0.839 | 0.893 | 0.932 | 0.908 | 0.862 | 0.909 | 0.931 | 0.921 | 0.827 | 0.891 |
| RefineNovo | 0.850 | 0.921 | 0.941 | 0.921 | 0.879 | 0.916 | 0.931 | 0.942 | 0.841 | 0.907 |
|
Nine-species benchmark revised (main) |
Peptide recall | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mus musculus | Homo sapiens | Saccharomyces cerevisiae | Methanosarcina mazei | Apis mellifera | Solanum lycopersicum | Vigna mungo | Bacillus subtilis | Candidatus Thiodiazotropha endoloripes | Average | |
| Casanovo V2 · released | 0.555 | 0.712 | 0.837 | 0.754 | 0.669 | 0.783 | 0.772 | 0.793 | 0.558 | 0.714 |
| π-PrimeNovo · released | 0.627 | 0.795 | 0.884 | 0.812 | 0.742 | 0.824 | 0.837 | 0.849 | 0.626 | 0.777 |
| RefineNovo | 0.637 | 0.805 | 0.895 | 0.827 | 0.762 | 0.829 | 0.862 | 0.856 | 0.637 | 0.790 |
In the paper: column amino acid precision, Ricebean: the original table only bolded RefineNovo (Ours) (0.931); π-PrimeNovo (Prime. (Zhang et al., 2025)) (0.931) ties with it and is bolded here too.
Table6 (HC-PT)
Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing, page 18: Performance comparison on the NovoBench benchmark (yeast test species). Scores for models marked with * are quoted from the NovoBench paper or original publications. CV denotes cross-validation results from the original PrimeNovo paper. “–” indicates data not available.
|
ProteomeTools HC-PT (NovoBench) |
Peptide precision |
|---|---|
| HC-PT | |
| Casanovo · quoted | 0.21 |
| InstaNovo · quoted | 0.57 |
| AdaNovo · quoted | 0.21 |
| π-HelixNovo · quoted | 0.21 |
| SearchNovo · quoted | 0.45 |
| π-PrimeNovo | 0.85 |
| RefineNovo | 0.88 |
Table6 (Nine-species)
Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing, page 18: Performance comparison on the NovoBench benchmark (yeast test species). Scores for models marked with * are quoted from the NovoBench paper or original publications. CV denotes cross-validation results from the original PrimeNovo paper. “–” indicates data not available.
| Nine-species benchmark | Peptide precision |
|---|---|
| Saccharomyces cerevisiae | |
| Casanovo · quoted | 0.48 |
| InstaNovo · quoted | 0.53 |
| AdaNovo · quoted | 0.50 |
| π-HelixNovo · quoted | 0.52 |
| SearchNovo · quoted | 0.55 |
| π-PrimeNovo CV · quoted | 0.58 |
| Casanovo pretrained · released | 0.60 |
| π-PrimeNovo | 0.70 |
| RefineNovo | 0.71 |
Table6 (Seven-species)
Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing, page 18: Performance comparison on the NovoBench benchmark (yeast test species). Scores for models marked with * are quoted from the NovoBench paper or original publications. CV denotes cross-validation results from the original PrimeNovo paper. “–” indicates data not available.
| Seven-species benchmark | Peptide precision |
|---|---|
| Saccharomyces cerevisiae | |
| Casanovo · quoted | 0.12 |
| AdaNovo · quoted | 0.17 |
| π-HelixNovo · quoted | 0.23 |
| SearchNovo · quoted | 0.26 |
| Casanovo pretrained · released | 0.05 |
| π-PrimeNovo | 0.09 |
| RefineNovo | 0.09 |
Papers describing it (2)
- Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing (2025, arXiv, preprint)
- Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing (2025, ICML 2025, ML conference)