Casanovo

algorithm · Transformer (AR)

Casanovo: algorithm · Transformer (AR). First Transformer

First Transformer

Kind algorithm
Deep learning yes
Acquisition DDA
Family Transformer (AR)

Code

Live stars, open issues and last-push figures are on the Code activity chart.

Checkpoints

Version Trained on Host Licence Size Checked
nine-species — Zenodo archival CC-BY-4.0 5.1 GB live 2026-10-02
data set and weights — Zenodo archival Apache-2.0 3.9 GB live 2026-10-02
MassIVE-KB splits — institutional not stated — live 2026-10-02
v4.0.0 MassIVE-KB v1 GitHub release Apache-2.0 — live 2026-10-02
v4.2.0 ~2M PSMs from MassIVE-KB v1 + v2.0.15 GitHub release Apache-2.0 — live 2026-10-02
v5.0.0 — GitHub release Apache-2.0 — live 2026-10-02
v5.2.0 the default –model orbitrap selector from v5.2.0 onward GitHub release Apache-2.0 — live 2026-10-02
v5.2.0 the –model timstof selector from v5.2.0 onward GitHub release Apache-2.0 575 MB live 2026-10-06

A host marked archival has a DOI and keeps what it is given. The others can move or disappear, which is why they are checked rather than merely listed. verified means the bytes were fetched and hashed on the date shown; live means only that the host answered when last asked.

Benchmarks

  • denovo_benchmarks: median peptide-level average precision 0.777 over 86 datasets, median rank 7 of 14 (version 5.0.0).
  • ProteoBench, on the nine-species benchmark, ProteoBench selection: peptide-level AUC 0.901; precision 0.650 at 100% coverage; amino-acid AUC 0.951 (version 4.0, beam search, submitted 2026-08-07).

Both are mass-based matches on the tool’s most recent run. What these numbers mean.

Reported comparison

The comparison table this method’s own papers print, standardised: every value on a 0-1 scale, methods down the side, the measure and then the species across. These are numbers papers report about themselves and their baselines. They are not a leaderboard, and they do not compare across tables: each was produced by a different group, on the dataset named in its corner, with each baseline either retrained, run from released weights or quoted from another paper. Where the paper says which, it follows the method’s name (hover it for the sentence); most papers do not say. Bold is the best value in a column and underline the runner-up, our ranking rather than the paper’s own marks.

Table2

De novo mass spectrometry peptide sequencing with a transformer model, page 6: Empirical comparison of Casanovo, DeepNovo and PointNovo. The table lists the peptide-level and amino acid-level precision of three competing models and coverage of Casanovo with precursor m/z filtering on all nine benchmark cross-validation folds. Each fold’s test set contains spectra from a single species, with nearly disjoint sets of peptides between species. For cross-validation folds corresponding to mouse and human, five models were trained with different random initializations. For these species, we report standard deviation of the performance measures.

Nine-species benchmark Peptide precision
Mus musculus Homo sapiens Saccharomyces cerevisiae Methanosarcina mazei Apis mellifera Solanum lycopersicum Vigna mungo Bacillus subtilis Candidatus Thiodiazotropha endoloripes
DeepNovo · released 0.286 0.293 0.462 0.422 0.330 0.454 0.436 0.449 0.253
PointNovo · quoted 0.355 0.351 0.534 0.478 0.396 0.513 0.511 0.518 0.298
Casanovo 0.665 0.683 0.824 0.771 0.732 0.771 0.798 0.805 0.695
Nine-species benchmark Peptide coverage
Mus musculus Homo sapiens Saccharomyces cerevisiae Methanosarcina mazei Apis mellifera Solanum lycopersicum Vigna mungo Bacillus subtilis Candidatus Thiodiazotropha endoloripes
Casanovo 0.666 0.537 0.681 0.630 0.557 0.557 0.547 0.671 0.534
Nine-species benchmark Peptide precision at coverage 1
Mus musculus Homo sapiens Saccharomyces cerevisiae Methanosarcina mazei Apis mellifera Solanum lycopersicum Vigna mungo Bacillus subtilis Candidatus Thiodiazotropha endoloripes
Casanovo 0.443 0.367 0.561 0.486 0.408 0.460 0.437 0.540 0.371
Nine-species benchmark Amino acid precision
Mus musculus Homo sapiens Saccharomyces cerevisiae Methanosarcina mazei Apis mellifera Solanum lycopersicum Vigna mungo Bacillus subtilis Candidatus Thiodiazotropha endoloripes
DeepNovo · released 0.623 0.610 0.750 0.694 0.630 0.731 0.679 0.742 0.602
PointNovo · quoted 0.626 0.606 0.779 0.712 0.644 0.733 0.730 0.768 0.589
Casanovo 0.899 0.898 0.952 0.935 0.920 0.929 0.920 0.943 0.908
Nine-species benchmark Amino acid precision at coverage 1
Mus musculus Homo sapiens Saccharomyces cerevisiae Methanosarcina mazei Apis mellifera Solanum lycopersicum Vigna mungo Bacillus subtilis Candidatus Thiodiazotropha endoloripes
Casanovo 0.562 0.424 0.591 0.518 0.461 0.471 0.442 0.573 0.405

Papers describing it (7)

Papers using it (9)

Applications and evaluations that ran this method. They are not counted among its authors below.

Authors (17)

Melih Yilmaz, William E. Fondrie, Wout Bittremieux, Sewoong Oh, William Stafford Noble, Carlo F. Melendez, Rowan Nelson, Varun Ananth, Justin Sanders, Gwenneth Straub, Chris Hsu, Daniela Klaproth-Andrade, Michael Riffle, Bo Wen, Lingwen Xu, Michael J. MacCoss, Marina Pominova

Seen in the charts

Back to the full map

Back to top