InstaNovo-FM

algorithm · Transformer (encoder-only)

InstaNovo-FM: algorithm · Transformer (encoder-only). Self-supervised foundation model for bottom-up proteomics: an encoder-only transformer trained to reconstruct masked regions of tandem mass spectra under a physics-aware…

Self-supervised foundation model for bottom-up proteomics: an encoder-only transformer trained to reconstruct masked regions of tandem mass spectra under a physics-aware objective, over a corpus of 1.47 billion MS/MS spectra with 184.6 million high-confidence annotations. The embeddings capture fragmentation method, sequence properties and post-translational modifications without peptide labels, and support de novo sequencing, database-free identification and analytical run classification as downstream tasks.

Kind algorithm
Deep learning yes
Acquisition DDA
Family Transformer (encoder-only)

Code

Live stars, open issues and last-push figures are on the Code activity chart.

Checkpoints

Version Trained on Host Licence Size Checked
releases — GitHub release Apache-2.0 — live 2026-10-02

A host marked archival has a DOI and keeps what it is given. The others can move or disappear, which is why they are checked rather than merely listed. verified means the bytes were fetched and hashed on the date shown; live means only that the host answered when last asked.

Reported comparisons (2)

The comparison tables this method’s own papers print, standardised: every value on a 0-1 scale, methods down the side, the measure and then the species across. These are numbers papers report about themselves and their baselines. They are not a leaderboard, and they do not compare across tables: each was produced by a different group, on the dataset named in its corner, with each baseline either retrained, run from released weights or quoted from another paper. Where the paper says which, it follows the method’s name (hover it for the sentence); most papers do not say. Bold is the best value in a column and underline the runner-up, our ranking rather than the paper’s own marks.

Table S12

Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics, page 84: Peptide recall across the biological validation datasets for the de novo sequencing benchmark. The best model per dataset is in bold.

one per row (InstaNovo's application datasets) Peptide recall
HeLa degradome Candidatus Scalindua brodae Snake venoms HeLa single-shot Nanobodies Wound exudates
InstaNovo-FM fine-tuned 0.821 0.740 0.230 0.662 0.528 0.366
InstaNovo v1.2 0.868 0.764 0.227 0.708 0.517 0.431
Casanovo 0.767 0.743 0.217 0.531 0.465 0.262
XuanjiNovo 0.808 0.741 0.233 0.673 0.529 0.466

No accession is printed for the six validation sets. They are mapped from the Methods, which call them InstaNovo's application datasets minus 'the Immuno and Herceptin datasets': GluC is InstaNovo's 'HeLa GluC degradome', TPL Antibodies its nanobodies, and Hela QC, by elimination, its HeLa single-shot set.

Table S13

Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics, page 84: Peptide recall across the biological validation datasets for the de novo sequencing InstaNovo-FM and InstaNovo model variant benchmarks. The best model per dataset is in bold. We abbreviate “InstaNovo (FM size matched)” to “IN (FM-SM)” to conserve page space.

one per row (InstaNovo's application datasets) Peptide recall
HeLa degradome Candidatus Scalindua brodae Snake venoms HeLa single-shot Nanobodies Wound exudates
InstaNovo-FM fine-tuned 0.821 0.740 0.230 0.662 0.528 0.366
InstaNovo-FM from scratch 0.819 0.747 0.239 0.657 0.522 0.354
InstaNovo-FM frozen 0.775 0.726 0.156 0.577 0.474 0.266
InstaNovo 0.821 0.756 0.219 0.652 0.520 0.345
InstaNovo FM size matched 0.823 0.748 0.226 0.650 0.516 0.362

No accession is printed for the six validation sets. They are mapped from the Methods, which call them InstaNovo's application datasets minus 'the Immuno and Herceptin datasets': GluC is InstaNovo's 'HeLa GluC degradome', TPL Antibodies its nanobodies, and Hela QC, by elimination, its HeLa single-shot set.

Paper describing it

Authors (12)

Mechiel Nieuwoudt, Marco Reverenna, Divanisha Patel, Rachel Catzel, Isaac H.J. Houngue, Jemma Daniel, Kevin Eloff, Alberto Santos, Nicolas Lopez Carranza, Timothy P. Jenkins, Jeroen Van Goey, Konstantinos Kalogeropoulos

Seen in the charts

Back to the full map

Back to top