Transformer (encoder-only)
3 methods · 2024–2026
Transformer (encoder-only): Transformer spectrum encoders trained without a decoder, usually self-supervised, to produce representations of a spectrum that other tasks reuse. The output is an embedding, not a peptide.
Transformer spectrum encoders trained without a decoder, usually self-supervised, to produce representations of a spectrum that other tasks reuse. The output is an embedding, not a peptide.
The earliest of its 3 methods is CPC spectrum encoder pretraining (2024); 2 more have followed.
| Methods | 3 |
| Papers describing them | 3 |
| Authors | 14 |
| Active | 2024-06-16 to 2026-09-03 |
| Deep learning | 3 of 3 |
| Kinds | adjacent (2), algorithm |
| Acquisition | DDA (3) |
Methods (3)
Oldest first, by the paper that describes each one.
- CPC spectrum encoder pretraining (2024): Unsupervised pretraining of a transformer spectrum encoder by Contrastive Predictive Coding, aimed at the two things that limit supervised de novo sequencing: scarce training data for post-translational modifications, and noisy or incomplete spectra. Exploits the large volume of unlabelled tandem mass spectra that supervised training cannot use. Evaluated on spectral library search rather than sequencing, where on 9-species-V2 it beats OpenMS by 3.73% average amino-acid precision and 4.15% recall, gains 3.8% peptide-level recall, and scales better on inference speed across dataset sizes.
- Fragment-ion and amino-acid probability models (2026): Two local probability-prediction tasks for tandem mass spectra, aimed at telling sequence-informative fragment evidence from noise and from merely unobserved fragments. Fragment-Ion Probability scores whether a peak-supported mass position is a sequence-informative fragment ion; Amino Acid Probability scores whether two mass-consistent positions are adjacent fragments joined by a candidate residue. Encoder-only Transformers outperform CNN baselines on both, and feeding the probabilities to a separate sequencer raises amino-acid recall, amino-acid precision and peptide recall.
- InstaNovo-FM (2026): Self-supervised foundation model for bottom-up proteomics: an encoder-only transformer trained to reconstruct masked regions of tandem mass spectra under a physics-aware objective, over a corpus of 1.47 billion MS/MS spectra with 184.6 million high-confidence annotations. The embeddings capture fragmentation method, sequence properties and post-translational modifications without peptide labels, and support de novo sequencing, database-free identification and analytical run classification as downstream tasks.
Papers describing them (3)
- Enhancing Peptide Mass Spectra Encoder through Pretraining using Contrastive Predictive Coding (2024, thesis)
- Spectra Fragment-Ion and Amino Acid Probability Prediction for Peptide Sequencing (2026, thesis)
- Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics (2026, bioRxiv, preprint)