Transformer (encoder-only)

3 methods · 2024–2026

Transformer (encoder-only): Transformer spectrum encoders trained without a decoder, usually self-supervised, to produce representations of a spectrum that other tasks reuse. The output is an embedding, not a peptide.

Transformer spectrum encoders trained without a decoder, usually self-supervised, to produce representations of a spectrum that other tasks reuse. The output is an embedding, not a peptide.

The earliest of its 3 methods is CPC spectrum encoder pretraining (2024); 2 more have followed.

Methods 3
Papers describing them 3
Authors 14
Active 2024-06-16 to 2026-09-03
Deep learning 3 of 3
Kinds adjacent (2), algorithm
Acquisition DDA (3)

Methods (3)

Oldest first, by the paper that describes each one.

  • CPC spectrum encoder pretraining (2024): Unsupervised pretraining of a transformer spectrum encoder by Contrastive Predictive Coding, aimed at the two things that limit supervised de novo sequencing: scarce training data for post-translational modifications, and noisy or incomplete spectra. Exploits the large volume of unlabelled tandem mass spectra that supervised training cannot use. Evaluated on spectral library search rather than sequencing, where on 9-species-V2 it beats OpenMS by 3.73% average amino-acid precision and 4.15% recall, gains 3.8% peptide-level recall, and scales better on inference speed across dataset sizes.
  • Fragment-ion and amino-acid probability models (2026): Two local probability-prediction tasks for tandem mass spectra, aimed at telling sequence-informative fragment evidence from noise and from merely unobserved fragments. Fragment-Ion Probability scores whether a peak-supported mass position is a sequence-informative fragment ion; Amino Acid Probability scores whether two mass-consistent positions are adjacent fragments joined by a candidate residue. Encoder-only Transformers outperform CNN baselines on both, and feeding the probabilities to a separate sequencer raises amino-acid recall, amino-acid precision and peptide recall.
  • InstaNovo-FM (2026): Self-supervised foundation model for bottom-up proteomics: an encoder-only transformer trained to reconstruct masked regions of tandem mass spectra under a physics-aware objective, over a corpus of 1.47 billion MS/MS spectra with 184.6 million high-confidence annotations. The embeddings capture fragmentation method, sequence properties and post-translational modifications without peptide labels, and support de novo sequencing, database-free identification and analytical run classification as downstream tasks.

Papers describing them (3)

Authors (14)

Abel Legese Shibiru, Alberto Santos, Divanisha Patel, Isaac H.J. Houngue, Jemma Daniel, Jeroen Van Goey, Kevin Eloff, Konstantinos Kalogeropoulos, Marco Reverenna, Mechiel Nieuwoudt, Nicolas Lopez Carranza, Rachel Catzel, Shichao Wang, Timothy P. Jenkins

Seen in the charts

Back to the full map

Back to top