Foundation model for mass spectrometry proteomics

preprint · arXiv · 2025

preprint · arXiv · 2025. Justin Sanders et al. Mass spectrometry is the dominant technology in the field of proteomics, enabling high-throughput analysis of…
Date 2025-05-16
Type preprint
Venue arXiv
Publisher arXiv
Contribution algorithm
DOI 10.48550/arXiv.2505.10848
Citations (OpenAlex) 3

Abstract

Mass spectrometry is the dominant technology in the field of proteomics, enabling high-throughput analysis of the protein content of complex biological samples. Due to the complexity of the instrumentation and resulting data, sophisticated computational methods are required for the processing and interpretation of acquired mass spectra. Machine learning has shown great promise to improve the analysis of mass spectrometry data, with numerous purpose-built methods for improving specific steps in the data acquisition and analysis pipeline reaching widespread adoption. Here, we propose unifying various spectrum prediction tasks under a single foundation model for mass spectra. To this end, we pre-train a spectrum encoder using de novo sequencing as a pre-training task. We then show that using these pre-trained spectrum representations improves our performance on the four downstream tasks of spectrum quality prediction, chimericity prediction, phosphorylation prediction, and glycosylation status prediction. Finally, we perform multi-task fine-tuning and find that this approach improves the performance on each task individually. Overall, our work demonstrates that a foundation model for tandem mass spectrometry proteomics trained on de novo sequencing learns generalizable representations of spectra, improves performance on downstream tasks where training data is limited, and can ultimately enhance data acquisition and analysis in proteomics experiments.

Authors

  1. Justin Sanders · University of Washington
  2. Melih Yilmaz · University of Washington
  3. Jacob H. Russell · University of Washington
  4. Wout Bittremieux · Indiana University, University of Antwerp, University of California San Diego
  5. William E. Fondrie · Talus Bioscience
  6. Nicholas M. Riley · University of Washington
  7. Sewoong Oh · University of Washington
  8. William Stafford Noble · University of Washington

Methods and tools

  • Casanovo Foundation: Foundation model for tandem-MS proteomics: pre-trained the Casanovo spectrum encoder on 30M labelled spectra from MassIVE-KB, then reused the encoder off-the-shelf for downstream tasks (de novo sequencing, spectrum quality, chimericity, phosphorylation, glycosylation prediction). Successor in spirit to Casanovo v1/v2/v5.

Cites (7)

Seen in the charts

Back to the full map

Back to top