Protein Language Model-Aligned Spectra Embeddings for De Novo Peptide Sequencing

preprint · bioRxiv · 2025

preprint · bioRxiv · 2025. Navid NaderiAlizadeh et al. We consider the problem of de novo peptide sequencing in tandem mass spectrometry, where the goal is to…
Date 2025-10-03
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution algorithm
DOI 10.1101/2025.10.01.679857
Citations (OpenAlex) 0

Abstract

We consider the problem of de novo peptide sequencing in tandem mass spectrometry, where the goal is to predict the underlying peptide sequence given a spectrum’s fragment peaks and precursor information. We present PLMNovo, a constrained learning framework that leverages pre-trained protein language models (PLMs) to guide the training process. In particular, we cast peptide-spectrum matching as a constrained optimization problem that enforces alignment between spectrum and peptide embeddings produced by a spectrum encoder and a PLM, respectively. We use a Lagrangian primal-dual algorithm to train the spectrum encoder and the peptide decoder by solving the proposed constrained learning problem, while optionally fine-tuning the pre-trained PLM. Through numerical experiments on established benchmarks, we demonstrate that PLMNovo outperforms several state-of-the-art deep learning-based de novo sequencing algorithms.

Authors

  1. Navid NaderiAlizadeh · Duke University
  2. Christian Dallago · Duke University
  3. Erik J. Soderblom · Duke University
  4. Scott H. Soderling · Duke University

Methods and tools

  • PLMNovo: Casts peptide-spectrum matching as a constrained optimisation that aligns the spectrum encoder’s embeddings with those of a pre-trained protein language model, and trains encoder and decoder with a Lagrangian primal-dual algorithm, optionally fine-tuning the language model.

Data used

Cites (20)

Seen in the charts

Back to the full map

Back to top