Tandem Mass Spectra Representation Design for Transformer-Based De Novo Peptide Sequencing

thesis · 2026

thesis · 2026. Xinlue Shen. De novo peptide sequencing predicts peptide sequences directly from tandem mass spectra without relying on a…
Date 2026-07-27
Type thesis
Publisher MSc thesis
Contribution benchmark
Supervisor Kaizhong Zhang
Link https://hdl.handle.net/20.500.14721/40077

Abstract

De novo peptide sequencing predicts peptide sequences directly from tandem mass spectra without relying on a predefined peptide database. This thesis studies how spectral peaks should be represented for Transformer-based de novo peptide sequencing. A fixed Transformer backbone is used while the peak representation is changed, including additive m/z-intensity embedding, separated feature embeddings, precursor-normalized peak-mass features, zero-vector controls, and probability-aware embeddings. Experiments are conducted on a non-probability spectral-library dataset and a probability-augmented spectral dataset, with model dimensionality kept fixed across embedding variants. The results show that separating peak features improves performance over the additive baseline. Precursor-normalized mass information further improves the non-probability setting, while the best probability-augmented result is obtained when externally generated signal probability is combined with intensity and mass-based features. These results show that peak representation design meaningfully affects Transformer-based de novo peptide sequencing.

Authors

  1. Xinlue Shen · University of Western Ontario

Methods and tools

  • Spectrum representation ablation study: MSc thesis ablation of how MS/MS peaks should be represented for a Transformer de novo sequencer. Holds a Casanovo-style encoder-decoder backbone fixed and varies only the peak embedding: additive m/z-intensity, separated/concatenated m/z and intensity, precursor-normalised (relative) m/z, learnable relative m/z, zero-vector controls, and probability-aware embeddings fed with DpNovo’s signal-vs-noise peak probabilities. Finds that separating peak features beats the additive baseline, relative m/z adds a further gain, and the best result combines external signal probability with intensity and mass features: evidence that input representation, not just architecture, moves the needle on Transformer de novo accuracy.

Methods it uses

  • Casanovo: First Transformer
  • DpNovo: Transformer + dynamic programming

Seen in the charts

Back to the full map

Back to top