Tandem Mass Spectra Representation Design for Transformer-Based De Novo Peptide Sequencing
thesis · 2026
| Date | 2026-07-27 |
| Type | thesis |
| Publisher | MSc thesis |
| Contribution | benchmark |
| Supervisor | Kaizhong Zhang |
| Link | https://hdl.handle.net/20.500.14721/40077 |
Abstract
De novo peptide sequencing predicts peptide sequences directly from tandem mass spectra without relying on a predefined peptide database. This thesis studies how spectral peaks should be represented for Transformer-based de novo peptide sequencing. A fixed Transformer backbone is used while the peak representation is changed, including additive m/z-intensity embedding, separated feature embeddings, precursor-normalized peak-mass features, zero-vector controls, and probability-aware embeddings. Experiments are conducted on a non-probability spectral-library dataset and a probability-augmented spectral dataset, with model dimensionality kept fixed across embedding variants. The results show that separating peak features improves performance over the additive baseline. Precursor-normalized mass information further improves the non-probability setting, while the best probability-augmented result is obtained when externally generated signal probability is combined with intensity and mass-based features. These results show that peak representation design meaningfully affects Transformer-based de novo peptide sequencing.
Methods and tools
- Spectrum representation ablation study: MSc thesis ablation of how MS/MS peaks should be represented for a Transformer de novo sequencer. Holds a Casanovo-style encoder-decoder backbone fixed and varies only the peak embedding: additive m/z-intensity, separated/concatenated m/z and intensity, precursor-normalised (relative) m/z, learnable relative m/z, zero-vector controls, and probability-aware embeddings fed with DpNovo’s signal-vs-noise peak probabilities. Finds that separating peak features beats the additive baseline, relative m/z adds a further gain, and the best result combines external signal probability with intensity and mass features: evidence that input representation, not just architecture, moves the needle on Transformer de novo accuracy.