Latent Imputation before Prediction: A New Computational Paradigm for De Novo Peptide Sequencing

ML conference · ICML 2025 · 2025

ML conference · ICML 2025 · 2025. Ye Du et al. De novo peptide sequencing is a fundamental computational technique for ascertaining amino acid sequences of…
Date 2025-07-13
Type ML conference
Venue ICML 2025
Publisher PMLR
Contribution algorithm
Link https://proceedings.mlr.press/v267/du25g.html

Abstract

De novo peptide sequencing is a fundamental computational technique for ascertaining amino acid sequences of peptides directly from tandem mass spectrometry data, eliminating the need for reference databases. Cutting-edge models encode the observed mass spectra into latent representations from which peptides are predicted auto-regressively. However, the issue of missing fragmentation, attributable to factors such as suboptimal fragmentation efficiency and instrumental constraints, presents a formidable challenge in practical applications. To tackle this obstacle, we propose a novel computational paradigm called \(\\underline{\\textbf{L}}\)atent \(\\underline{\\textbf{I}}\)mputation before \(\\underline{\\textbf{P}}\)rediction (LIPNovo). LIPNovo is devised to compensate for missing fragmentation information within observed spectra before executing the final peptide prediction. Rather than generating raw missing data, LIPNovo performs imputation in the latent space, guided by the theoretical peak profile of the target peptide sequence. The imputation process is conceptualized as a set-prediction problem, utilizing a set of learnable peak queries to reason about the relationships among observed peaks and directly generate the latent representations of theoretical peaks through optimal bipartite matching. In this way, LIPNovo manages to supplement missing information during inference and thus boosts performance. Despite its simplicity, experiments on three benchmark datasets demonstrate that LIPNovo outperforms state-of-the-art methods by large margins. Code is available at https://github.com/usr922/LIPNovo.

Authors

  1. Ye Du · The Hong Kong Polytechnic University
  2. Chen Yang · The Hong Kong Polytechnic University
  3. Nanxi Yu · The Hong Kong Polytechnic University
  4. Wanyu Lin · The Hong Kong Polytechnic University
  5. Qian Zhao · The Hong Kong Polytechnic University
  6. Shujun Wang · The Hong Kong Polytechnic University

Methods and tools

Data used

  • ProteomeTools (HC-PT (NovoBench)) · no public address
  • Seven-species benchmark (NovoBench split) · no public address

Cited by (1)

Seen in the charts

Back to the full map

Back to top