LIPNovo+: Self-Reflective Latent Imputation for Robust De Novo Peptide Sequencing

preprint · SSRN Electronic Journal · 2026

preprint · SSRN Electronic Journal · 2026. Ye Du et al. De novo peptide sequencing is an important pattern recognition task in computational proteomics, enabling…
Date 2026-06-06
Type preprint
Venue SSRN Electronic Journal
Publisher Elsevier BV (SSRN)
Contribution algorithm
DOI 10.2139/ssrn.6890054
Citations (OpenAlex) 0

Abstract

De novo peptide sequencing is an important pattern recognition task in computational proteomics, enabling direct identification of peptide sequences from tandem mass spectrometry (MS/MS) data for biomarker discovery, antibody characterization, and analysis of peptides absent from reference databases. State-of-the-art models encode observed spectra into latent representations for peptide prediction. However, the issue of missing fragmentation, attributable to factors such as suboptimal fragmentation efficiency and instrumental constraints, presents a formidable challenge in practical applications. To tackle this obstacle, we propose Latent Imputation before Prediction (LIPNovo), a new computational paradigm that compensates for missing fragmentation information before peptide prediction. Instead of generating raw missing data, LIPNovo performs imputation in the latent space, guided by the theoretical peak profile of the target peptide. The imputation process is formulated as a set-prediction problem, where learnable peak queries reason over observed peaks and generate latent representations of theoretical peaks through optimal bipartite matching. To move beyond static missingness patterns seen during training, we further propose LIPNovo+, a self-reflective extension that identifies fragmentation sites with unreliable latent imputation and refocuses learning on these vulnerable regions through a reflection-guided curriculum. Across four benchmark datasets, LIPNovo and LIPNovo+ consistently outperform state-of-the art methods, with gains of up to +15% in peptide precision. Code is available at https://github.com/usr922/LIPNovo.

Authors

  1. Ye Du · The Hong Kong Polytechnic University
  2. Qian Niu
  3. Chen Yang · The Hong Kong Polytechnic University
  4. Nanxi Yu · The Hong Kong Polytechnic University
  5. Shujun Wang · The Hong Kong Polytechnic University

Methods and tools

  • LIPNovo: Latent imputation
  • LIPNovo+: Self-reflective extension of LIPNovo that finds the fragmentation sites where latent imputation is unreliable and refocuses training on them through a reflection-guided curriculum, so the model is not limited to the missingness patterns seen during training.

Cites (19)

Seen in the charts

Back to the full map

Back to top