π-PrimeNovo: an accurate and efficient non-autoregressive deep learning model for de novo peptide sequencing

preprint · bioRxiv · 2024

preprint · bioRxiv · 2024. Xiang Zhang (Shanghai AI Lab) et al. Peptide sequencing via tandem mass spectrometry (MS/MS) is fundamental in proteomics data analysis, playing a…
Date 2024-05-17
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution algorithm
DOI 10.1101/2024.05.17.594647
Citations (OpenAlex) 1

Abstract

Peptide sequencing via tandem mass spectrometry (MS/MS) is fundamental in proteomics data analysis, playing a pivotal role in unraveling the complex world of proteins within biological systems. In contrast to conventional database searching methods, deep learning models excel in de novo sequencing peptides absent from existing databases, thereby facilitating the identification and analysis of novel peptide sequences. Current deep learning models for peptide sequencing predominantly use an autoregressive generation approach, where early errors can cascade, largely affecting overall sequence accuracy. And the usage of sequential decoding algorithms such as beam search suffers from the low inference speed. To address this, we introduce{pi} -PrimeNovo, a non-autoregressive Transformer-based deep learning model designed to perform accurate and efficient de novo peptide sequencing. With the proposed novel architecture,{pi} -PrimeNovo achieves significantly higher accuracy and up to 69x faster sequencing compared to the state-of-the-art methods. This remarkable speed makes it highly suitable for computation-extensive peptide sequencing tasks such as metaproteomic research, where{pi} -PrimeNovo efficiently identifies the microbial species-specific peptides. Moreover,{pi} -PrimeNovo has been demonstrated to have a powerful capability in accurately mining phosphopeptides in a non-enriched phosphoproteomic dataset, showing an alternative solution to detect low-abundance post-translational modifications (PTMs). We suggest that this work not only advances the development of peptide sequencing techniques but also introduces a transformative computational model with wide-range implications for biological research.

Authors

  1. Xiang Zhang (Shanghai AI Lab) · Fudan University, Shanghai Artificial Intelligence Laboratory, University of British Columbia
  2. Tianze Ling · Beijing Institute of Lifeomics, State Key Laboratory of Medical Proteomics, Tsinghua University
  3. Zhi Jin · Shanghai Artificial Intelligence Laboratory, Soochow University
  4. Sheng Xu · Fudan University, Shanghai Artificial Intelligence Laboratory
  5. Zhiqiang Gao · Shanghai Artificial Intelligence Laboratory
  6. Boyan Sun · Beijing Institute of Lifeomics, State Key Laboratory of Medical Proteomics
  7. Zijie Qiu · Fudan University, Shanghai Artificial Intelligence Laboratory
  8. Nanqing Dong · Shanghai Artificial Intelligence Laboratory
  9. Guangshuai Wang · Shanghai Artificial Intelligence Laboratory
  10. Guibin Wang · Beijing Institute of Lifeomics
  11. Leyuan Li · Beijing Institute of Lifeomics, State Key Laboratory of Medical Proteomics
  12. Muhammad Abdul-Mageed · Mohamed bin Zayed University of Artificial Intelligence, University of British Columbia
  13. Laks V.S. Lakshmanan · University of British Columbia
  14. Wanli Ouyang · Shanghai Artificial Intelligence Laboratory
  15. Cheng Chang · Beijing Institute of Lifeomics, International Academy of Phronesis Medicine (Guangdong), National Center for Protein Sciences (Beijing), State Key Laboratory of Medical Proteomics
  16. Siqi Sun · Fudan University, Shanghai Artificial Intelligence Laboratory

Methods and tools

Cites (17)

Cited by (9)

Seen in the charts

Back to the full map

Back to top