DyCoNovo: a De Novo Peptide Prediction Model Based on Dynamic Convolution and Phased Contrastive Learning
peer-reviewed · 2025 18th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI) · 2025
| Date | 2025-10-25 |
| Type | peer-reviewed |
| Venue | 2025 18th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI) |
| Publisher | IEEE |
| Contribution | algorithm |
| DOI | 10.1109/cisp-bmei68103.2025.11259143 |
| Citations (OpenAlex) | 0 |
Abstract
De novo sequencing is an important research field in proteomics, aiming to directly infer the amino acid sequence of unknown polypeptides solely based on tandem mass spectrometry (MS/MS) without relying on databases. Traditional peptide sequence prediction methods often rely on manual feature extraction and statistical models, which have certain limitations. In recent years, end-to-end models based on deep learning have significantly improved prediction accuracy. However, most existing models focus on the global Transformer architecture, with insufficient characterization of local peak cluster details. Additionally, the presence of considerable noise in mass spectrometry data and the large differences in amino acid sequences between different species leave room for improvement in accuracy on species-specific datasets. The key innovation of DyCoNovo is the introduction of a lightweight dynamic convolution network module on the basis of CasaNovo, providing spectral representations that combine global context and local details for de novo peptide sequencing research. Meanwhile, a phased contrastive learning strategy is adopted, which enhances the model’s generalization ability across different species datasets and under low signal-to-noise ratio data, as well as improves training stability. On a standard benchmark covering nine species and approximately 1.5 million spectra, compared with CasaNovo, DyCoNovo achieves an average increase of 13.9 % in amino acid accuracy and 19 % in peptide accuracy. Compared with ContraNovo, the latest model adopting contrastive learning strategies, DyCoNovo shows an average increase of 2.4 % in amino acid accuracy and 2.1 % in peptide accuracy. DyCoNovo demonstrates more significant advantages in low signal-to-noise ratio data, verifying the effectiveness of dynamic convolution in robust modeling of local peak shapes. This model effectively improves the accuracy of de novo peptide sequencing.
Methods and tools
- DyCoNovo: A deep learning de novo sequencing model that adds lightweight dynamic convolution for local peak clusters and phased contrastive learning to a Transformer.