π-MNovo improves de novo peptide sequencing through microbial-domain adaptation and evidence-guided candidate selection

preprint · bioRxiv · 2026

preprint · bioRxiv · 2026. Jingwen Ye et al. The high taxonomic and strain-level diversity of microbial communities makes it difficult for reference…
Date 2026-09-24
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution algorithm
DOI 10.64898/2026.09.22.753549
Citations (OpenAlex) 0

Abstract

The high taxonomic and strain-level diversity of microbial communities makes it difficult for reference databases to fully represent the protein sequences present in metaproteomic samples, limiting database-dependent peptide identification. De novo peptide sequencing can recover peptide sequences directly from tandem mass spectra without relying on reference databases, providing complementary peptide evidence for metaproteomics. However, most existing de novo sequencing models were developed largely from non-microbial proteomic data and lack specific adaptation to microbial spectra. Here, we established a dedicated microbial spectral resource comprising more than 10 million annotated high-quality tandem mass spectra from 72 cultured microbial isolates, with peptide digests from each isolate separated into five high-pH reversed-phase fractions to increase the opportunity for deeper and more diverse peptide sampling. Using this resource, we developed {pi}-MNovo, which combines adaptation to microbial spectra with evidence-guided candidate selection to improve full-length peptide sequencing precision. On a seven-species external benchmark, {pi}-MNovo achieved 66.30% complete-peptide recall under the same residue-mass-based complete-peptide matching criterion applied to all models, exceeding four published models by 6.9-21.7% in relative recall, with consistent gains across species and peptide-length groups. In two independent metaproteomic datasets, {pi}-MNovo increased reference-supported taxonomic coverage, while synthetic-community analysis recapitulated the designed abundance ranking. From pFind-unidentified spectra, {pi}-MNovo recovered 1,587 pFind-unreported, reference-matched peptides that passed spectrum-evidence filtering. These results show that {pi}-MNovo expands recoverable peptide evidence and enhances the utility of de novo sequencing for metaproteomic analysis.

Authors

  1. Jingwen Ye · Dalian Minzu University, National Center for Protein Sciences (Beijing)
  2. Xin Zhang (NCPSB) · National Center for Protein Sciences (Beijing)
  3. Boyan Sun · Beijing Institute of Lifeomics, State Key Laboratory of Medical Proteomics
  4. Tianze Ling · Beijing Institute of Lifeomics, State Key Laboratory of Medical Proteomics, Tsinghua University
  5. Zhendong Liang · Peng Cheng Laboratory, Tsinghua Shenzhen International Graduate School, Tsinghua University
  6. Jingyi Dong · Hebei University, National Center for Protein Sciences (Beijing)
  7. Yuan Zheng · National Center for Protein Sciences (Beijing)
  8. Hui Jin · National Center for Protein Sciences (Beijing)
  9. Liming Jin · Dalian Minzu University
  10. Zikai Hao · Beijing Institute of Technology
  11. Leyuan Li · Beijing Institute of Lifeomics, State Key Laboratory of Medical Proteomics
  12. Cheng Chang · Beijing Institute of Lifeomics, International Academy of Phronesis Medicine (Guangdong), National Center for Protein Sciences (Beijing), State Key Laboratory of Medical Proteomics

Methods and tools

  • π-MNovo: De novo sequencer adapted to microbial spectra, addressing the fact that existing models were trained largely on non-microbial proteomic data. Built on a purpose-made resource of over 10 million annotated spectra from 72 cultured microbial isolates, each digest split into five high-pH reversed-phase fractions for deeper peptide sampling, and paired with evidence-guided candidate selection. Reports 66.30% complete-peptide recall on a seven-species benchmark, 6.9-21.7% relative above four published models, and recovered 1,587 reference-matched peptides from spectra pFind left unidentified.

Seen in the charts

Back to the full map

Back to top