MassNet: billion-scale AI-friendly mass spectral corpus enables robust de novo peptide sequencing
preprint · bioRxiv · 2025
| Date | 2025-06-20 |
| Type | preprint |
| Venue | bioRxiv |
| Publisher | Cold Spring Harbor Laboratory |
| Contribution | algorithm |
| DOI | 10.1101/2025.06.20.660691 |
| Citations (OpenAlex) | 2 |
Abstract
Breakthroughs in artificial intelligence (AI) for natural language processing and computer vision have been largely driven by high-quality, large-scale datasets such as OpenWebText and ImageNet. Inspired by this, we present MassNet, a foundational resource for proteomics designed to accelerate deep learning applications. MassNet is the largest known corpus of data-dependent acquisition (DDA) mass spectrometry (MS) data, derived from ~30 TB of raw files and comprising 1.54 billion MS/MS spectra, resulting in 558 million peptide-spectrum matches (PSMs) across 35 species, including animals, plants, and microbes. Within the human subset, MassNet includes more than 1.7 million precursors and 19,966 proteins, covering 98% of annotated human proteins. To enable efficient AI training, we developed the Mass Spectrometry Data Tensor (MSDT), a structured format based on Parquet that enables standardized, high-performance batch access and seamless integration with GPU and TPU platforms for distributed training. We further extended MassNet to support de novo peptide sequencing, which infers peptide sequences directly from MS/MS spectra without reference databases, and is critical for discovering novel proteins, characterizing non-model organisms, and identifying post-translational modifications (PTMs). We introduce XuanjiNovo, a non-autoregressive Transformer model that leverages a curriculum learning strategy to enhance training stability. By dynamically adjusting learning difficulty based on model performance, XuanjiNovo achieves smooth convergence on complex, multi-distributional data without manual hyperparameter tuning. Trained on 100 million PSMs from the MassNet, it consistently outperforms state-of-the-art methods across diverse benchmarking tasks. Peptide recall exceeds 0.8 on the Bacteroides thetaiotaomicron and Zea mays datasets. On human data acquired using the Orbitrap Astral platform, XuanjiNovo achieves achieves 38.8% to 144.3% improvement over existing models. MassNet represents the first large-scale, standardized foundational dataset in proteomics, marking a critical milestone in the integration of artificial intelligence into proteomics research.
Methods and tools
- XuanjiNovo: Billion-scale pretraining
Cites (26)
- InstaNovo enables diffusion-powered de novo peptide sequencing in large-scale proteomics experiments (2025) crossref
- π-PrimeNovo: an accurate and efficient non-autoregressive deep learning model for de novo peptide sequencing (2025) crossref
- A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models (2024) crossref
- Transforming de novo peptide sequencing by explainable AI (2024) crossref
- Sequence-to-sequence translation from mass spectra to peptides with a transformer model (2024) crossref
- π-PrimeNovo: an accurate and efficient non-autoregressive deep learning model for de novo peptide sequencing (2024) semanticscholar
- ContraNovo: A Contrastive Learning Approach to Enhance De Novo Peptide Sequencing (2024) crossref
- AdaNovo: Adaptive De Novo Peptide Sequencing with Conditional Mutual Information (2024) crossref
- Bidirectional de novo peptide sequencing using a transformer model (2024) crossref
- Introducing π-HelixNovo for practical large-scale de novo peptide sequencing (2024) crossref
- Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing (2024) crossref
- Accurate de novo peptide sequencing using fully convolutional neural networks (2023) crossref
- Mitigating the missing-fragmentation problem in de novo peptide sequencing with a two-stage graph-based deep learning model (2023) crossref
- SeqNovo: De Novo Peptide Sequencing Prediction in IoMT via Seq2Seq (2023) crossref
- Introducing PandaNovo for practical large-scale de novo peptide sequencing (2023) semanticscholar
- BiATNovo: A Self-Attention based Bidirectional Peptide Sequencing Method (2023) crossref
- PGPointNovo: an efficient neural network-based tool for parallel de novo peptide sequencing (2023) crossref
- Denovo-GCN: De Novo Peptide Sequencing by Graph Convolutional Neural Networks (2023) crossref
- DPST: De Novo Peptide Sequencing with Amino-Acid-Aware Transformers (2022) crossref
- DePS: An improved deep learning model for de novo peptide sequencing (2022) crossref
- De novo mass spectrometry peptide sequencing with a transformer model (2022) crossref
- Computationally instrument-resolution-independent de novo peptide sequencing for high-resolution devices (2021) crossref
- Uncovering Thousands of New Peptides with Sequence-Mask-Search Hybrid De Novo Peptide Sequencing Framework (2019) crossref
- pNovo 3: precise de novo peptide sequencing using a learning-to-rank framework (2019) crossref
- Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning (2019) crossref
- De novo peptide sequencing by deep learning (2017) crossref
Cited by (2)
- AI proteomics: from protein identification to virtual cells (2026) crossref
- Bidirectional Representations Augmented Autoregressive Biological Sequence Generation (2025) semanticscholar