MassNet: billion-scale AI-friendly mass spectral corpus enables robust de novo peptide sequencing

preprint · bioRxiv · 2025

preprint · bioRxiv · 2025. Jun A et al. Breakthroughs in artificial intelligence (AI) for natural language processing and computer vision have been…
Date 2025-06-20
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution algorithm
DOI 10.1101/2025.06.20.660691
Citations (OpenAlex) 2

Abstract

Breakthroughs in artificial intelligence (AI) for natural language processing and computer vision have been largely driven by high-quality, large-scale datasets such as OpenWebText and ImageNet. Inspired by this, we present MassNet, a foundational resource for proteomics designed to accelerate deep learning applications. MassNet is the largest known corpus of data-dependent acquisition (DDA) mass spectrometry (MS) data, derived from ~30 TB of raw files and comprising 1.54 billion MS/MS spectra, resulting in 558 million peptide-spectrum matches (PSMs) across 35 species, including animals, plants, and microbes. Within the human subset, MassNet includes more than 1.7 million precursors and 19,966 proteins, covering 98% of annotated human proteins. To enable efficient AI training, we developed the Mass Spectrometry Data Tensor (MSDT), a structured format based on Parquet that enables standardized, high-performance batch access and seamless integration with GPU and TPU platforms for distributed training. We further extended MassNet to support de novo peptide sequencing, which infers peptide sequences directly from MS/MS spectra without reference databases, and is critical for discovering novel proteins, characterizing non-model organisms, and identifying post-translational modifications (PTMs). We introduce XuanjiNovo, a non-autoregressive Transformer model that leverages a curriculum learning strategy to enhance training stability. By dynamically adjusting learning difficulty based on model performance, XuanjiNovo achieves smooth convergence on complex, multi-distributional data without manual hyperparameter tuning. Trained on 100 million PSMs from the MassNet, it consistently outperforms state-of-the-art methods across diverse benchmarking tasks. Peptide recall exceeds 0.8 on the Bacteroides thetaiotaomicron and Zea mays datasets. On human data acquired using the Orbitrap Astral platform, XuanjiNovo achieves achieves 38.8% to 144.3% improvement over existing models. MassNet represents the first large-scale, standardized foundational dataset in proteomics, marking a critical milestone in the integration of artificial intelligence into proteomics research.

Authors

  1. Jun A · Westlake University
  2. Xiang Zhang (Shanghai AI Lab) · Fudan University, Shanghai Artificial Intelligence Laboratory, University of British Columbia
  3. Xiaofan Zhang · Westlake University
  4. Jiaqi Wei · Shanghai Artificial Intelligence Laboratory, Zhejiang University
  5. Te Zhang · Westlake University
  6. Yamin Deng · Westlake University
  7. Pu Liu · Westlake Omics (Hangzhou) Biotechnology Co., Ltd.
  8. Zongxiang Nie · Westlake University
  9. Yi Chen · Westlake University
  10. Nanqing Dong · Shanghai Artificial Intelligence Laboratory
  11. Zhiqiang Gao · Shanghai Artificial Intelligence Laboratory
  12. Siqi Sun · Fudan University, Shanghai Artificial Intelligence Laboratory
  13. Tiannan Guo · Westlake Institute for Advanced Study, Westlake University

Methods and tools

Cites (26)

Cited by (2)

Seen in the charts

Back to the full map

Back to top