Accurate and ultra-fast de novo HLA-I immunopeptide sequencing with FoxNovo
preprint · LangTaoSha (LTS) Preprint · 2026
| Date | 2026-08-03 |
| Type | preprint |
| Venue | LangTaoSha (LTS) Preprint |
| Publisher | LangTaoSha Preprint Server |
| Contribution | algorithm |
| DOI | 10.65215/LTSpreprints.2026.08.02.000299 |
| Citations (OpenAlex) | 0 |
Abstract
We present FoxNovo, a hybrid deep learning-combinatorial framework for de novo sequencing of immunopeptides trained on a large-scale HLA-I immunopeptidomics dataset assembled and reprocessed from public mass spectrometry (MS) repositories. This integration achieved >90% peptide accuracy on the reported benchmarks while enabling repository-scale analysis at ~2,800 spectra per second—more than 100-fold faster than the evaluated beam-search baseline under the reported benchmark conditions. To mimic the heterogeneous spectral quality encountered in experimental MS analyses, we constructed controlled peak-removal stress tests, in which FoxNovo retained higher accuracy than the evaluated methods at different simulation levels. We subsequently re-analyzed 168 million spectra from all collected 4,423 MS raw files in only 18 hours on a single GPU, equivalent to ~245 raw files per hour. This repository-scale application yielded score-filtered canonical and putative ncORF-mapped peptide predictions and recovered 41 of 42 non-canonical HLA-I peptides previously validated by targeted MS. FoxNovo demonstrates the potential of integrating AI with combinatorial decoding for scalable immunopeptidomics. The source code is available at https://github.com/fennomix/fennomix.novo.
Methods and tools
- FoxNovo: Non-autoregressive de novo sequencer specialised for HLA-I immunopeptides. Uses dual-token m/z encoding (integer + decimal vocabularies, ~4k tokens instead of ~3M fine-grained bins) in the spectrum encoder, then dynamic-programming top-K decoding over the NAR probability matrix to enforce exact precursor-mass constraints. Reaches >90% peptide accuracy at ~2,800 spectra/s, over 100x faster than the beam-search baseline, and was used to re-analyse 168M spectra from 4,423 raw files in 18 h on one GPU, recovering 41 of 42 targeted-MS-validated non-canonical peptides.