Reference-free protein sequencing by consensus assembly of redundant de novo peptide reads
preprint · bioRxiv · 2026
| Date | 2026-08-13 |
| Type | preprint |
| Venue | bioRxiv |
| Publisher | Cold Spring Harbor Laboratory |
| Contribution | post-processor |
| DOI | 10.64898/2026.08.13.744110 |
| Citations (OpenAlex) | 0 |
Abstract
Reading a protein’s sequence from tandem mass spectra without a reference is limited by single-spectrum accuracy, most acutely across the hypervariable complementarity-determining regions of antibodies. Broadly specific proteases tile a protein with long, overlapping peptides, so every residue is covered by many independent de novo reads. borgonovo assembles their per-step probability profiles into a reference-free per-residue consensus, seeding templates from mass-closure-consistent reads and recruiting the rest by substitution-tolerant alignment and per-column voting. Re-decoding each spectrum with a prior from its consensus position lifts amino acid accuracy on placed spectra from 0.80 to 0.87. On the therapeutic antibody trastuzumab, nine proteases cover its heavy and light chains completely at 0.88 fixed-window identity and 0.93 on the pruned assembly once local indels are accommodated. Applied unchanged to five secretome proteins and trastuzumab with three proteases, it reaches 0.87 mean fixed-window identity over 82% coverage. borgonovo is open source and works with most de novo sequencers, so redundant digestion turns any of them into a protein sequencer where no reference exists.
Methods and tools
- borgonovo: Reference-free protein sequencer: multi-protease digestion tiles a protein with overlapping peptides, and the redundant de novo reads are assembled into a per-residue consensus by substitution-tolerant alignment and per-column voting. Re-decoding each spectrum under a prior from its consensus position raises amino-acid accuracy. Wraps Casanovo by default but is backend-agnostic.
Methods it uses
- Casanovo: First Transformer
Data used
Cites (21)
- Deep coverage and extended sequence reads obtained with a single archaeal protease expedite de novo protein sequencing by mass spectrometry (2026) crossref
- AbNovoBench: a resource and benchmarking platform for monoclonal antibody de novo sequencing (2026) both
- Pairwise Attention: Leveraging Mass Differences to Enhance De Novo Sequencing of Mass Spectra (2025) both
- InstaNovo enables diffusion-powered de novo peptide sequencing in large-scale proteomics experiments (2025) both
- A Handle on Mass Coincidence Errors in De Novo Sequencing of Antibodies by Bottom-up Proteomics (2024) both
- Sequence-to-sequence translation from mass spectra to peptides with a transformer model (2024) both
- ContraNovo: A Contrastive Learning Approach to Enhance De Novo Peptide Sequencing (2024) crossref
- Comprehensive evaluation of peptide de novo sequencing tools for monoclonal antibody assembly (2023) both
- De novo mass spectrometry peptide sequencing with a transformer model (2022) crossref
- De novo mass spectrometry peptide sequencing with a transformer model (2022) semanticscholar
- Highly Robust de Novo Full-Length Protein Sequencing (2021) both
- De novo peptide sequencing by deep learning (2017) both
- De novo Peptide Sequencing (2016) both
- Complete De Novo Assembly of Monoclonal Antibody Sequences (2016) both
- Shotgun Protein Sequencing with Meta-contig Assembly (2012) both
- PEAKS DB: De Novo Sequencing Assisted Database Search for Sensitive and Accurate Peptide Identification (2012) both
- SPIDER: software for protein identification from sequence tags with de novo sequencing error (2005) crossref
- Shotgun Protein Sequencing by Tandem Mass Spectra Assembly (2004) both
- High-Throughput Identification of Proteins and Unanticipated Sequence Modifications Using a Mass-Based Alignment Algorithm for MS/MS de Novo Sequencing Results (2004) both
- PEAKS: powerful software for peptide de novo sequencing by tandem mass spectrometry (2003) both
- Sequence database searches via de novo peptide sequencing by tandem mass spectrometry (1997) both