De novo peptide sequencing with InstaNovo: Accurate, database-free peptide identification for large scale proteomics experiments

preprint · bioRxiv · 2023

preprint · bioRxiv · 2023. Kevin Eloff et al. Bottom-up mass spectrometry-based proteomics is challenged by the task of identifying the peptide that…
Date 2023-08-30
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution algorithm
DOI 10.1101/2023.08.30.555055
Citations (OpenAlex) 30

Peer-reviewed version: InstaNovo enables diffusion-powered de novo peptide sequencing in large-scale proteomics experiments (2025-04-01, Nature Machine Intelligence)

Abstract

Bottom-up mass spectrometry-based proteomics is challenged by the task of identifying the peptide that generates a tandem mass spectrum. Traditional methods that rely on known peptide sequence databases are limited and may not be applicable in certain contexts. De novo peptide sequencing, which assigns peptide sequences to the spectra without prior information, is valuable for various biological applications; yet, due to a lack of accuracy, it remains challenging to apply this approach in many situations. Here, we introduce InstaNovo, a transformer neural network with the ability to translate fragment ion peaks into the sequence of amino acids that make up the studied peptide(s). The model was trained on 28 million labelled spectra matched to 742k human peptides from the ProteomeTools project. We demonstrate that InstaNovo outperforms current state-of-the-art methods on benchmark datasets and showcase its utility in several applications. Building upon human intuition, we also introduce InstaNovo+, a multinomial diffusion model that further improves performance by iterative refinement of predicted sequences. Using these models, we could de novo sequence antibody-based therapeutics with unprecedented coverage, discover novel peptides, and detect unreported organisms in different datasets, thereby expanding the scope and detection rate of proteomics searches. Finally, we could experimentally validate tryptic and non-tryptic peptides with targeted proteomics, demonstrating the fidelity of our predictions. Our models unlock a plethora of opportunities across different scientific domains, such as direct protein sequencing, immunopeptidomics, and exploration of the dark proteome.

Authors

  1. Kevin Eloff · InstaDeep Ltd
  2. Konstantinos Kalogeropoulos · Delft University of Technology, Kavli Institute of Nanoscience, Technical University of Denmark
  3. Oliver Morell · Technical University of Denmark
  4. Amandla Mabona · InstaDeep Ltd
  5. Jakob Berg Jespersen · Technical University of Denmark
  6. Wesley Williams · InstaDeep Ltd
  7. Sam P. B. van Beljouw · Delft University of Technology, Kavli Institute of Nanoscience
  8. Marcin J. Skwark · InstaDeep Ltd
  9. Andreas Hougaard Laustsen · Technical University of Denmark
  10. Stan J. J. Brouns · Delft University of Technology, Kavli Institute of Nanoscience
  11. Anne Ljungars · Technical University of Denmark
  12. Erwin M. Schoof · Technical University of Denmark
  13. Jeroen Van Goey · InstaDeep Ltd
  14. Ulrich auf dem Keller · Technical University of Denmark
  15. Karim Beguir · InstaDeep Ltd
  16. Nicolas Lopez Carranza · InstaDeep Ltd
  17. Timothy P. Jenkins · Technical University of Denmark

Methods and tools

Cites (15)

Cited by (17)

Seen in the charts

Back to the full map

Back to top