Evaluating de novo sequencing in proteomics: already an accurate alternative to database-driven peptide identification?

peer-reviewed · Briefings in Bioinformatics · 2017

peer-reviewed · Briefings in Bioinformatics · 2017. Thilo Muth et al. While peptide identifications in mass spectrometry (MS)-based shotgun proteomics are mostly obtained using…
Date 2017-03-21
Type peer-reviewed
Venue Briefings in Bioinformatics
Publisher Oxford University Press
Contribution review
DOI 10.1093/bib/bbx033
Citations (OpenAlex) 129
Venue 2-year citedness 6.23

Abstract

While peptide identifications in mass spectrometry (MS)-based shotgun proteomics are mostly obtained using database search methods, high-resolution spectrum data from modern MS instruments nowadays offer the prospect of improving the performance of computational de novo peptide sequencing. The major benefit of de novo sequencing is that it does not require a reference database to deduce full-length or partial tag-based peptide sequences directly from experimental tandem mass spectrometry spectra. Although various algorithms have been developed for automated de novo sequencing, the prediction accuracy of proposed solutions has been rarely evaluated in independent benchmarking studies. The main objective of this work is to provide a detailed evaluation on the performance of de novo sequencing algorithms on high-resolution data. For this purpose, we processed four experimental data sets acquired from different instrument types from collision-induced dissociation and higher energy collisional dissociation (HCD) fragmentation mode using the software packages Novor, PEAKS and PepNovo. Moreover, the accuracy of these algorithms is also tested on ground truth data based on simulated spectra generated from peak intensity prediction software. We found that Novor shows the overall best performance compared with PEAKS and PepNovo with respect to the accuracy of correct full peptide, tag-based and single-residue predictions. In addition, the same tool outpaced the commercial competitor PEAKS in terms of running time speedup by factors of around 12-17. Despite around 35% prediction accuracy for complete peptide sequences on HCD data sets, taken as a whole, the evaluated algorithms perform moderately on experimental data but show a significantly better performance on simulated data (up to 84% accuracy). Further, we describe the most frequently occurring de novo sequencing errors and evaluate the influence of missing fragment ion peaks and spectral noise on the accuracy. Finally, we discuss the potential of de novo sequencing for now becoming more widely used in the field.

Authors

  1. Thilo Muth · Federal Institute for Materials Research and Testing (BAM), Max Planck Institute for Dynamics of Complex Technical Systems, Robert Koch Institute
  2. Bernhard Y. Renard · Hasso Plattner Institute, Robert Koch Institute

Methods and tools

Cites (37)

Cited by (28)

Seen in the charts

Back to the full map

Back to top