NovoRank: Machine Learning Based Post-processing for Performance Improvement in De Novo Peptide Sequencing
thesis · 2022
| Date | 2022-08-01 |
| Type | thesis |
| Publisher | MSc thesis |
| Contribution | post-processor |
| Supervisor | Eunok Paek |
| Link | https://repository.hanyang.ac.kr/handle/20.500.11754/174145 |
Abstract
To identify peptides in mass spectrometry-based proteomics, tandem mass (MS/MS) spectra are analyzed using database search or de novo sequencing tools. In contrast to database search approaches, de novo sequencing directly deduces peptide sequences from MS/MS spectra without any reference to sequence databases. De novo sequencing method often generates incorrect peptide identifications due to its practically unlimited search space and its peptide identification performance does not reach that of database search methods. Instead, de novo sequencing has the advantage of finding novel peptides that are not a part of the sequence database, thus is an essential method for discovering peptides of as yet unknown, biologically important functions. Here, we propose a machine learning based post-processer for de novo sequencing tools, named NovoRank, that can improve the performance of de novo sequencing and is applicable with any de novo peptide sequencing tools. NovoRank uses DBSCAN, a well-known density-based clustering algorithm, and adopts deep learning techniques so that candidate peptide reordering can give a better top-ranked sequence. Given a large-scale synthetic peptide dataset (ProteomeTools), NovoRank increased the peptide recall by 8.63~12.66% when applied with de novo sequencing results from three different software tools.
Methods and tools
- NovoRank: Spectral clustering refinement