Sequence similarity‐driven proteomics in organisms with unknown genomes by LC‐MS/MS and automated de novo sequencing
peer-reviewed · PROTEOMICS · 2007
| Date | 2007-07-01 |
| Type | peer-reviewed |
| Venue | PROTEOMICS |
| Publisher | Wiley |
| Contribution | algorithm |
| DOI | 10.1002/pmic.200700003 |
| Citations (OpenAlex) | 99 |
Abstract
LC-MS/MS analysis on a linear ion trap LTQ mass spectrometer, combined with data processing, stringent, and sequence-similarity database searching tools, was employed in a layered manner to identify proteins in organisms with unsequenced genomes. Highly specific stringent searches (MASCOT) were applied as a first layer screen to identify either known (i.e. present in a database) proteins, or unknown proteins sharing identical peptides with related database sequences. Once the confidently matched spectra were removed, the remainder was filtered against a nonannotated library of background spectra that cleaned up the dataset from spectra of common protein and chemical contaminants. The rectified spectral dataset was further subjected to rapid batch de novo interpretation by PepNovo software, followed by the MS BLAST sequence-similarity search that used multiple redundant and partially accurate candidate peptide sequences. Importantly, a single dataset was acquired at the uncompromised sensitivity with no need of manual selection of MS/MS spectra for subsequent de novo interpretation. This approach enabled a completely automated identification of novel proteins that were, otherwise, missed by conventional database searches.
Methods and tools
- Layered sequence-similarity proteomics: Identifies proteins in organisms with unsequenced genomes by layering the search: a stringent database search first, then automated de novo sequencing of what it leaves behind, then sequence-similarity searching of those de novo reads against related species.
Cites (11)
- Performance Evaluation of Existing De Novo Sequencing Algorithms (2006) crossref
- Rapid Validation of Protein Identifications with the Borderline Statistical Confidence via De Novo Sequencing and MS BLAST Searches (2006) crossref
- PepNovo: de novo peptide sequencing via probabilistic network modeling (2005) crossref
- High-Throughput Identification of Proteins and Unanticipated Sequence Modifications Using a Mass-Based Alignment Algorithm for MS/MS de Novo Sequencing Results (2004) crossref
- GutenTag: High-Throughput Sequence Tagging via an Empirically Derived Fragmentation Model (2003) crossref
- Peptide and protein de novo sequencing by mass spectrometry (2003) crossref
- MultiTag: Multiple Error-Tolerant Sequence Tag Search for the Sequence-Similarity Identification of Proteins by Mass Spectrometry (2003) crossref
- Implementation and Uses of Automated de Novo Peptide Sequencing by Tandem Mass Spectrometry (2001) crossref
- Charting the Proteomes of Organisms with Unsequenced Genomes by MALDI-Quadrupole Time-of-Flight Mass Spectrometry and BLAST Homology Searching (2001) crossref
- Sequence database searches via de novo peptide sequencing by tandem mass spectrometry (1997) crossref
- Error-Tolerant Identification of Peptides in Sequence Databases by Peptide Sequence Tags (1994) crossref
Cited by (7)
- Uncovering Hidden Members and Functions of the Soil Microbiome Using De Novo Metaproteomics (2022) semanticscholar
- Venomics and antivenomics of Indian spectacled cobra (Naja naja) from the Western Ghats (2022) crossref
- Mass spectrometry-assisted venom profiling of Hypnale hypnale found in the Western Ghats of India incorporating de novo sequencing approaches (2018) crossref
- Lessons in de novo peptide sequencing by tandem mass spectrometry (2015) both
- De Novo Sequencing and Homology Searching (2012) both
- Algorithm Development of de novo Peptide Sequencing Via Tandem Mass Spectrometry (2010) crossref
- A high-throughput de novo sequencing approach for shotgun proteomics using high-resolution tandem mass spectrometry (2010) both