Quick and clean: Cracking sentences encoded in E. coli by LC–MS/MS, de novo sequencing, and dictionary search

peer-reviewed · EuPA Open Proteomics · 2019

peer-reviewed · EuPA Open Proteomics · 2019. Lili Niu et al. In this study, we faced the challenge of deciphering a protein that has been designed and expressed by E…
Date 2019-03-01
Type peer-reviewed
Venue EuPA Open Proteomics
Publisher Elsevier BV
Contribution downstream-application
DOI 10.1016/j.euprot.2019.07.010
Citations (OpenAlex) 8

Abstract

In this study, we faced the challenge of deciphering a protein that has been designed and expressed by E. coli in such a way that the amino acid sequence encodes two concatenated English sentences. The letters ‘O’ and ‘U’ in the sentence are both replaced by ‘K’ in the protein. The sequence cannot be found online and carried to-be-discovered modifications. With limited information in hand, to solve the challenge, we developed a workflow consisting of bottom-up proteomics, de novo sequencing and a bioinformatics pipeline for data processing and searching for frequently appearing words. We assembled a complete first question: “Have you ever wondered what the most fundamental limitations in life are?” and validated the result by sequence database search against a customized FASTA file. We also searched the spectra against an E. coli proteome database and found close to 600 endogenous, co-purified E. coli proteins and contaminants introduced during sample handling, which made the inference of the sentence very challenging. We conclude that E. coli can express English sentences, and that de novo sequencing combined with clever sequence database search strategies is a promising tool for the identification of uncharacterized proteins.

Authors

  1. Lili Niu · Novo Nordisk Foundation, University of Copenhagen
  2. Matthias Mann · European Molecular Biology Laboratory, Max Planck Institute of Biochemistry, Max Planck Institute of Molecular Cell Biology and Genetics, Novo Nordisk Foundation, University of California San Diego, University of Copenhagen

Methods and tools

  • Encoded-sentence protein decoding: Decodes an E. coli-expressed protein whose sequence spells two English sentences, by bottom-up proteomics, de novo sequencing and a dictionary search.

Seen in the charts

Back to the full map

Back to top