Sequence assembly

5 methods · 2016–2026

Sequence assembly: Sequencing a whole protein rather than a peptide. Overlapping peptides are called first, then assembled into full-length sequence, which is how antibodies and protein therapeutics are sequenced without a reference.

Sequencing a whole protein rather than a peptide. Overlapping peptides are called first, then assembled into full-length sequence, which is how antibodies and protein therapeutics are sequenced without a reference.

The earliest of its 5 methods is ALPS (2016); 4 more have followed.

Methods 5
Papers describing them 6
Authors 36
Active 2016-08-26 to 2026-08-13
Kinds post-processor (5)
Acquisition DDA (4)

Methods (5)

Oldest first, by the paper that describes each one.

  • ALPS (2016): Assembles de novo sequenced peptides and their per-residue confidence scores into a de Bruijn graph to reconstruct complete monoclonal antibody heavy and light chains without a template.
  • Stitch (2024): Assembles de novo peptides from Casanovo, PEAKS, pNovo and MaxNovo into full antibody sequences, and corrects the two error classes that assembly alone cannot: mass coincidences, where a different residue combination matches the same mass, and I/L ambiguity.
  • InstaNexus (2025): End-to-end workflow for reference-free sequencing of full-length protein therapeutics. Multi-protease digestion yields overlapping peptides, InstaNovo sequences them de novo and Winnow rescores, then greedy overlap or de Bruijn graph assembly (default k=7, min overlap 3) reconstructs contigs ranked by a composite score over coverage, N50, scaffold count and identity. Validated on nanobodies, monoclonal antibodies and de novo mini-binders.
  • SequenceAssembler (2025): Post-identification tool that assembles full-length protein sequences by unifying peptide-spectrum matching (PSM) and de novo sequencing outputs from Novor Cloud, PEAKS Studio, and PatternLab for Proteomics; one-click GUI and comparable in performance to Stitch.
  • borgonovo (2026): Reference-free protein sequencer: multi-protease digestion tiles a protein with overlapping peptides, and the redundant de novo reads are assembled into a per-residue consensus by substitution-tolerant alignment and per-column voting. Re-decoding each spectrum under a prior from its consensus position raises amino-acid accuracy. Wraps Casanovo by default but is backend-agnostic.

Papers describing them (6)

Authors (36)

Alberto Santos, Alfred Nilsson, Ana Gisele da Costa Neves-Ferreira, Andreas Hougaard Laustsen, Anne Ljungars, Baozhen Shan, Celso Vitor A. Q. Calomeno, Darian Stephan Wolff, Douwe Schulte, Elpida Lytra, Emil Sporre, Erwin M. Schoof, Fredrik Edfors, Hulyana Brum, Jemma Daniel, Jeroen Van Goey, Joost Snijder, Konstantinos Kalogeropoulos, Lei Xin, Lin He, Luis Miguel Muñoz-Gómez, Lukas Käll, M. Ziaur Rahman, Maike Wennekers Nielsen, Marco Reverenna, Marie V. Lukassen, Marlon D. M. Santos, Michel Batista, Ming Li, Ngoc Hieu Tran, Pasquale D. Colaianni, Paulo C. Carvalho, Richard Hemmi Valente, Rodrigo S. C. Brant, Suthimon Thumtecho, Timothy P. Jenkins

Seen in the charts

Back to the full map

Back to top