Awesome De Novo Peptide Sequencing

A comprehensive, interactive map of the field. Algorithms, post-processors, downstream applications and adjacent tools, deep-learning and classical alike.

A comprehensive, interactive map of the de novo peptide sequencing field: algorithms, post-processors, downstream applications, and adjacent tools, deep-learning and classical alike.

Scope. A comprehensive map of de novo peptide sequencing covering core algorithms, post-processors (re-rankers / FDR / refinement), downstream applications (immunopeptidomics, metaproteomics, cyclopeptides), adjacent tools (database-search hybrids, glycopeptide pipelines), reviews / surveys, and benchmarks. Both deep-learning and classical methods are tracked; the filters below let you slice by approach, acquisition mode (DDA / DIA), and paper kind. Want a paper added? See Contributing.

The wave

The architectures

De novo sequencing has cycled through several methodological families: first hand-engineered dynamic programming and learning-to-rank, then a long stretch of CNN+RNN models, then transformers, GNNs, NAR variants, and most recently diffusion. Use the filters to focus on one slice of the field; hover a dot to read the method’s signature contribution, or click a family name on the left for every method in it.

The long view

How they score

Everything above is what the methods ARE. This is how they do, on the field’s two public benchmarks. Neither is a leaderboard someone can game by reporting their own numbers: both run the tools themselves and publish the results as data.

denovo_benchmarks runs every tool in its own container over 84 datasets, against the same ground-truth PSMs and the same metric code, and asks how a method holds up across instruments, organisms, digests and modifications. That is the first four charts here.

ProteoBench takes one dataset, the published nine-species benchmark, and accepts submitted runs with their parameters recorded. That is the last two. It answers a different question: with these settings, on this data, where does the tool land.

Rank across every dataset

Which of those differences are real

Precision against coverage

Every tool on every dataset

The table

The other benchmark: ProteoBench

The data underneath

Every number on this page was produced by running software over spectra, and the spectra have versions. “Trained on nine-species” names four different objects, and a third of the papers that say it do not say which. These two charts are about that.

An amber segment is not sloppiness, usually. A paper that cites the nine per-species PRIDE submissions has told you exactly where its spectra came from. What it has not told you is which curated re-release it ran, and since the revised benchmark removed cross-species peptide redundancy that the original leaked, the two are not interchangeable for measuring generalisation. The catalog records this as a null version rather than guessing.

Where the nine-species benchmark comes from

The left-hand ribbons carry no count, and are drawn pale for that reason. The nine source submissions are unrelated proteomics studies, of honeybees and tomatoes and bottlenose dolphins, whose spectra were re-curated into one deposit; the catalog stores no per-species spectrum count, and sizing those ribbons would have invented one. The right-hand ribbons are weighted by papers. A version with no papers still gets a visible box, because it exists whether or not anyone cited it.

Deposits that travel together

Six of these have no shared provenance at all. They are unrelated studies, of dolphin tissue and bronchoalveolar lavage fluid and diabetic beta-cells, that nobody ever merged into one object. What they share is that the same papers keep reaching for all of them at once, which makes them a benchmark suite in practice and nothing in name. The catalog keeps them as six deposits for that reason, and this chart is where the convention becomes visible.

Every ribbon here weighs the same, deliberately. A paper either used a deposit or it did not, and unlike the nine-species flow above there is no spectrum count to scale by, so node height is simply degree. The teal block is computed from an identical citing-paper signature rather than from pairwise similarity, which is what makes it a claim worth printing: these deposits are not used by overlapping sets of papers, they are used by the same set. The lighter shade is a deposit matching that signature plus one extra paper.

The dataset network

Every dataset

Application areas

De novo peptide sequencing gets picked up by a handful of distinct scientific communities, each with its own workflow conventions. Each dot below is one application-focused method or workflow, placed at its first publication date and stacked into a lane by sub-domain.

Application → sequencer flow

Which sequencing tools does each application community actually reach for? Two signals can drive the Sankey and they’re different strengths:

  • Uses: the paper has a curated publication_algorithm link to the tool — usually because the methods section names it explicitly. These are the links whose role is uses, plus the rare case of an application paper that also introduces a tool.
  • Cites: the paper’s Crossref reference list contains a publication describing the tool — a weaker signal (the tool might just be mentioned as prior art). Rebuilt monthly by build_citations.py.

Use the toggle to switch between them, or view the union of both. An edge from metaproteomics to PEAKS means at least one metaproteomics-tagged paper satisfies the selected signal. Edge thickness is the paper count.

Where the work happens

Authors and the organizations behind them span countries. Pan and zoom the map. Zoomed out, nearby work collapses to country-level pies; zoom in past about 2× and those aggregates split into city-level pies. Circle area scales with the selected impact metric; pie colours group organizations by type. Organization type is inferred from institution names and should be read as a practical display category.

Top institutions

Who’s driving it

The chart below shows the twenty most-published authors.

The collaboration network

The network below shows how authors with ≥ 3 papers are connected through co-authorship; drag a node to reshape the layout, or hover to highlight a neighborhood.

Edge thickness is Newman fractional collaboration strength: a pair sharing an n-author paper earns 1/(n−1), summed over every paper they share. So an intimate two-author collaboration scores a full 1.0 while each pair in a 53-author consortium scores ~0.02. Without this correction one community benchmark paper alone contributes a 33-node clique that swamps the whole graph. Use min strength to peel away weak ties and expose the dense research groups, and max authors/paper to drop mega-author papers outright.

Models and the authors behind them

This second view rewires the same network as a bipartite graph: every prolific author (≥ 3 papers) is linked to the models they helped publish. Algorithm nodes are diamonds colored by their architecture family, so you can see which research groups own which slice of the architectural landscape.

How the field cites itself

A chronological citation arc diagram. Papers are placed left-to-right by publication date and stratified vertically by kind; within each row, the most-cited papers float to the top. Each arc connects a citing paper (right end) to a paper it cites (left end), curving upward above the row. Hover any paper to highlight the citations into it (red) and out of it (blue), and dim everything else.

Edges resolved from Crossref (by DOI) and Semantic Scholar (by DOI or title-search fallback), matched back to publications via DOI-exact (and, for refs without a DOI, fuzzy-title with token-set ratio ≥ 92). Only intra-catalog citations are drawn; references to papers outside the catalog are filtered out. Every arrow runs citing → cited, so the arrowhead always lands on the older paper.

Academic impact by citation count

Publication-level global citation counts from OpenAlex cited_by_count. Counts are matched to catalog publications by DOI first, then by high-confidence title search when DOI lookup is unavailable.

Code activity

Open-source uptake snapshot from the GitHub API: stars, fork count, open / closed issues + PRs, the most recent push, and the latest released tag (falling back to the latest plain tag if the project doesn’t formally release). Pulled offline by build_repo_metrics.py against algorithm_repository.url. Non-GitHub repos (PyPI, project home pages, anonymised review repos) aren’t counted here.

Plot popularity (★, log scale) against staleness (months since last push). Vibrant repos cluster on the right at higher star counts; abandoned-but-historically-popular projects drift toward the left. Dot size scales with open issues + open PRs. Hover any dot for the repo name, url, and counts.

Where it appears

Most papers in this space appear first on bioRxiv or arXiv. Toggle preprints vs. peer-reviewed to see how the venue distribution shifts.

Venue citedness (open-data analog of the Impact Factor)

Two-year mean citedness from OpenAlex (summary_stats.2yr_mean_citedness). Methodologically equivalent to the Clarivate Impact Factor formula (mean citations in year t to articles published in years t-1 and t-2), but computed over OpenAlex’s open Crossref-aggregated citation graph rather than the paywalled Web of Science one. Conferences and preprint servers are omitted (their non-rolling publication schedule makes the metric misleading). Built offline via build_journal_metrics.py; refresh annually.

Publication lifecycle

How a method goes from arXiv / bioRxiv preprint to a peer-reviewed publication. Each pairing is recorded explicitly rather than inferred from titles or dates, taken from bioRxiv’s own record of where a preprint was published, Crossref’s preprint relations, or hand-checked author overlap. Peer-reviewed, ML-conference and thesis publications all count as “post-preprint”. The Status column tells you whether each row is paired (lifecycle complete), preprint-only (still in flight), or peer-reviewed-only (published without a preprint we have on file).

Browse all papers

Recently added

Every paper

Browse all authors

Aggregated from the currently-filtered set of papers. Searching is case-insensitive across every column (name, affiliation, country, methods).

Contributing

Easiest path: open a GitHub issue with a link to the paper (DOI / arXiv / bioRxiv / OpenReview / …) and I’ll wire it into the database. Corrections are equally welcome: wrong author lists, missing affiliations, mis-classified kind / DL / acquisition, broken hyperlinks, anything that looks off.

Advanced: edit the database directly

The site is generated from denovo.db (SQLite, the source of truth). If you’re comfortable with SQL:

  1. Edit denovo.db with any SQLite tool (sqlite3 CLI, DB Browser for SQLite, DataGrip, …). A new paper typically needs rows in publication, publication_author, and publication_algorithm (set role to uses for a tool the paper runs rather than introduces, which keeps it out of that tool’s authors and papers); a new model also needs a row in algorithm (set kind, is_deep_learning, acquisition_mode). Affiliations cascade through country → city → affiliation and link to authors via author_affiliation.

  2. Regenerate the human-readable SQL dump so the diff is reviewable:

    sqlite3 denovo.db .dump > denovo.sql
  3. Open a PR with both denovo.db and denovo.sql. The GitHub Action rebuilds the site and publishes to gh-pages on merge; typically live within ~3 minutes.

Cite this catalog

If you use this catalog, please cite it as:

Van Goey, J. Awesome De Novo Peptide Sequencing. Zenodo. https://doi.org/10.5281/zenodo.20825737

Machine-readable metadata is in CITATION.cff; GitHub’s “Cite this repository” button exports BibTeX / APA.


This page is a comprehensive map of de novo peptide sequencing covering algorithms, post-processors, downstream applications and adjacent tools, deep-learning and classical alike. Source data and code: GitHub, rebuilt automatically on every push to main.

Back to top