Awesome De Novo Peptide Sequencing
A comprehensive, interactive map of the field. Algorithms, post-processors, downstream applications and adjacent tools, deep-learning and classical alike.
plot_tip_style = ({
fontSize: 13,
lineHeight: 1.35,
textPadding: 10,
// Plot truncates each tip line at lineWidth (20em by default, roughly 40
// characters) and the ellipsis eats the END of the line, which is the value.
// Several x labels here are already 29-37 characters, so the default left no
// room for the number. None of these tips carry free text like a paper title,
// so a wider limit costs nothing.
lineWidth: 34,
fill: "white",
stroke: "#d0d7de"
})
pubs_t = transpose(pubs).map(r => ({ ...r, year: +r.year, date: r.date ? new Date(r.date) : null }))
// Shared country colour scale. Both the Top-institutions bars and the
// most-published-authors bars use this, so a country keeps the same colour in
// both. It has to be an explicit scale over the FULL country list: Plot's
// default ordinal colours are assigned per-chart from whatever domain that
// chart happens to contain, so two charts with different country sets would
// otherwise paint the same country differently.
country_color = {
const all = Array.from(
new Set(transpose(pub_authorship).map(r => r.country).filter(Boolean))
).sort()
const palette = d3.schemeTableau10
.concat(d3.schemeSet3, d3.schemeSet2, d3.schemePaired)
return d3.scaleOrdinal()
.domain(all)
.range(all.map((_, i) => palette[i % palette.length]))
.unknown("#c9d1d9")
}
// Shared algorithm-family colour scale, for the same reason country_color
// exists: Plot's default ordinal colours are positional, so a family's colour
// silently changes whenever the set of families present changes. That bit us
// once already: adding one 'Tooling' category to the code-activity scatter
// reshuffled every colour and turned Transformer (AR) brown. An explicit scale
// over the full family list keeps a family the same colour wherever it appears.
// The first 13 hexes are the architectures swim-lane's canonical band colours,
// so the scatter and the swim-lane now agree.
family_color_scale = {
// Heuristic, Graph / DP and Learning-to-rank used to be pinned to three near
// identical greys (#8a96a0, #6c757d, #5f6b7a). They are not here any more, so
// they get a generated colour like any other family. Heuristic and Graph / DP
// are the two oldest and therefore the two biggest lanes at the TOP of the
// swim lane, 12.8 deltaE apart: the chart opened on two grey slabs, which is
// the main reason it read as colourless.
const canonical = {
"HMM": "#4d6a8c",
"Decision tree": "#7b6f43",
"Random Forest": "#a5673f",
"CNN + RNN": "#4C72B0",
"Transformer (AR)": "#DD8452",
"GNN": "#937860",
"CNN": "#8172B3",
"Transformer (NAR)": "#55A868",
"Diffusion": "#C44E52"
}
// Flow is not pinned either. Its notebook colour, #937DC2, sat 5.8 deltaE
// from CNN's #8172B3, which was the worst pair on the whole page and a real
// ambiguity in a 71-entry legend.
// The other 36 families get a GENERATED colour rather than a slot in a d3
// scheme. No scheme is long enough: cycling `i % palette.length` repeats a
// colour, which is exactly the ambiguity this scale exists to prevent, and the
// schemes carry near-white pastels (#ffffb3, #ffed6f) that vanish as a small
// dot. Colours are drawn instead from a 24-hue grid at FIXED CIE lightness and
// chroma, `d3.lch(46, 58, h)`, which keeps every one equally dark: these have
// to carry a 10 px bold label as well as a dot. LCh and not HSL because HSL at
// one lightness renders yellow far paler than blue.
//
// Assignment walks families in order of FIRST APPEARANCE and picks, for each,
// the grid hue furthest in CIE Lab from its neighbours in that order. Both
// swim lanes and "The long view" lay families out chronologically, so a
// family's neighbours in this walk are the lanes it will actually sit next
// to, which is the only adjacency the eye compares. Alphabetical assignment
// measured four adjacent pairs under the deltaE >= 20 bar, worst 11.1; this
// measured none. The order is a property of the data, not of one chart, so
// the scale stays chart-independent.
const lab = hex => {
const [r, g, b] = [1, 3, 5].map(i => parseInt(hex.slice(i, i + 2), 16) / 255)
.map(c => c <= 0.04045 ? c / 12.92 : ((c + 0.055) / 1.055) ** 2.4)
const X = (0.4124 * r + 0.3576 * g + 0.1805 * b) / 0.95047
const Y = 0.2126 * r + 0.7152 * g + 0.0722 * b
const Z = (0.0193 * r + 0.1192 * g + 0.9505 * b) / 1.08883
const f = t => t > 0.008856 ? Math.cbrt(t) : 7.787 * t + 16 / 116
return [116 * f(Y) - 16, 500 * (f(X) - f(Y)), 200 * (f(Y) - f(Z))]
}
const dE = (p, q) => {
const a = lab(p), b = lab(q)
return Math.hypot(a[0] - b[0], a[1] - b[1], a[2] - b[2])
}
const first_pub = d3.rollup(
transpose(algorithms).filter(r => r.family && r.first_pub),
v => d3.min(v, r => r.first_pub),
r => r.family
)
// The author-model graph labels a FAMILY-LESS algorithm by its kind, so
// 'Reviews', 'Adjacent tools (misc)' and 'Application: venomics' are families
// as far as a legend is concerned. They have to be coloured here or they all
// fall through to the single unknown fallback: consolidating that chart's own
// palette into this scale is what exposed it, with 20-odd legend entries
// coming out the same brown. Built from the same columns the graph reads, so
// the two cannot drift; they sort last because they have no first-appearance
// date of their own.
const pseudo = ["Reviews", "Benchmarks", "Meta / catalogs",
"Post-processors (misc)", "Adjacent tools (misc)"]
.concat(Array.from(new Set(
transpose(algorithms)
.filter(r => r.kind === "downstream-application")
.map(r => `Application: ${r.subdomain ?? "misc"}`))).sort())
const ordered = Array.from(new Set(
transpose(algorithms).map(r => r.family).filter(Boolean)
)).sort((a, b) =>
d3.ascending(first_pub.get(a) ?? "9999", first_pub.get(b) ?? "9999") ||
d3.ascending(a, b)).concat(pseudo)
// The candidate palette is 24 hues at two lightnesses and two chromas, 96
// colours, not 24 hues at one of each. With 71 families to colour (49 real
// plus the graph's 22 pseudo-families) a single-hue-ring palette runs out
// after 24 and the picker below starts reusing: the legend came out as
// alternating greens and purples. Lightness 40 and 54 both carry a bold
// label on white, which is the constraint that keeps this grid small.
const grid = d3.cross(d3.range(24), [40, 54], [42, 62])
.map(([h, l, c]) => d3.lch(l, c, h * 15).formatHex())
const map = new Map(Object.entries(canonical))
const taken = Object.values(canonical)
ordered.forEach((f, i) => {
if (map.has(f)) return
// Farthest-point assignment: the dominant term is the distance to EVERY
// colour already assigned, which spreads 71 of them as far apart as the
// palette allows. The second term is the distance to this family's
// neighbours in the chronological walk -- the one before, and the one after
// when it is canonical -- which breaks ties in favour of lanes that will
// sit next to each other. Weighted the other way round, the walk produced
// good lane contrast and a legend full of near-duplicates.
const near = [map.get(ordered[i - 1]), canonical[ordered[i + 1]]].filter(Boolean)
const score = c => d3.min(taken, t => dE(c, t))
+ 0.35 * (near.length ? d3.min(near, n => dE(c, n)) : 0)
const pick = d3.greatest(grid, score)
map.set(f, pick)
taken.push(pick)
})
return f => map.get(f) ?? "#9c755f" // fallback also covers the no-family group
}
top_authors_t = transpose(top_authors)
geo_t = transpose(geo)
institutions_t = transpose(institutions)
pub_authorship_t = transpose(pub_authorship)
coauth_edges_t = transpose(coauth_edges).map(r => ({ ...r, n_authors: +r.n_authors, papers: +r.papers }))
// Newman (2001) fractional collaboration strength. Each pair on an n-author
// paper earns 1/(n-1), so a 2-author paper contributes a full 1.0 to its single
// pair while a 53-author consortium contributes ~0.019 to each of its 1378
// pairs. `max_authors` lets the reader drop mega-author papers entirely.
coauth_strength = (max_authors = Infinity) => {
// Author names contain spaces, so the composite key needs a delimiter that
// cannot occur inside a name (same '<<|>>' convention as the bipartite chart).
const SEP = "<<|>>"
const w = new Map() // "source<<|>>target" -> summed fractional weight
for (const r of coauth_edges_t) {
if (r.n_authors > max_authors) continue
const k = r.source + SEP + r.target
w.set(k, (w.get(k) ?? 0) + r.papers / (r.n_authors - 1))
}
return Array.from(w, ([k, weight]) => {
const [source, target] = k.split(SEP)
return { source, target, weight }
})
}
// All prolific authors, regardless of the network's reactive filters — used by
// the bipartite author↔model chart so its node set stays stable.
coauth_all_authors = new Set(coauth_edges_t.flatMap(e => [e.source, e.target]))
author_affs_t = transpose(author_affs)
algorithms_t = transpose(algorithms).map(r => ({
...r,
first_pub: r.first_pub ? new Date(r.first_pub) : null,
is_dl: r.is_dl == null ? null : !!r.is_dl // SQLite INTEGER 0/1 → JS boolean
}))
// Version-aware view of algorithms: any algorithm whose joined publications
// carry a `version` tag (currently Casanovo v1 / v2 / v5) is expanded into one
// row per version, each with its own first_pub date and a label like
// "Casanovo v2". Drives the architectures timeline so successive releases of
// the same method appear as distinct dots instead of collapsing onto the
// earliest one. All other downstream cells (counters, bipartite network,
// browse table) keep using `algorithms_t` and remain one-row-per-algorithm.
algorithms_versioned_t = {
const pubs_for = new Map()
for (const p of pubs_t) {
for (const m of (p.models ?? "").split(",").map(s => s.trim()).filter(Boolean)) {
if (!pubs_for.has(m)) pubs_for.set(m, [])
pubs_for.get(m).push(p)
}
}
const out = []
for (const r of algorithms_t) {
const its_pubs = pubs_for.get(r.model) ?? []
const versions = Array.from(new Set(its_pubs.map(p => p.version).filter(Boolean))).sort()
if (versions.length === 0) { out.push({ ...r, base_model: r.model }); continue }
for (const v of versions) {
const dates = its_pubs.filter(p => p.version === v).map(p => p.date).filter(Boolean)
const earliest = dates.length ? new Date(Math.min(...dates.map(d => +d))) : r.first_pub
out.push({ ...r, model: `${r.model} ${v}`, version: v, base_model: r.model, first_pub: earliest })
}
const unversioned = its_pubs.filter(p => !p.version)
if (unversioned.length) {
const dates = unversioned.map(p => p.date).filter(Boolean)
const earliest = dates.length ? new Date(Math.min(...dates.map(d => +d))) : r.first_pub
out.push({ ...r, base_model: r.model, first_pub: earliest })
}
}
return out
}
venues_t = transpose(venues)
author_details_t = transpose(author_details)
// --- links into the generated entity pages ------------------------------
// entity_slugs comes from slugs.py, the same module build_pages.py uses, so a
// link here and the page it targets can never disagree.
entity_slugs_t = transpose(entity_slugs)
slug_map = {
const m = new Map()
for (const r of entity_slugs_t) {
if (!m.has(r.type)) m.set(r.type, new Map())
m.get(r.type).set(String(r.key), r.slug)
}
return m
}
// page_href("authors", "Lukas Käll") -> "pages/authors/lukas-kall.html", or
// null when there is no such page. Callers MUST handle null: an author who is
// filtered out, or a venue string with no page, should render as plain text
// rather than a broken link.
page_href = (type, key) => {
if (key == null) return null
const slug = slug_map.get(type)?.get(String(key))
return slug ? `pages/${type}/${slug}.html` : null
}
// Markup helper for Inputs.table cell formatters.
page_link = (type, key, label) => {
const href = page_href(type, key)
const text = label ?? key
return href ? htl.html`<a href=${href}>${text}</a>` : text
}
// Wrap already-rendered d3 nodes in an SVG <a>, the way Plot does internally.
// The selection keeps referencing the ORIGINAL elements, so force-simulation
// tick handlers that set `transform` on them keep working.
//
// Both force graphs use d3.drag(), and a drag that ends on an anchor would
// otherwise navigate. So record the pointer position and swallow the click if
// it moved more than a few pixels.
// A repo can back two algorithms, so `model` reads "A / B"; link the first.
algo_href = m => page_href("algorithms", String(m ?? "").split(" / ")[0].trim())
// Where a family label should go. 26 of 52 families have a page of their own;
// the other 26 hold exactly one method, so their page would have carried that
// method's papers, authors and dates and nothing else. Rather than leave a
// lane label dead, a singleton family links straight to its one method: every
// label in the swim lane and every row of the long-view table is clickable, and
// which of the two kinds of target it has is not something a reader needs to
// know.
family_sole_method = new Map(
d3.rollups(algorithms_t.filter(a => a.family), v => v, a => a.family)
.filter(([, v]) => v.length === 1)
.map(([f, v]) => [f, v[0].model])
)
family_href = f => page_href("families", f)
?? algo_href(family_sole_method.get(f))
svg_link_wrap = (selection, href_of) => {
selection.each(function (d) {
const href = href_of(d)
if (!href) return
const node = this
const parent = node.parentNode
const a = document.createElementNS("http://www.w3.org/2000/svg", "a")
a.setAttributeNS("http://www.w3.org/1999/xlink", "xlink:href", href)
a.setAttribute("href", href)
a.style.cursor = "pointer"
parent.insertBefore(a, node)
a.appendChild(node)
let start = null
a.addEventListener("pointerdown", ev => { start = [ev.clientX, ev.clientY] })
a.addEventListener("click", ev => {
if (!start) return
const moved = Math.hypot(ev.clientX - start[0], ev.clientY - start[1])
if (moved > 4) ev.preventDefault() // it was a drag, not a click
})
})
return selection
}
// Linkify the tick labels of a Plot ordinal axis. Matched by TEXT, never by
// index, so there is no ordering assumption to get wrong: a label that does not
// resolve simply stays plain text.
linkify_tick_labels = (chart, href_by_label) => {
const groups = chart.querySelectorAll('g[aria-label$="axis tick label"]')
for (const g of groups) {
for (const text of g.querySelectorAll("text")) {
const href = href_by_label.get(text.textContent)
if (!href) continue
const a = document.createElementNS("http://www.w3.org/2000/svg", "a")
a.setAttributeNS("http://www.w3.org/1999/xlink", "xlink:href", href)
a.setAttribute("href", href)
a.style.cursor = "pointer"
text.parentNode.insertBefore(a, text)
a.appendChild(text)
}
}
return chart
}
// Several tables show a comma-joined list of names in one cell (authors,
// models). Link each name separately, leaving unmatched names as plain text.
page_links = (type, joined, { sep = ", ", max = Infinity } = {}) => {
const names = String(joined ?? "").split(sep).map(s => s.trim()).filter(Boolean)
const shown = names.slice(0, max)
const parts = []
shown.forEach((n, i) => {
if (i) parts.push(sep)
parts.push(page_link(type, n, n))
})
if (names.length > shown.length) parts.push(` +${names.length - shown.length}`)
return htl.html`${parts}`
}
// preprint_id -> published_id, from the explicit publication_version table.
pub_versions_t = transpose(pub_versions).map(r => ({
...r, preprint_id: +r.preprint_id, published_id: +r.published_id
}))
published_of = new Map(pub_versions_t.map(r => [r.preprint_id, r.published_id]))
preprint_of = new Map(pub_versions_t.map(r => [r.published_id, r.preprint_id]))
citations_t = transpose(citations)
journal_impact_t = transpose(journal_impact)
// Lookup table: journal name → 2yr citedness, for fast joins in OJS cells.
journal_impact_by_name = new Map(journal_impact_t.map(j => [j.journal, j]))
publication_impact_t = transpose(publication_impact)
publication_impact_by_pub = new Map(publication_impact_t.map(r => [r.publication_id, r]))
repo_metrics_t = transpose(repo_metrics).map(r => ({
...r,
last_pushed: r.last_pushed ? new Date(r.last_pushed) : null,
fetched_at: r.fetched_at ? new Date(r.fetched_at) : null,
}))
// Some repositories back multiple catalogued algorithms (e.g. InstaNovo +
// InstaNovo-P share instadeepai/instanovo). For the two Code-activity plots
// we want ONE dot per repo — otherwise the shared bar/point appears twice
// on top of itself and the label collides. Fold rows by URL, joining the
// per-repo model names with ' / ' so the tooltip still shows which
// algorithms live in that repo. The Browse-all table below keeps the
// un-deduplicated repo_metrics_t so per-algorithm columns (family, kind)
// stay intact.
repo_metrics_by_url = {
const grouped = new Map()
for (const r of repo_metrics_t) {
if (!r.url) continue
if (!grouped.has(r.url)) grouped.set(r.url, {...r, models: [r.model]})
else grouped.get(r.url).models.push(r.model)
}
return Array.from(grouped.values()).map(r => ({
...r,
// 'model' overrides the single-value field with a joined label used by
// Plot.text and the tooltip title. 'family' picks the first algorithm's
// family for the dot color (families are usually the same when they
// share a repo; where they differ, the first one seen wins).
model: r.models.join(" / "),
}))
}
// Derived counters. Everything flows from the data, no hardcoded numbers.
n_papers = pubs_t.length
n_models = algorithms_t.length
n_authors = new Set(pubs_t.flatMap(p => (p.authors ?? "").split(", ").filter(Boolean))).size
n_countries = geo_t.length
// Datasets replaced countries in the hero. n_countries stays defined because
// the geography section's own prose still reads it.
n_datasets = all_dataset_rows.length
years_with_pubs = pubs_t.map(p => p.year).filter(y => Number.isFinite(y))
first_year = Math.min(...years_with_pubs)
last_year = Math.max(...years_with_pubs)
// First year a paper using deep learning appears in the catalog. Anchors the
// "wave" prose so the DL inflection point reads from data, not a constant.
dl_algo_names = new Set(algorithms_t.filter(a => a.is_dl === true).map(a => a.model))
first_dl_year = Math.min(...pubs_t
.filter(p => Number.isFinite(p.year) && (p.models ?? "").split(",").map(s => s.trim()).some(m => dl_algo_names.has(m)))
.map(p => p.year)
)
n_preprints = pubs_t.filter(p => p.type === "preprint").length
n_peer_reviewed = pubs_t.filter(p => p.type === "peer-reviewed").length
families = Array.from(new Set(algorithms_t.map(a => a.family).filter(Boolean)))
// Classification breakdowns (kind / DL / acquisition).
n_dl = algorithms_t.filter(a => a.is_dl === true).length
n_non_dl = algorithms_t.filter(a => a.is_dl === false).length
kinds_present = Array.from(new Set(algorithms_t.map(a => a.kind).filter(Boolean))).sort()
acq_modes_present = Array.from(new Set(algorithms_t.map(a => a.acquisition).filter(Boolean))).sort()
kind_counts = kinds_present.map(k => ({ kind: k, n: algorithms_t.filter(a => a.kind === k).length }))
acq_counts = acq_modes_present.map(m => ({ acquisition: m, n: algorithms_t.filter(a => a.acquisition === m).length })){
// Each badge jumps to the section that lets you actually browse that number.
// Methods point at the architectures swim lane rather than a table because
// there is no per-method table on the page; that chart is the only view
// showing every method. Plain <a href> rather than a click handler, so
// keyboard, middle-click and open-in-new-tab all behave normally.
const stat = (n, label, href, title) => html`<a class="hero-stat-link" href="${href}" title="${title}">
<div class="hero-stat"><div class="hero-num">${n}</div><div class="hero-lbl">${label}</div></div>
</a>`
return html`<div class="hero-grid">
${stat(n_papers, "papers", "#every-paper", "Browse all papers")}
${stat(n_models, "methods", "#the-architectures", "See every method on the architectures timeline")}
${stat(n_authors, "authors", "#browse-all-authors", "Browse all authors")}
${stat(n_datasets, "datasets", "#the-data-underneath", "See what the field runs on")}
</div>`
}// Lets the breakdown rows below drive the three global filters.
//
// It references the viewof ELEMENTS, not their values. In OJS that is a
// dependency on the element (stable for the life of the page) rather than on
// the value, so this cell does not re-run every time a filter changes, and
// there is no reactive loop back into the breakdown that reads it.
breakdown_apply = {
const set = (view, value) => {
view.value = value
view.dispatchEvent(new Event("input", { bubbles: true }))
// The filter panel is collapsed by default on narrow screens, where a
// filter set from here would change page-wide state invisibly. Open it so
// the change is legible and reversible.
const wrap = document.querySelector(".sticky-filters")
if (wrap) wrap.setAttribute("aria-expanded", "true")
}
return {
kind: k => set(viewof kind_filter, [k]),
approach: v => set(viewof dl_filter, v),
acq: m => set(viewof acq_filter, [m])
}
}{
const kind_label = {
"algorithm": "Algorithms",
"post-processor": "Post-processors",
"downstream-application": "Downstream apps",
"adjacent": "Adjacent",
"review": "Reviews / surveys",
"benchmark": "Benchmarks",
"meta": "Meta"
}
// Every row sets the matching global filter and lands on the papers table.
// href does the navigation so keyboard and middle-click still work; a click
// listener sets the filter. Only the clicked dimension is touched, so an
// approach or acquisition choice already made is preserved.
//
// Handlers are attached with addEventListener AFTER building the markup,
// not as onclick=${fn} inside the template. Quarto's OJS `html` is the
// observablehq/stdlib one, which builds through innerHTML and cannot take a
// function: it stringifies it, leaving a literal onclick attribute that
// throws "Failed to read the 'onclick' property ... Unexpected token ')'".
// Only the data-* strings survive the template, so the wiring reads them.
//
// Note the counts are of METHODS while the destination lists PAPERS, so the
// number on the row is not the number of rows you land on. They answer
// different questions; the filter is what they share.
const row = (n, label, colour, dim, val, hint) => html`<a
class="breakdown-row breakdown-row-link"
href="#every-paper"
title="${hint}"
data-dim="${dim}" data-val="${val}">
<span class="bk-num">${n}</span>
<span class="bk-bar"><span class="bk-fill" style="width:${100 * n / n_models}%; background:${colour}"></span></span>
<span class="bk-lbl">${label}</span>
</a>`
const root = html`<div class="breakdown-grid">
<div class="breakdown-cell">
<div class="breakdown-title">By kind</div>
${kind_counts.map(k => row(k.n, kind_label[k.kind] ?? k.kind, "#1f6feb", "kind", k.kind,
`Filter the page to ${kind_label[k.kind] ?? k.kind} and jump to the papers table`))}
</div>
<div class="breakdown-cell">
<div class="breakdown-title">By approach</div>
${row(n_dl, "Deep learning", "#1f6feb", "approach", "DL only",
"Filter the page to deep-learning methods and jump to the papers table")}
${row(n_non_dl, "Classical", "#6f42c1", "approach", "Classical only",
"Filter the page to classical methods and jump to the papers table")}
</div>
<div class="breakdown-cell">
<div class="breakdown-title">By acquisition</div>
${acq_counts.map(m => row(m.n, m.acquisition, "#1a7f37", "acq", m.acquisition,
`Filter the page to ${m.acquisition} and jump to the papers table`))}
</div>
</div>`
for (const a of root.querySelectorAll("a.breakdown-row-link")) {
a.addEventListener("click", () => {
const { dim, val } = a.dataset
if (dim === "kind") breakdown_apply.kind(val)
else if (dim === "approach") breakdown_apply.approach(val)
else if (dim === "acq") breakdown_apply.acq(val)
})
}
return root
}md`Since **${first_year}**, **${n_papers}** papers have introduced **${n_models}** methods for *de novo* peptide sequencing, written by **${n_authors}** authors across **${n_countries}** countries. Of those papers, **${n_preprints}** are preprints and **${n_peer_reviewed}** are peer-reviewed. A snapshot of a field where the conversation moves faster than the journals.`Scope. A comprehensive map of de novo peptide sequencing covering core algorithms, post-processors (re-rankers / FDR / refinement), downstream applications (immunopeptidomics, metaproteomics, cyclopeptides), adjacent tools (database-search hybrids, glycopeptide pipelines), reviews / surveys, and benchmarks. Both deep-learning and classical methods are tracked; the filters below let you slice by approach, acquisition mode (DDA / DIA), and paper kind. Want a paper added? See Contributing.
🎚️ Filters · Kind · Approach · Acquisition. Apply across the whole page; pinned to the top while you scroll. ▸
pubs_matches_filter = p => {
if (p.kind && !kind_filter.includes(p.kind)) return false
if (dl_filter === "DL only" && !(p.is_dl === 1 || p.is_dl === true)) return false
if (dl_filter === "Classical only" && !(p.is_dl === 0 || p.is_dl === false)) return false
if (p.acquisition && !acq_filter.includes(p.acquisition)) return false
return true
}
pubs_filtered = pubs_t.filter(pubs_matches_filter)
n_papers_filtered = pubs_filtered.lengthThe wave
md`The earliest paper tracked here appeared in **${first_year}**; the first deep-learning method shows up in **${first_dl_year}**. Activity has accelerated sharply since. **${n_peer_reviewed}** papers have made it through peer review, alongside **${n_preprints}** preprints still in the publication pipeline.`Plot.plot({
marginLeft: 50,
width: 1100,
height: 360,
marginBottom: 48,
// labelOffset + the matching marginBottom put the axis label on its own line
// below the ticks. Plot centres it on the tick baseline for a band scale,
// where it touched the 2005 tick by 3 px.
x: { label: "Year", tickFormat: "d", interval: 1, labelOffset: 40 },
y: { label: `Papers (${n_papers_filtered} shown)`, grid: true },
color: { legend: true, scheme: "blues", domain: ["preprint", "postprint", "peer-reviewed", "ML conference", "thesis", "commentary", "abstract", "presentation"] },
marks: [
Plot.barY(
// Drop kind='meta' rows: the catalog's self-entry (no
// publication_type, no real venue) plus any other meta artefacts
// like commentaries-without-method. They'd otherwise render as
// transparent stack segments and skew the publication-volume story.
pubs_filtered.filter(p => Number.isFinite(p.year) && p.kind !== "meta"),
Plot.groupX(
{ y: "count" },
{ x: "year", fill: "type", tip: plot_tip_style }
)
),
Plot.ruleY([0])
]
})The architectures
De novo sequencing has cycled through several methodological families: first hand-engineered dynamic programming and learning-to-rank, then a long stretch of CNN+RNN models, then transformers, GNNs, NAR variants, and most recently diffusion. Use the filters to focus on one slice of the field; hover a dot to read the method’s signature contribution, or click a family name on the left for every method in it.
innovations_timeline = {
// Colour comes from the SHARED family scale, which every family-coloured chart
// on the page uses, so a family is the same colour in the swim lane, the
// code-activity scatter and the bipartite graph. It used to stop at these 13:
// the other 36 shared one neutral grey, so most of the chart read as a single
// undifferentiated family. See family_color_scale for how the rest are
// generated. Defining a second local generator here would be the bug that
// scale exists to prevent.
const color_of = f => family_color_scale(f)
const family_first = d3.rollup(
algorithms_versioned_t.filter(a => a.family && a.first_pub),
v => d3.min(v, a => a.first_pub),
a => a.family
)
// Lanes stack UPWARD from y = 0, so the array runs latest-first to put the
// earliest family at the top, matching "The long view". Family name breaks a
// tie so the order is deterministic; an undated family sorts as latest.
const LATEST = new Date(8640000000000)
const lane_sort = (a, b) =>
d3.descending(family_first.get(a) ?? LATEST, family_first.get(b) ?? LATEST)
|| d3.descending(a, b)
const all_families = Array.from(new Set(
algorithms_versioned_t.filter(a => a.family).map(a => a.family))).sort(lane_sort)
// Lane geometry is measured in PIXELS, because the thing that must not
// collide is a rendered label. It used to be assigned from a date gap against
// a per-family tier list sized to the BUSIEST family, which handed a 2-unit
// band 38 tier positions ~4 px apart and piled its labels on top of each
// other: 7 label-label and 15 label-dot collisions, measured in the DOM.
// Packing by measured label width removes the class rather than retuning it,
// and it makes the hand-tuned band_heights table unnecessary -- a lane is now
// exactly as tall as the rows its labels need.
const ROW_PX = 43 // 24 px for up to two label lines, 3 gap, 16 dot
const DOT_IN_ROW_PX = 35 // dot centre from the row top, label in the 24 above
const CHAR_PX = 5.7 // mean advance of 10 px bold Source Sans Pro
const LABEL_GAP_PX = 8 // clear air between two labels sharing a row
const PLOT_WIDTH_PX = 1130 // width 1400 minus marginLeft 210 and marginRight 60
// Apply every filter (family, kind, approach, acquisition) up front.
const matches_filters = a => {
if (!a.first_pub || !a.family) return false
if (!family_filter.includes(a.family)) return false
if (!kind_filter.includes(a.kind)) return false
if (dl_filter === "DL only" && a.is_dl !== true) return false
if (dl_filter === "Classical only" && a.is_dl !== false) return false
if (a.acquisition && !acq_filter.includes(a.acquisition)) return false
return true
}
// Only include bands that have at least one surviving model. Every present
// family gets a lane: the lane list used to be intersected with the thirteen
// hand-tuned ones, which dropped the other 36 entirely and made checking any
// one of them produce an empty chart.
const surviving = algorithms_versioned_t.filter(matches_filters)
const present = new Set(surviving.map(a => a.family))
const visible_bands = Array.from(present).sort(lane_sort)
// Wrap the display label onto two lines at a word boundary when it exceeds
// the max width (Plot renders \n as separate tspans). Same helper as the
// application-areas chart below -- labels here can be long too
// ('MALDI-QIT de novo sequencing', 'Robust FL-Sequencing', etc).
const wrap_label = (s, max_chars = 22) => {
if (!s || s.length <= max_chars) return s
const words = s.split(' ')
if (words.length === 1) return s
let line1 = ''
let i = 0
while (i < words.length && (line1 ? line1.length + 1 + words[i].length : words[i].length) <= max_chars) {
line1 = line1 ? `${line1} ${words[i]}` : words[i]
i++
}
if (!line1) { line1 = words[0]; i = 1 }
const line2 = words.slice(i).join(' ')
return line2 ? `${line1}\n${line2}` : line1
}
// Half the width a label will occupy, floored at the dot so a one-word label
// still reserves the dot's own footprint.
const half_px = label => Math.max(9,
d3.max(String(label).split('\n'), l => l.length) * CHAR_PX / 2)
const rows = surviving
.filter(a => present.has(a.family))
.map(a => ({ ...a, label: wrap_label(a.model) }))
.sort((a, b) => a.first_pub - b.first_pub)
// The x domain is padded by exactly the widest label's half width, solved for
// rather than guessed: with plot width W, data span S days and pad P days,
// P * W / (S + 2P) = half, so P = half * S / (W - 2 * half). An eight-month
// guess left 'PepSeq' 7 px outside the frame.
const x_min = d3.min(rows, a => a.first_pub) ?? new Date(Date.UTC(1984, 0, 1))
const x_max = d3.max(rows, a => a.first_pub) ?? new Date()
const max_half = d3.max(rows, a => half_px(a.label)) ?? 9
const span_days = Math.max(365, (x_max - x_min) / 86400000)
const pad_days = Math.min(3650,
max_half * span_days / Math.max(1, PLOT_WIDTH_PX - 2 * max_half))
const x_pad_min = new Date(+x_min - pad_days * 86400000)
const x_pad_max = new Date(+x_max + pad_days * 86400000)
const px_per_day = PLOT_WIDTH_PX / ((x_pad_max - x_pad_min) / 86400000)
const x_px = d => (d - x_pad_min) / 86400000 * px_per_day
// Greedy first-fit row packing per family: a row is reused only when the two
// labels cannot touch. Rows are filled in date order, so an early paper never
// gets pushed down by a later one.
const row_right = new Map() // `${family}|${row}` -> right edge in px
const rows_used = new Map() // family -> how many rows it needed
const placed = []
for (const row of rows) {
const cx = x_px(row.first_pub)
const half = half_px(row.label)
let r = 0
for (;; r++) {
const key = `${row.family}|${r}`
const right = row_right.get(key)
if (right === undefined || cx - half >= right + LABEL_GAP_PX) {
row_right.set(key, cx + half)
break
}
}
rows_used.set(row.family, Math.max(rows_used.get(row.family) ?? 0, r + 1))
placed.push({ ...row, row: r })
}
// Stack bands bottom-to-top, each exactly as tall as its packed rows.
// visible_bands is already latest-first and y grows upward, so walking it in
// order puts the latest family at the bottom and the earliest at the top,
// which is the order "The long view" uses. Reversing here would insist on
// the opposite and did: the first cut of this put Sparse autoencoder on top.
let y_cursor = 0
const bands = visible_bands.map(family => {
const h = Math.max(1, rows_used.get(family) ?? 1) * ROW_PX
const band = { family, y0: y_cursor, y1: y_cursor + h, center: y_cursor + h / 2, height: h }
y_cursor += h
return band
})
const total_height = y_cursor
const band_index = new Map(bands.map(b => [b.family, b]))
// Row 0 is the TOP row of its band (y grows upward here), so within a lane the
// rows also read earliest-first.
const items = placed.map(d => {
const band = band_index.get(d.family)
return { ...d, y: band.y1 - d.row * ROW_PX - DOT_IN_ROW_PX }
})
const chart = Plot.plot({
width: 1400,
// total_height is in pixels, so 1 y unit = 1 px: the row packing above
// computed real label geometry and the plot must not rescale it.
height: Math.max(240, total_height + 60),
marginLeft: 210, // measured 49 px, then 10 px still over at 195
marginRight: 60,
marginTop: 20,
marginBottom: 40,
x: { type: "time", label: "First publication →", domain: [x_pad_min, x_pad_max], grid: true },
y: { domain: [0, total_height], axis: null },
color: {
domain: all_families,
range: all_families.map(color_of),
legend: false
},
marks: [
// Background band per family, extended to the padded x domain so edge labels stay inside.
Plot.rect(bands, {
x1: () => x_pad_min, x2: () => x_pad_max,
y1: "y0", y2: "y1",
fill: "family",
fillOpacity: 0.09
}),
// Thin separators between bands
Plot.ruleY(bands.flatMap(b => [b.y0, b.y1]), { stroke: "#ddd", strokeWidth: 0.5 }),
// Left-edge family label, and a link: a family with two or more methods
// has a page collecting them, and a one-method family sends the reader
// to its single method instead. See family_href.
Plot.text(bands, {
x: () => x_pad_min,
y: "center",
text: "family",
textAnchor: "end",
dx: -8,
fontSize: 12,
fontWeight: "bold",
fill: "family",
href: d => family_href(d.family),
target: "_self"
}),
// Rich HTML tooltips are wired below via a scoped
// chart.querySelectorAll('g[aria-label="dot"] circle') selector on the
// rendered SVG; no Plot tip is used here.
Plot.dot(items, {
x: "first_pub",
y: "y",
fill: "family",
r: 7,
stroke: "white",
strokeWidth: 1.5,
// Plot wraps each mark in an SVG <a> itself, which sidesteps the
// descending-radius re-sort that makes DOM-index binding fragile here.
href: d => algo_href(d.base_model ?? d.model),
target: "_self"
}),
// One label per dot, always directly above it. The old above-or-below
// split was there to buy vertical space from a fixed band height; the row
// packing gives every label its own 24 px instead, so a single anchor is
// both simpler and collision-free. Names over ~22 chars wrap onto two
// lines; the full name stays in the tooltip.
Plot.text(items, {
x: "first_pub",
y: "y",
text: "label",
textAnchor: "middle",
lineAnchor: "bottom",
dy: -11,
fontSize: 10,
fontWeight: "bold",
fill: "family"
})
]
})
// Rich .map-tooltip on the Plot.dot circles — matches the HTML tooltip style
// used across the world-map / co-auth / bipartite / Sankey / citation-arc
// charts. Plot renders dots in the order of the input `items` array, so we
// bind by DOM order.
const arch_tooltip = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_arch_tip = ev => {
const pad = 12
const box = arch_tooltip.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
arch_tooltip.style.left = `${Math.max(pad, left)}px`
arch_tooltip.style.top = `${Math.max(pad, top)}px`
}
// Plot silently ignores the `className: 'arch-dot'` mark option we passed —
// it labels dot marks with aria-label='dot' instead. querySelectorAll is
// scoped to *this* chart's DOM so we won't pick up other Plot.dot charts.
const arch_dot_circles = chart.querySelectorAll('g[aria-label="dot"] circle')
arch_dot_circles.forEach((circle, i) => {
const d = items[i]
if (!d) return
circle.style.cursor = "pointer"
circle.setAttribute("aria-label", `${d.model}. ${d.family}. First publication ${d.first_pub.toISOString().slice(0,10)}.`)
circle.addEventListener("mouseenter", ev => {
arch_tooltip.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.model}</div>
<div class="map-tooltip-meta">${d.first_pub.toISOString().slice(0,10)}</div>
<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${color_of(d.family)}`}></span>
<span>${d.family}</span>
</div>
${d.description ? html`<div class="map-tooltip-meta" style="margin-top:6px; max-width:280px; white-space:normal;">${d.description}</div>` : ""}
</div>`)
arch_tooltip.classList.add("visible")
arch_tooltip.setAttribute("aria-hidden", "false")
move_arch_tip(ev)
})
circle.addEventListener("mousemove", move_arch_tip)
circle.addEventListener("mouseleave", () => {
arch_tooltip.classList.remove("visible")
arch_tooltip.setAttribute("aria-hidden", "true")
})
})
// Wrap in a horizontally-scrollable container with a fullscreen button.
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${chart}</div>
${arch_tooltip}
</div>`
}The long view
md`Every methodological family in the catalog, placed at the moment it first
appeared. The span runs from **${ff_first_year}** to **${ff_last_year}**,
**${ff_span}** years, across **${ff_rows.length}** distinct families. This is the
only view that fits them on one screen: the chart above gives each family a lane
too, but spends 20 of its 91 height units on Transformer (AR) alone, and the
volume chart below counts papers per year, where the early era is a one-paper
sliver against a linear axis. Families appear here in the same order as the
lanes above.`ff_label_flips = d => {
const lo = new Date(`${ff_first_year - 1}-01-01`).getTime()
const hi = new Date(`${ff_last_year + 1}-01-01`).getTime()
return (d.date.getTime() - lo) / (hi - lo) > 0.62
}
ff_approach = r => r.first_is_dl === true || r.first_is_dl === 1 ? "Deep learning"
: r.first_is_dl === false || r.first_is_dl === 0 ? "Classical"
: "Not applicable"Plot.plot({
width: 1100,
height: 22 * ff_rows.length + 70,
marginLeft: 250,
marginRight: 100, // flipping late labels left cut 182 px of overflow to
// 68; this clears the rest without crushing the axis
marginTop: 34,
x: {
label: "First appearance",
grid: true,
// Fixed to whole years so the 1984-to-1994 gap reads as a real decade of
// quiet rather than being compressed away by a data-driven domain.
domain: [new Date(`${ff_first_year - 1}-01-01`), new Date(`${ff_last_year + 1}-01-01`)]
},
y: { label: null, domain: ff_rows.map(r => r.family), tickSize: 0 },
color: {
legend: true,
domain: ["Classical", "Deep learning", "Not applicable"],
range: ["#6c757d", "#C44E52", "#b0b7bd"]
},
marks: [
// A rule from the axis to the dot: without it, a 49-row dot plot is hard to
// read across 250px of label gutter.
Plot.ruleY(ff_rows, {
y: "family",
x1: new Date(`${ff_first_year - 1}-01-01`),
x2: "date",
stroke: "#e3e6e9",
strokeWidth: 1
}),
Plot.dot(ff_rows, {
x: "date",
y: "family",
r: r => 3 + Math.min(r.n_methods, 9) * 0.55,
fill: ff_approach,
stroke: "white",
strokeWidth: 0.8,
tip: true,
channels: {
"introduced by": "first_method",
"paper": "first_title",
"venue": "first_journal",
"methods in this family": "n_methods"
}
}),
// Labels sit to the RIGHT of the dot, except for late families, whose
// labels would run off the chart: measured 182 px of overflow, worst case
// "Sparse autoencoder interpretability of InstaNovo" at 2026-06-25.
// Flipping them to the left of the dot costs nothing, because the right
// side of a late dot is empty by definition, whereas reserving 215 px of
// marginRight would squeeze 42 years of timeline into less width.
//
// TWO MARKS, not one with functions. Plot's `dx`, `dy` and `textAnchor` are
// CONSTANT options, not channels: passing a function silently applies
// neither, so the first attempt at this left every label centred on its own
// dot with dx 0. That is what put a dot in the middle of all 49 names, and
// it was invisible in the code because a function is a perfectly legal
// value. Measured before: 49 label-over-dot collisions, 9 px each.
//
// dx of 11 clears the largest dot, whose radius is 3 + 9 * 0.55 = 7.95.
Plot.text(ff_rows.filter(d => !ff_label_flips(d)), {
x: "date",
y: "family",
text: "first_method",
dx: 11,
textAnchor: "start",
fontSize: 10,
fill: "#444"
}),
Plot.text(ff_rows.filter(ff_label_flips), {
x: "date",
y: "family",
text: "first_method",
dx: -11,
textAnchor: "end",
fontSize: 10,
fill: "#444"
})
]
})Inputs.table(
ff_rows.map(r => ({
"First seen": r.first_date,
"Family": r.family,
"Approach": ff_approach(r),
"Introduced by": r.first_method,
"Methods since": r.n_methods,
"Paper": r.first_title,
"Venue": r.first_journal || ""
})),
{
rows: 20,
sort: "First seen",
reverse: false,
layout: "auto",
// Same cell-linking convention as the other tables: an internal detail page,
// plus an arrow to the publisher where there is one. Inputs.table passes the
// index into the ORIGINAL array, not the sorted position, so ff_rows[i] is the
// right row however the reader sorts the table.
format: {
// The family name links to the family's own page where it has one, and to
// its single method where it does not; same rule as the swim-lane labels.
"Family": f => {
const href = family_href(f)
return href ? htl.html`<a href=${href}>${f}</a>` : f
},
"Introduced by": (m, i) => page_link("algorithms", ff_rows[i].first_method, m),
"Paper": (t, i) => {
const row = ff_rows[i]
const internal = page_href("publications", row.first_pub_id)
const external = row.first_url || (row.first_doi ? `https://doi.org/${row.first_doi}` : null)
const label = internal ? htl.html`<a href=${internal}>${t}</a>` : t
return external
? htl.html`${label} <a href="${external}" target="_blank" rel="noopener"
title="Publisher / DOI" style="text-decoration:none">↗</a>`
: label
},
"Venue": v => v ? page_link("venues", v, v) : ""
},
width: { "First seen": 92, "Methods since": 92, "Approach": 100 }
}
)How they score
Everything above is what the methods ARE. This is how they do, on the field’s two public benchmarks. Neither is a leaderboard someone can game by reporting their own numbers: both run the tools themselves and publish the results as data.
denovo_benchmarks runs every tool in its own container over 84 datasets, against the same ground-truth PSMs and the same metric code, and asks how a method holds up across instruments, organisms, digests and modifications. That is the first four charts here.
ProteoBench takes one dataset, the published nine-species benchmark, and accepts submitted runs with their parameters recorded. That is the last two. It answers a different question: with these settings, on this data, where does the tool land.
// A tool that ran on every dataset can be ranked against the others; one that
// did not, cannot. The only such entry is casanovo-scaling, a scaling
// experiment over a 24-dataset subset rather than a released tool, so it stays
// in the heatmap, where every cell stands on its own, and out of the two charts
// that compare tools across datasets.
bench_full = bench_t.filter(d => d.n_datasets === bench_n_datasets)bench_provenance = {
const src = benchmark_source
const tools = new Set(bench_full.map(d => d.display_name))
const linked = new Set(bench_full.filter(d => d.algorithm_id != null)
.map(d => d.display_name))
return `**${tools.size} tools** over **${bench_n_datasets} datasets**, ${linked.size}
of them with a page in this catalog. Numbers as committed to
[\`${src.repo}\`](https://github.com/${src.repo}) at
[\`${src.commit.slice(0, 7)}\`](https://github.com/${src.repo}/commit/${src.commit})
(${src.commit_date.slice(0, 10)}), read on ${src.fetched} by \`build_benchmarks.py\`.`
}// Both charts label an axis with a method name, so both get the tick-label
// links every other chart on the page uses.
bench_linkify = chart => linkify_tick_labels(
chart,
new Map(bench_t.filter(d => d.algorithm_id != null)
.map(d => [d.display_name, page_href("algorithms", d.display_name)])
.filter(([, href]) => href))
)Rank across every dataset
bench_ranked = {
const rows = bench_full.filter(d => d[bench_ap_key] != null)
const out = []
for (const [, group] of d3.group(rows, d => d.dataset)) {
group.slice()
.sort((a, b) => d3.descending(a[bench_ap_key], b[bench_ap_key]))
.forEach((d, i) => out.push({ ...d, ap: d[bench_ap_key], rank: i + 1 }))
}
return out
}bench_order = {
// By median rank, best first, with median AP breaking ties: with 17 tools
// over 84 datasets two medians do land on the same half-integer.
const med = d3.rollup(bench_ranked,
v => [d3.median(v, d => d.rank), -d3.median(v, d => d.ap)],
d => d.display_name)
return Array.from(med)
.sort((a, b) => d3.ascending(a[1][0], b[1][0]) || d3.ascending(a[1][1], b[1][1]))
.map(d => d[0])
}bench_rank_counts = {
// One dot per (method, rank): how many datasets put that method there and the
// median AP of those datasets.
const out = []
for (const [, rows] of d3.group(bench_ranked, d => `${d.display_name}|${d.rank}`)) {
out.push({
display_name: rows[0].display_name,
family: rows[0].family,
rank: rows[0].rank,
n: rows.length,
total: bench_n_datasets,
median_ap: d3.median(rows, d => d.ap)
})
}
return out
}{
// Box plus dots, the shape of plot_ranking_boxplot.py in the upstream
// visualisation PR, turned on its side because the tool names are long
// enough that the vertical version needs rotated labels.
//
// The dots are SIZED BY COUNT rather than jittered into a strip. A rank is an
// integer, so 84 datasets land on 17 x positions and a strip plot stacks
// about five dots per position; jittering them inside a 30 px row separates
// nothing and the alpha pile-up reads as a colour, not a count. One dot whose
// area is the number of datasets says the same thing exactly.
const n = bench_order.length
const chart = Plot.plot({
width: 1100,
height: Math.max(320, n * 30 + 90),
marginLeft: 135,
marginRight: 30,
marginBottom: 48,
x: { domain: [0.5, n + 0.5], grid: true, label: "Rank (1 = best)",
labelOffset: 40 },
r: { range: [2, 9] },
y: { domain: bench_order, label: null },
color: { domain: bench_family_domain, range: bench_family_range, legend: false },
marks: [
Plot.boxX(bench_ranked, {
x: "rank", y: "display_name",
fill: "#e3e9ef", stroke: "#57606a", strokeWidth: 1,
// The dots below already show every dataset, so the box's own outlier
// dots would be a second, differently sized copy of some of them.
r: 0
}),
// Grouped in bench_rank_counts rather than by Plot.group, because a
// `channels` entry naming a field the GROUP transform produces ("count")
// does not resolve: the tooltip came out empty, having been told to hide
// x, y, fill and r and then given three channels it could not find.
Plot.dot(bench_rank_counts, {
x: "rank", y: "display_name", r: "n", fill: "family", fillOpacity: 0.85,
stroke: "white", strokeWidth: 0.6,
href: d => page_href("algorithms", d.display_name), target: "_self",
tip: { format: { x: false, y: false, r: false, fill: false } },
channels: {
"Method": "display_name",
"Rank": d => `${d.rank} of ${bench_order.length}`,
"Datasets at this rank": d => `${d.n} of ${d.total}`,
"Median AP here": d => d.median_ap.toFixed(3)
// No list of dataset names: plot_tip_style caps a tip line at 34em
// and the ellipsis eats the end of it, which is the whole list, and a
// method can hold one rank on 72 of the 84 datasets anyway. The
// heatmap below is the chart that answers "which datasets".
}
}),
// A method's own median, drawn on top of its box: the box's median line
// is the same value, but at this row height it is one pixel wide and sits
// under the dots.
Plot.tickX(bench_ranked, Plot.groupY({ x: "median" }, {
x: "rank", y: "display_name", stroke: "#24292f", strokeWidth: 2
}))
],
tip: plot_tip_style
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${bench_linkify(chart)}</div>
</div>`
}Which of those differences are real
md`The box plot above is a description. This is a test: a **critical-difference
diagram**, the standard way to compare many methods over many datasets
([Demšar 2006](https://www.jmlr.org/papers/volume7/demsar06a/demsar06a.pdf)). Two methods are drawn connected by a bar when the
evidence does **not** separate them.
Read it in three steps. **The axis is average rank**, best on the left, so
position alone is the ranking. **The bars are the finding**: any two methods
joined by one are statistically indistinguishable over these
${bench_n_datasets} datasets, at the conventional 5% error rate (α = 0.05),
meaning the test tolerates a one-in-twenty chance of calling a difference real
when it is not. And **the width of a bar is the
critical difference**, here ${bench_cd.cd.toFixed(2)} rank positions: two
methods whose average ranks differ by less than that could have swapped places
by chance.`bench_cd = {
// Nemenyi post-hoc test over the per-dataset ranks, following Demšar (2006),
// "Statistical Comparisons of Classifiers over Multiple Data Sets", JMLR 7:
// https://www.jmlr.org/papers/volume7/demsar06a/demsar06a.pdf
//
// CD = q_alpha * sqrt(k(k+1) / 6N), with k methods, N datasets, and q_alpha
// the Studentized-range critical value divided by sqrt(2). Those come from a
// table, not a formula; these are Demšar's table 5(a) for alpha = 0.05,
// indexed by k.
const Q05 = [NaN, NaN, 1.960, 2.343, 2.569, 2.728, 2.850, 2.949, 3.031,
3.102, 3.164, 3.219, 3.268, 3.313, 3.354, 3.391, 3.426, 3.458,
3.489, 3.517, 3.544]
// The family comes along because the marks colour by it. Leaving it out cost
// an hour: `fill: "family"` with an undefined value silently drops the datum,
// so the leads rendered and the dots and labels did not.
const tools = Array.from(d3.group(bench_ranked, d => d.display_name),
([name, v]) => ({ name, avg: d3.mean(v, d => d.rank),
family: v[0].family ?? "Unknown" }))
.sort((a, b) => d3.ascending(a.avg, b.avg))
const k = tools.length
const N = new Set(bench_ranked.map(d => d.dataset)).size
const q = Q05[k] ?? Q05[Q05.length - 1]
const cd = q * Math.sqrt(k * (k + 1) / (6 * N))
// Friedman test first: the post-hoc comparison is only licensed if the
// omnibus test rejects "all methods are equivalent". chi2_F is Demšar's
// equation (3).
const sum_sq = d3.sum(tools, d => d.avg ** 2)
const chi2 = 12 * N / (k * (k + 1)) * (sum_sq - k * (k + 1) ** 2 / 4)
const df = k - 1
// Maximal cliques of methods that the test cannot separate. Walking the
// rank-sorted list, each clique is the longest run starting at i whose span
// is within CD; a run contained in the previous one adds nothing.
const cliques = []
for (let i = 0; i < k; i++) {
let j = i
while (j + 1 < k && tools[j + 1].avg - tools[i].avg <= cd) j++
if (j > i && !(cliques.length && cliques.at(-1)[1] >= j)) cliques.push([i, j])
}
return { tools, k, N, q, cd, chi2, df, cliques,
p: chi2_sf(chi2, df) }
}// Upper tail of the chi-square distribution, as the regularised incomplete
// gamma Q(df/2, x/2), by series below the crossover and continued fraction
// above it (Numerical Recipes). Needed because the Friedman statistic is the
// licence for the post-hoc test, and quoting a statistic without its p-value
// leaves the reader to look up a table.
chi2_sf = (x, df) => {
if (!(x > 0)) return 1
const a = df / 2, xx = x / 2
const ln_gamma = z => {
// Lanczos approximation; plenty for the two decimal places this is used at.
const g = [676.5203681218851, -1259.1392167224028, 771.32342877765313,
-176.61502916214059, 12.507343278686905, -0.13857109526572012,
9.9843695780195716e-6, 1.5056327351493116e-7]
if (z < 0.5) return Math.log(Math.PI / Math.sin(Math.PI * z)) - ln_gamma(1 - z)
z -= 1
let s = 0.99999999999980993
for (let i = 0; i < g.length; i++) s += g[i] / (z + i + 1)
const t = z + g.length - 0.5
return 0.5 * Math.log(2 * Math.PI) + (z + 0.5) * Math.log(t) - t + Math.log(s)
}
const lead = -xx + a * Math.log(xx) - ln_gamma(a)
if (xx < a + 1) { // series for P, then Q = 1 - P
let sum = 1 / a, term = sum
for (let n = 1; n < 500; n++) {
term *= xx / (a + n)
sum += term
if (Math.abs(term) < Math.abs(sum) * 1e-15) break
}
return 1 - Math.exp(lead + Math.log(sum))
}
let b = xx + 1 - a, c = 1e300, d = 1 / b, h = d // continued fraction for Q
for (let i = 1; i < 500; i++) {
const an = -i * (i - a)
b += 2
d = an * d + b; if (Math.abs(d) < 1e-300) d = 1e-300
c = b + an / c; if (Math.abs(c) < 1e-300) c = 1e-300
d = 1 / d
const del = d * c
h *= del
if (Math.abs(del - 1) < 1e-15) break
}
return Math.exp(lead) * h
}{
// Demšar's layout: a rank axis along the top, each method's average rank
// marked and led out to a label on the nearer side, and one bar per clique of
// methods the test cannot separate.
//
// y is in "label rows" hanging BELOW the axis, so 0 is the axis and the rows
// run negative: the axis is the baseline of this diagram, not a separate
// strip at the top. A linear y scale with the axis switched off, because the
// rows are layout and not data.
const { tools, k, cd, cliques } = bench_cd
const half = Math.ceil(k / 2)
const rows = tools.map((d, i) => ({
...d, i,
left: i < half,
row: -(i < half ? i + 1 : k - i)
}))
const deepest = d3.min(rows, d => d.row)
const bottom = deepest - cliques.length - 1
const x_pad = 0.6
const href = d => page_href("algorithms", d.name)
const chart = Plot.plot({
width: 1100,
height: 70 + (1 - bottom) * 26,
marginLeft: 155,
marginRight: 155,
marginTop: 52,
marginBottom: 16,
x: { domain: [1 - x_pad, k + x_pad], axis: "top", ticks: k, grid: true,
label: "Average rank across the datasets (1 = best) →",
labelAnchor: "center", labelOffset: 38 },
y: { domain: [bottom - 0.5, 0.5], axis: null },
color: { domain: bench_family_domain, range: bench_family_range, legend: false },
marks: [
Plot.ruleY([0], { stroke: "#57606a" }),
// Vertical lead from the axis down to the method's row, then out to the
// label. Two link marks rather than one path: a link is a straight line.
Plot.link(rows, { x1: "avg", x2: "avg", y1: 0, y2: "row",
stroke: "#8c959f", strokeWidth: 1 }),
Plot.link(rows.filter(d => d.left), {
x1: "avg", x2: () => 1 - x_pad, y1: "row", y2: "row",
stroke: "#8c959f", strokeWidth: 1 }),
Plot.link(rows.filter(d => !d.left), {
x1: "avg", x2: () => k + x_pad, y1: "row", y2: "row",
stroke: "#8c959f", strokeWidth: 1 }),
Plot.dot(rows, { x: "avg", y: "row", fill: "family", r: 4,
stroke: "white", strokeWidth: 1, href, target: "_self" }),
Plot.text(rows.filter(d => d.left), {
x: () => 1 - x_pad, y: "row", text: d => `${d.name} ${d.avg.toFixed(2)}`,
textAnchor: "end", dx: -6, fontSize: 11, fontWeight: "bold",
fill: "family", href, target: "_self" }),
Plot.text(rows.filter(d => !d.left), {
x: () => k + x_pad, y: "row", text: d => `${d.avg.toFixed(2)} ${d.name}`,
textAnchor: "start", dx: 6, fontSize: 11, fontWeight: "bold",
fill: "family", href, target: "_self" }),
// One bar per clique: a thick line spanning the average ranks of the
// methods the test cannot tell apart, stacked below the leads.
Plot.link(cliques.map(([a, b], n) => ({ a, b, n })), {
x1: d => tools[d.a].avg, x2: d => tools[d.b].avg,
y1: d => deepest - 1 - d.n, y2: d => deepest - 1 - d.n,
stroke: "#24292f", strokeWidth: 5, strokeLinecap: "round" }),
Plot.text(cliques.map(([a, b], n) => ({ a, b, n })), {
x: d => tools[d.b].avg, y: d => deepest - 1 - d.n,
text: d => `${d.b - d.a + 1} methods, ranks ${tools[d.a].avg.toFixed(2)}-${tools[d.b].avg.toFixed(2)}`,
dx: 8, textAnchor: "start", fontSize: 10, fill: "#57606a" }),
// The critical difference itself, drawn to scale just under the axis.
Plot.link([0], { x1: () => 1, x2: () => 1 + cd, y1: -0.45, y2: -0.45,
stroke: "#cf222e", strokeWidth: 3,
strokeLinecap: "round" }),
Plot.text([0], { x: () => 1 + cd / 2, y: -0.45, dy: -9, fontSize: 11,
fill: "#cf222e", fontWeight: "bold",
text: () => `critical difference = ${cd.toFixed(2)}` })
]
})
return html`<div class="chart-wrap">
<div class="chart-scroll">${chart}</div>
</div>`
}md`**Ask whether ANY of them differ before asking which.** That first question
is the *omnibus* test, and it is not a formality: ${bench_cd.k} methods make
${bench_cd.k * (bench_cd.k - 1) / 2} pairs, and running that many comparisons on
data where nothing is happening will still turn up a few that look convincing.
Clearing the omnibus test is what licenses the bars above.
**Friedman's test** is the one for this shape of data. It treats each dataset as
a judge that ranks all ${bench_cd.k} methods, and asks whether those rankings
agree with each other more than judges assigning ranks at random would. Working
on ranks rather than on the AP values themselves is what makes it safe here: it
assumes nothing about how AP is distributed, and a dataset where everything
scores around 0.9 counts exactly as much as one where everything scores 0.3.
**χ² ("chi-squared") is the number that test produces**: one summary of how far
the ${bench_cd.k} average ranks sit from all-equal. On its own it means nothing:
it has to be read against a reference distribution, here with
${bench_cd.df} degrees of freedom, one fewer than the number of methods. That
conversion is what turns χ² = ${bench_cd.chi2.toFixed(0)} into
**p ${bench_cd.p < 2.3e-16 ? "< 2.3 × 10⁻¹⁶" : "= " + bench_cd.p.toExponential(1)}**,
the probability of rankings this consistent if every method were equally good.
They are not equally good, so asking which pairs differ is a fair question.
**What a bar does and does not say.** A bar means the test did not find a
difference; it does not mean the two methods are equally good. With
${bench_cd.N} datasets the test is sensitive enough that a surviving bar is a
genuine statement about similarity, but the reverse reading, "not significantly
different therefore the same", is the one mistake this diagram invites. The
other is to forget that a rank throws away the size of a difference: a method
that loses every dataset by 0.001 AP ranks last as decisively as one that loses
by 0.3.
And the test sees only the ${bench_cd.N} datasets it was given. They are not a
random sample of proteomics experiments; ${bench_group_share}% of them come from
the single largest group in the heatmap below.`Precision against coverage
md`Every tool scores every spectrum it is handed, so precision is a function of
how many of those answers you keep. These curves are that trade-off, averaged
over all ${bench_n_datasets} datasets: down the y axis is "how right is it when
it answers", across the x axis is "how often does it answer". **AP** in the
legend is the area under the curve, and the greyed lines are the tools you have
not ticked.`viewof bench_show = Inputs.checkbox(bench_order, {
label: "Highlight",
// Everything on by default. The grey-behind treatment is what the benchmark
// manuscript's figure 1b does, and it is still one click away, but a chart
// that starts by hiding twelve of seventeen tools answers a narrower question
// than the reader asked.
value: bench_order
})bench_ap_of_curve = {
// Trapezoid over the 101 grid points, per tool and level. This is the number
// the legend quotes, and it is NOT the median of the per-dataset APs: it is
// the AP of the AVERAGED curve, which is what the benchmark manuscript's
// figure 1b reports. The two differ by up to 0.12 (InstaNovo: median 0.854,
// averaged curve 0.736) because averaging curves is not averaging areas.
const out = new Map()
for (const [key, pts] of d3.group(bench_curves_t, d => `${d.level}|${d.tool}`)) {
const p = pts.slice().sort((a, b) => d3.ascending(a.coverage, b.coverage))
let area = 0
for (let i = 1; i < p.length; i++) {
area += (p[i].precision + p[i - 1].precision) / 2
* (p[i].coverage - p[i - 1].coverage)
}
out.set(key, area)
}
return out
}{
const rows = bench_curves_t.filter(d => d.level === bench_level
&& d.n_datasets === bench_n_datasets)
const shown = new Set(bench_show)
const ap = d => bench_ap_of_curve.get(`${bench_level}|${d.tool}`) ?? 0
const label = d => `${d.display_name} (AP = ${ap(d).toFixed(3)})`
const highlighted = rows.filter(d => shown.has(d.display_name))
// Legend order follows AP, so the legend doubles as the ranking. Built from
// the tools rather than from the labels, which saves parsing the number back
// out of the string.
const domain = Array.from(d3.group(highlighted, d => d.tool))
.sort((a, b) => d3.descending(ap(a[1][0]), ap(b[1][0])))
.map(d => label(d[1][0]))
const chart = Plot.plot({
width: 1100,
height: 560,
marginLeft: 60,
marginBottom: 48,
x: { label: `${bench_level_label} coverage`, domain: [0, 1], grid: true,
labelOffset: 40 },
y: { label: `${bench_level_label} precision`, domain: [0, 1], grid: true },
// Seventeen legend entries at full width fit four to a row.
color: { domain, range: domain.map((_, i) => bench_curve_colour(i)),
legend: true, columns: "25%" },
marks: [
// Both line marks carry the same tip, so an unticked grey line still
// answers "which tool is this". `stroke` is hidden on the highlighted
// mark because its channel value is the legend string, which already
// ends in the AP the tip prints on its own row.
Plot.line(rows.filter(d => !shown.has(d.display_name)), {
x: "coverage", y: "precision", z: "tool",
stroke: "#c6cdd5", strokeWidth: 1,
tip: { format: { x: false, y: false, z: false } },
channels: bench_curve_tip(ap)
}),
Plot.line(highlighted, {
x: "coverage", y: "precision", z: "tool",
stroke: label, strokeWidth: 2,
tip: { format: { x: false, y: false, z: false, stroke: false } },
channels: bench_curve_tip(ap)
}),
Plot.ruleY([0]), Plot.ruleX([0])
]
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${chart}</div>
</div>`
}// The AP is the whole curve's; the coverage and precision are the point the
// cursor is nearest, which is what makes a hover on a bundle of 17 lines useful
// rather than just an identification.
bench_curve_tip = ap => ({
"Method": "display_name",
"AP": d => ap(d).toFixed(3),
"at coverage": d => d.coverage.toFixed(2),
"precision": d => d.precision.toFixed(3)
})// With every tool ticked by default there are 17 lines, and Tableau10 would
// repeat a colour every tenth. Extending it with Set2 and Dark2 gives 17
// distinct ones in a consistent register; the ORDER is by AP, so neighbours in
// the legend are neighbours in the ranking.
bench_curve_colour = i => (d3.schemeTableau10.concat(d3.schemeSet2, d3.schemeDark2))[i % 26]Every tool on every dataset
{
const rows = bench_t.filter(d => d[bench_ap_key] != null)
.map(d => ({ ...d, ap: d[bench_ap_key] }))
// Median AP orders everything: rows (tools), columns (datasets) and column
// groups (categories). Note the columns are NOT faceted by category: Plot
// shares the x scale across facets, so 17 facets would each carry all 84
// columns. The grouping is drawn instead, as a header per group plus a rule
// at its left edge.
const med_by = keyfn => {
const med = d3.rollup(rows, v => d3.median(v, d => d.ap), keyfn)
return Array.from(med).sort((a, b) => d3.descending(a[1], b[1])).map(d => d[0])
}
// Column and group order come from the tools that ran EVERYWHERE, so a
// partially evaluated tool cannot move a dataset's difficulty ranking. Row
// order uses every tool, because a partially evaluated tool still needs a
// row to sit in.
const full = rows.filter(d => d.n_datasets === bench_n_datasets)
const med_full = keyfn => {
const med = d3.rollup(full, v => d3.median(v, d => d.ap), keyfn)
return Array.from(med).sort((a, b) => d3.descending(a[1], b[1])).map(d => d[0])
}
const cat_order = med_full(d => d.cat_group)
const ds_med = d3.rollup(full, v => d3.median(v, d => d.ap), d => d.dataset)
const ds_cat = new Map(rows.map(d => [d.dataset, d.cat_group]))
const datasets = Array.from(ds_med.keys()).sort((a, b) =>
d3.ascending(cat_order.indexOf(ds_cat.get(a)), cat_order.indexOf(ds_cat.get(b)))
|| d3.descending(ds_med.get(a), ds_med.get(b)))
// One header per group, anchored on the group's middle column, plus a rule
// half a band to the left of its first column.
const groups = cat_order.map(category => {
const members = datasets.filter(d => ds_cat.get(d) === category)
return { category, n: members.length, first: members[0],
mid: members[Math.floor(members.length / 2)] }
})
const width = 1400, marginLeft = 135, marginRight = 20
const band = (width - marginLeft - marginRight) / datasets.length
const tools = med_by(d => d.display_name)
// A run that does not exist upstream would otherwise show the page
// background, and on a white-to-teal scale that is indistinguishable from an
// average precision of 0. Grey says "not run" and cannot be misread as a
// score. 60 of the 1512 cells, all of them casanovo-scaling's.
const present = new Set(rows.map(d => `${d.display_name}|${d.dataset}`))
const missing = tools.flatMap(t => datasets
.filter(ds => !present.has(`${t}|${ds}`))
.map(ds => ({ display_name: t, dataset: ds })))
const chart = Plot.plot({
width,
height: tools.length * 22 + 320,
marginLeft, marginRight,
marginTop: 34,
marginBottom: 230, // the rotated dataset names live in here
x: { domain: datasets, tickRotate: -90, label: null, tickSize: 0 },
y: { domain: tools, label: null },
color: { scheme: "BuGn", domain: [0, 1] },
style: { fontSize: "10px" },
marks: [
Plot.cell(missing, {
x: "dataset", y: "display_name", fill: "#e7eaee", inset: 0.5,
title: d => `${d.display_name} was not run on ${d.dataset}`
}),
Plot.cell(rows, {
x: "dataset", y: "display_name", fill: "ap", inset: 0.5,
tip: { format: { x: false, y: false, fill: false } },
channels: { "Method": "display_name", "Dataset": "dataset",
"Kind": "category", "AP": d => d.ap.toFixed(3) }
}),
Plot.text(groups, {
x: "mid", text: "category", frameAnchor: "top", dy: -20,
fontWeight: "bold", fontSize: 11, fill: "#57606a"
}),
Plot.ruleX(groups.slice(1), {
x: "first", dx: -band / 2, stroke: "#8c959f", strokeWidth: 1,
insetTop: -14
})
],
tip: plot_tip_style
})
// The ramp legend is rendered here rather than by `legend: true` on the
// scale: Plot sizes that one to the ramp exactly, so the 0.0 and 1.0 tick
// labels, which are centred on the ends, hang 7 and 8 px outside its frame.
const legend = Plot.legend({
color: { scheme: "BuGn", domain: [0, 1],
label: `${bench_level_label}-level average precision` },
width: 320, marginLeft: 18, marginRight: 18
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
${legend}
<div class="chart-scroll">${bench_linkify(chart)}</div>
</div>`
}The table
bench_summary = {
const rank_by = d3.rollup(bench_ranked, v => d3.median(v, d => d.rank),
d => d.display_name)
const rows = []
for (const [name, mine] of d3.group(bench_t, d => d.display_name)) {
const aps = mine.map(d => d[bench_ap_key]).filter(v => v != null)
if (!aps.length) continue
rows.push({
Method: name,
Family: mine[0].family ?? "",
Version: mine[0].version,
Datasets: mine[0].n_datasets,
"Median rank": rank_by.get(name) ?? null,
"Median AP": d3.median(aps),
"AP of mean curve": bench_ap_of_curve.get(`${bench_level}|${mine[0].tool}`) ?? null,
"Worst": d3.min(aps),
"Best": d3.max(aps)
})
}
return rows.sort((a, b) => d3.descending(a["Median AP"], b["Median AP"]))
}Inputs.table(bench_summary, {
rows: 20,
sort: "Median AP",
reverse: true,
layout: "auto",
format: {
Method: (m, i) => page_link("algorithms", bench_summary[i].Method, m),
"Median rank": v => v == null ? "" : v.toFixed(1),
"Median AP": v => v.toFixed(3),
"AP of mean curve": v => v == null ? "" : v.toFixed(3),
Worst: v => v.toFixed(3),
Best: v => v.toFixed(3)
},
width: { Method: 150, Family: 150, Version: 110, Datasets: 80 }
})bench_zeros = {
// Exact zeros are worth counting rather than hiding: they are what a tool
// scores on a condition it does not support at all.
const z = bench_full.filter(d => d[bench_ap_key] === 0)
const per_dataset = d3.rollups(z, v => v.length, d => d.dataset)
.sort((a, b) => d3.descending(a[1], b[1]))
return {
cells: z.length,
total: bench_full.filter(d => d[bench_ap_key] != null).length,
datasets: new Set(z.map(d => d.dataset)).size,
tools: new Set(z.map(d => d.display_name)).size,
worst: per_dataset[0] ?? ["-", 0]
}
}md`Four things worth knowing before quoting any of this.
**A tool's latest version, not its best.** Four tools have two versions in the
results: \`_old\` twins for PEAKS and Novor, and two container builds each for
InstaNovo and π-HelixNovo. The charts take the newest version that ran on each
dataset. Taking the higher-scoring one instead, as the upstream visualisation
PR's loader does, lets a tool use a different version on every dataset; it
changes 83 of the 1452 (tool, dataset) pairs in the results, with a median
difference of 0.000 and a maximum of 0.209.
**These are not the manuscript's numbers.** The benchmark preprint's figure 1
was drawn over 29 datasets. This is the ${bench_n_datasets} now in the
repository, run with newer containers. The shape of the story survives; the
individual numbers do not.
**A rank is only as fair as the dataset set.** Every tool ranked above ran on
all ${bench_n_datasets} datasets, which is why casanovo-scaling, evaluated on
24, appears in the heatmap but not in the ranking. Two tools can also sit close
enough that their order is noise, which is what the box plot is for; the
upstream PR carries a critical-difference diagram that tests it properly.
**A zero is a real answer, not a failed run.** ${bench_zeros.cells} of the
${bench_zeros.total} cells are exactly 0, spread over ${bench_zeros.datasets}
datasets and ${bench_zeros.tools} tools, and \`${bench_zeros.worst[0]}\` alone
accounts for ${bench_zeros.worst[1]} of them: a LysN digest with TMT labelling
that most tools do not support at all. Their amino-acid scores on the same
spectra are not zero, which is the tell. The benchmark deliberately keeps those
runs in, because the aggregate is meant to be practical performance over the
whole benchmark rather than performance on each method's own supported domain.`The other benchmark: ProteoBench
md`[ProteoBench](${pb_src.module_url}) asks a narrower question than the 84-dataset
sweep above, and answers it in more detail. One dataset, the published
nine-species benchmark of ${(+pb_t[0].n_spectra).toLocaleString("en-US")} spectra;
one point per **submitted run**, with the parameters that produced it recorded
alongside the numbers. ${pb_subs.length} runs so far, all
${pb_subs.filter(d => d.algorithm_id != null).length} of them methods in this
catalog. Results as committed to
[\`${pb_src.repo}\`](https://github.com/${pb_src.repo}) at
[\`${pb_src.commit.slice(0, 7)}\`](https://github.com/${pb_src.repo}/commit/${pb_src.commit})
(${pb_src.commit_date.slice(0, 10)}).`// One row per submission, carrying the four metrics of the selected level and
// stringency. The peptide and amino-acid rows are the two axes of the scatter.
pb_subs = {
const by = d3.group(pb_t.filter(d => d.match_type === pb_match), d => d.id)
const out = []
for (const [id, rows] of by) {
const pep = rows.find(d => d.level === "peptide")
const aa = rows.find(d => d.level === "aa")
if (!pep || !aa) continue
out.push({
...pep, id,
x: pep[pb_metric], y: aa[pb_metric],
pep_precision: pep.precision, pep_recall: pep.recall,
pep_auc: pep.auc, pep_coverage: pep.coverage,
aa_precision: aa.precision, aa_recall: aa.recall,
aa_auc: aa.auc, aa_coverage: aa.coverage
})
}
return out.sort((a, b) => d3.descending(a.x, b.x))
}{
// ProteoBench's own main plot: the same metric at peptide level against
// amino-acid level, on fixed [-0.05, 1.1] axes split into four named
// quadrants at 0.5. The names and their meanings are ProteoBench's.
//
// Not reproduced: the background gradient from light at the origin to dark in
// the top-right corner. Here the colour channel carries the architecture
// family, as it does in every other chart on this page, and two colour
// encodings in one frame is one too many.
const R = [-0.05, 1.1], MID = 0.5
const quadrants = [
{ x: (MID + R[1]) / 2, y: R[1] - 0.04, label: "Good performance",
note: "high at both levels" },
{ x: (R[0] + MID) / 2, y: R[1] - 0.04, label: "Near-miss",
note: "often wrong, but only by a residue or two" },
{ x: (R[0] + MID) / 2, y: R[0] + 0.05, label: "Low performance", note: "" },
{ x: (MID + R[1]) / 2, y: R[0] + 0.05, label: "Alternative candidate",
note: "when wrong, wrong about the whole peptide" }
]
const chart = Plot.plot({
width: 760,
height: 620,
marginLeft: 60,
marginBottom: 48,
x: { domain: R, grid: true, label: `Peptide-level ${pb_metric_label} →`,
labelOffset: 40 },
y: { domain: R, grid: true, label: `↑ Amino-acid-level ${pb_metric_label}` },
color: { domain: bench_family_domain, range: bench_family_range, legend: false },
marks: [
Plot.ruleX([MID], { stroke: "#8c959f", strokeDasharray: "4 3" }),
Plot.ruleY([MID], { stroke: "#8c959f", strokeDasharray: "4 3" }),
Plot.text(quadrants, {
x: "x", y: "y", text: d => d.note ? `${d.label}\n${d.note}` : d.label,
fill: "#8c959f", fontSize: 11, fontWeight: "bold", lineWidth: 16,
textAnchor: "middle"
}),
Plot.dot(pb_subs, {
x: "x", y: "y", fill: "family", r: 7, stroke: "white", strokeWidth: 1.5,
href: d => page_href("algorithms", d.display_name), target: "_self",
tip: { format: { x: false, y: false, fill: false } },
channels: {
"Method": "display_name", "Version": d => d.version ?? "—",
"Peptide": d => d.x.toFixed(3), "Amino acid": d => d.y.toFixed(3),
"Peptide coverage": d => d.pep_coverage.toFixed(3),
"Decoding": d => d.decoding ?? "—",
"Submitted": d => d.submitted ?? "—"
}
}),
// Labels above or below the dot, on one of two rows. FOUR marks, because
// `dy` and `lineAnchor` are constants and not channels: passing a
// function applies neither, silently. pb_placed decides which.
...[[true, 0, -11, "bottom"], [false, 0, 11, "top"],
[true, 1, -28, "bottom"], [false, 1, 28, "top"]].map(
([above, row, dy, lineAnchor]) => Plot.text(
pb_placed.filter(d => d.above === above && d.row === row), {
x: "x", y: "y", text: "display_name", dy, lineAnchor,
fontSize: 11, fontWeight: "bold", fill: "family"
}))
],
tip: plot_tip_style
})
return html`<div class="chart-wrap"><div class="chart-scroll">${chart}</div></div>`
}pb_placed = {
// Seven labels, so a greedy pass is enough: take them in metric order and put
// each above its dot when that box is free, below it otherwise. Obstacles are
// the labels already placed AND every dot, because two of these submissions
// sit a pixel apart (ContraNovo and π-PrimeNovo) and a label clear of its own
// dot can still land on its neighbour's.
//
// Boxes are in DATA units, and x and y need their own conversion: the frame
// is 680 px wide and 552 px tall over the same 1.15-unit domain.
const R = [-0.05, 1.1], span = R[1] - R[0]
const px_x = 680 / span, px_y = 552 / span
// 7.2 px per character at 11 px bold, plus a 3 px pad: the first cut used 6.4
// and no pad, and ContraNovo against π-PrimeNovo still overlapped by 4 px --
// an underestimated box is an invisible licence to overlap.
const LINE_PX = 15, CHAR_PX = 7.2, PAD_PX = 3, DOT_PX = 9
// Four candidate positions: above or below the dot, on the near row or a
// second row further out. Two rows are needed, not one: ContraNovo and
// π-PrimeNovo sit a pixel apart vertically, and AdaNovo's dot blocks the near
// row above both of them, so with two positions they both ended up below and
// overlapped each other.
const SLOTS = [[true, 0, 11], [false, 0, 11], [true, 1, 28], [false, 1, 28]]
const dots = pb_subs.map(d => ({
id: d.id,
x0: d.x - DOT_PX / px_x, x1: d.x + DOT_PX / px_x,
y0: d.y - DOT_PX / py(), y1: d.y + DOT_PX / py()
}))
function py() { return px_y }
const hits = (a, b) => a.x0 < b.x1 && a.x1 > b.x0 && a.y0 < b.y1 && a.y1 > b.y0
const placed = []
const out = []
for (const d of pb_subs) {
const w = (d.display_name.length * CHAR_PX + 2 * PAD_PX) / px_x
const h = LINE_PX / px_y
for (const [i, [above, row, gap_px]] of SLOTS.entries()) {
const gap = gap_px / px_y
const y0 = above ? d.y + gap : d.y - gap - h
const box = { x0: d.x - w / 2, x1: d.x + w / 2, y0, y1: y0 + h }
const clash = placed.some(o => hits(box, o))
|| dots.some(o => o.id !== d.id && hits(box, o))
// The last slot is taken whether it clashes or not: a label that is not
// drawn at all is worse than one that is 4 px from its neighbour, and the
// chart audit will say if that ever happens.
if (!clash || i === SLOTS.length - 1) {
placed.push(box)
out.push({ ...d, above, row })
break
}
}
}
return out
}md`**Which metric you pick changes the answer**, which is the whole argument in
ProteoBench's own design discussion. ${pb_flip.top_precision} has the highest
peptide-level precision of the ${pb_subs.length} runs, but it answers only
${(pb_flip.coverage * 100).toFixed(0)}% of the spectra; rank the same runs by AUC
and it comes ${pb_flip.rank_by_auc}, behind ${pb_flip.top_auc}. Precision alone
rewards a tool for keeping quiet:
- **Precision** is correct predictions over predictions *made*, at whatever
coverage the tool chose. A tool that answers only the easy spectra scores well.
- **Precision at full coverage** counts an unanswered spectrum as wrong. It is
the same number as recall, and it is what the discussion asked for on the
grounds that some tools skip more than 10% of spectra.
- **AUC** is the area under the precision-coverage curve, swept over every score
threshold, which is where that discussion landed as the default. It is the
metric the 84-dataset charts above use as well, so the two benchmarks can be
read against each other.
*Mass-based* matching accepts a prediction whose residue masses line up within
tolerance, so it cannot tell I from L or deamidated Q/N from E/D; *exact*
requires the sequence. Mass-based is therefore always the higher number, and the
gap is not small: ${pb_gap.median_pct}% of the peptide-level score on median
across these runs.`pb_flip = {
const rank = (key) => pb_t.filter(d => d.match_type === pb_match && d.level === "peptide")
.slice().sort((a, b) => d3.descending(a[key], b[key]))
const by_prec = rank("precision"), by_auc = rank("auc")
const top = by_prec[0]
const pos = by_auc.findIndex(d => d.id === top.id) + 1
const ordinal = ["", "first", "second", "third", "fourth", "fifth", "sixth",
"seventh", "eighth", "ninth", "tenth"][pos] ?? `${pos}th`
return { top_precision: top.display_name, coverage: top.coverage,
rank_by_auc: ordinal, top_auc: by_auc[0].display_name }
}pb_gap = {
// How much of the peptide-level score is mass-based matching worth, as a
// percentage of the mass-based number, on median across submissions.
const mass = new Map(pb_t.filter(d => d.level === "peptide" && d.match_type === "mass")
.map(d => [d.id, d[pb_metric]]))
const gaps = pb_t.filter(d => d.level === "peptide" && d.match_type === "exact")
.map(d => (mass.get(d.id) - d[pb_metric]) / mass.get(d.id) * 100)
.filter(v => isFinite(v))
return { median_pct: (d3.median(gaps) ?? 0).toFixed(0) }
}md`#### The curves behind the AUC
Each run's own precision-coverage curve, at the level and stringency selected
above. This is what a single precision number summarises, and what the
84-dataset chart averages: a curve that holds high precision and only gives way
near full coverage is a better tool than one that drops immediately and stays
flat, and the two can report the same precision.`{
const rows = pb_curves_t.filter(d => d.level === pb_curve_level
&& d.match_type === pb_match)
const chart = Plot.plot({
width: 760,
height: 420,
marginLeft: 60,
marginBottom: 48,
x: { domain: [0, 1.03], grid: true, labelOffset: 40,
label: `${pb_curve_level === "peptide" ? "Peptide" : "Amino acid"} coverage →` },
y: { domain: [0, 1], grid: true,
label: `↑ ${pb_curve_level === "peptide" ? "Peptide" : "Amino acid"} precision` },
color: { domain: pb_curve_domain, range: pb_curve_range, legend: true,
columns: "33%" },
marks: [
Plot.line(rows, {
x: "coverage", y: "precision", z: "display_name",
stroke: "display_name", strokeWidth: 1.8,
tip: { format: { z: false } },
channels: { "Method": "display_name" }
}),
Plot.ruleY([0]), Plot.ruleX([0])
],
tip: plot_tip_style
})
return html`<div class="chart-wrap"><div class="chart-scroll">${chart}</div></div>`
}// Seven curves need seven distinguishable colours, which the family palette
// cannot give: five of these seven runs are Transformer (AR) and would be the
// same orange. Ordered by AUC so the legend doubles as a ranking.
pb_curve_domain = pb_t
.filter(d => d.level === pb_curve_level && d.match_type === pb_match)
.slice().sort((a, b) => d3.descending(a.auc, b.auc)).map(d => d.display_name)md`Two caveats specific to this benchmark. A submission records **one
configuration**, so a run is evidence about a checkpoint and its parameters
rather than about a method at its best: the table's decoding column is part of
the result. And amino-acid coverage can exceed 1, because the denominator is the
ground truth's residue count and a prediction may be longer than the peptide it
is predicting.`pb_table = pb_subs.map(d => ({
Method: d.display_name,
Version: d.version ?? "",
Family: d.family ?? "",
Decoding: d.decoding ?? "",
"Peptide precision": d.pep_precision,
"Peptide @cov 1": d.pep_recall,
"Peptide AUC": d.pep_auc,
"Peptide coverage": d.pep_coverage,
"AA precision": d.aa_precision,
"AA AUC": d.aa_auc,
Submitted: d.submitted ?? "",
"ProteoBench": d.pb_version ?? ""
}))Inputs.table(pb_table, {
rows: 12,
sort: "Peptide AUC",
reverse: true,
layout: "auto",
format: {
Method: (m, i) => page_link("algorithms", pb_table[i].Method, m),
"Peptide precision": v => v.toFixed(3),
"Peptide @cov 1": v => v.toFixed(3),
"Peptide AUC": v => v.toFixed(3),
"Peptide coverage": v => v.toFixed(3),
"AA precision": v => v.toFixed(3),
"AA AUC": v => v.toFixed(3)
},
width: { Method: 130, Version: 80, Family: 130, Decoding: 150 }
})The data underneath
Every number on this page was produced by running software over spectra, and the spectra have versions. “Trained on nine-species” names four different objects, and a third of the papers that say it do not say which. These two charts are about that.
// One row per (dataset, version-or-null) with its paper count. The null is
// preserved as null all the way from SQL so that "not stated" cannot be
// confused with a version that happens to be called that.
dataset_version_use = {
const rows = new Map()
for (const u of dataset_usage_t) {
const key = `${u.dataset}|${u.version ?? ""}`
const r = rows.get(key) ?? { dataset: u.dataset, kind: u.kind,
version: u.version ?? null, papers: 0 }
r.papers += 1
rows.set(key, r)
}
return Array.from(rows.values())
}// A dataset earns a row only if naming it is actually ambiguous: it has more
// than one version in use, or at least one paper that named no version. A
// single-version deposit cited once is not evidence about anything and there
// are several hundred of them.
version_ambiguity_rows = {
const by_dataset = d3.group(dataset_version_use, d => d.dataset)
const keep = []
for (const [name, rows] of by_dataset) {
const total = d3.sum(rows, d => d.papers)
const unstated = d3.sum(rows.filter(d => d.version === null), d => d.papers)
const versions = new Set(rows.filter(d => d.version !== null).map(d => d.version))
// Two papers is the floor for "is naming this ambiguous": one paper cannot
// disagree with anything. The threshold is NOT a slider, because it would
// be inert -- measured, exactly three datasets in the catalog qualify at
// any floor from 2 to 4, so a control would imply a longer list than
// exists.
if (total < 2) continue
if (versions.size < 2 && unstated === 0) continue
for (const r of rows) keep.push({ ...r, total, unstated })
}
return keep.sort((a, b) => b.total - a.total || a.dataset.localeCompare(b.dataset)
|| (a.version ?? "").localeCompare(b.version ?? ""))
}html`<p style="margin:0.4rem 0 0.9rem; color:#57606a;">
Of the <b>${dataset_versions_t.length}</b> dataset versions the catalog records,
naming is genuinely ambiguous for only
<b>${d3.group(version_ambiguity_rows, d => d.dataset).size}</b> datasets: the
rest are cited by one paper, or have one version, or are always pinned. Those
three carry <b>${version_unstated_share.total}</b> uses, of which
<b>${version_unstated_share.unstated}</b> (<b>${version_unstated_share.pct}%</b>)
name no version. The amber segment is that: a citation that does not identify
which curated object was run.</p>`// Colour carries "did the paper pin a version", not which version, because
// version names are not comparable between datasets: 'revised (main)' and
// 'Bruker timsTOF' share nothing. The version itself is on the segment label
// and in the tooltip, where it does not need a shared scale.
//
// This is a CATEGORY, not a colour. Handing Plot `fill: d => "#d97706"` makes
// the hex string a channel value, which Plot then maps through the default
// categorical colour scale: the chart renders in scheme colours and the hex is
// never used. Measured -- the amber appeared nowhere in the DOM. A named
// category plus an explicit range is the fix, and it earns a legend.
version_status = d => d.version === null ? "version not stated" : "version stated"Plot.plot({
width: 1100,
height: Math.max(220, 46 * d3.group(version_ambiguity_rows, d => d.dataset).size + 80),
marginLeft: 250,
marginRight: 90,
marginBottom: 48,
// "Uses", not "papers": a paper that names two versions of one dataset
// contributes two rows here, so the bar total is link count and exceeds the
// distinct paper count (nine-species: 44 uses across 34 papers). The table
// below counts distinct papers, and the two disagreeing was a labelling bug.
x: { label: "Uses, counted once per version named", grid: true, nice: true,
labelOffset: 40 },
y: { label: null },
color: {
domain: ["version stated", "version not stated"],
range: ["#0f766e", "#d97706"],
legend: true
},
marks: [
// Unstated sorts last, so the amber block always ends every bar and can be
// compared across rows without reading a single label.
Plot.barX(version_ambiguity_rows, Plot.stackX({
order: d => d.version === null ? 1 : 0,
y: "dataset",
x: "papers",
fill: version_status,
stroke: "white",
strokeWidth: 1,
channels: { "Dataset": "dataset", "Version": d => d.version ?? "not stated",
"Papers": "papers" },
tip: { ...plot_tip_style, format: { x: false, y: false, fill: false } }
})),
// The version name on the segment, but only where it fits. 6.4 px per
// character at 10 px, the same basis as the other charts here; anything
// narrower stays unlabelled rather than overlapping its neighbour. The
// x-per-paper factor is derived from the widest bar, not assumed.
Plot.text(version_ambiguity_rows, Plot.stackX({
order: d => d.version === null ? 1 : 0,
y: "dataset",
x: "papers",
text: d => d.version ?? "not stated",
fill: "white",
fontSize: 10,
fontWeight: 600,
filter: d => d.papers * (760 / d3.max(version_ambiguity_rows, r => r.total))
> ((d.version ?? "not stated").length * 6.4 + 10)
})),
Plot.text(d3.rollups(version_ambiguity_rows, v => d3.sum(v, d => d.papers), d => d.dataset)
.map(([dataset, papers]) => ({ dataset, papers })), {
y: "dataset", x: "papers", text: d => d.papers, dx: 14,
textAnchor: "start", fontSize: 11, fill: "#57606a"
}),
Plot.ruleX([0])
]
})An amber segment is not sloppiness, usually. A paper that cites the nine per-species PRIDE submissions has told you exactly where its spectra came from. What it has not told you is which curated re-release it ran, and since the revised benchmark removed cross-species peptide redundancy that the original leaked, the two are not interchangeable for measuring generalisation. The catalog records this as a null version rather than guessing.
Where the nine-species benchmark comes from
nine_species_flow = {
const DATASET = "Nine-species benchmark"
const versions = dataset_versions_t.filter(d => d.dataset === DATASET)
const addrs = dataset_addresses_t.filter(d => d.dataset === DATASET)
const provenance = addrs.filter(a => a.is_provenance)
const use = dataset_version_use.filter(d => d.dataset === DATASET)
const papers_of = new Map(use.map(d => [d.version ?? "__none__", d.papers]))
const width = 1100
const PAD = 10
const row_px = 34
const n_rows = Math.max(provenance.length, versions.length + 1)
const height = Math.max(430, 60 + n_rows * row_px)
// left 300: 'Candidatus Thiodiazotropha endoloripes . PXD004536' is the
// longest provenance label and was clipped by 42 px at 232, measured by
// check_chart_overlap.py. Widening beats truncating a species name.
const margin = { top: 44, right: 210, bottom: 34, left: 300 }
const inner_w = width - margin.left - margin.right
const inner_h = height - margin.top - margin.bottom
// Three columns: the nine source submissions, the curated deposit, then the
// versions. Column x positions are fixed; only the version column carries a
// weight, which is why the two halves are drawn differently (see below).
const col_x = [0, inner_w * 0.46, inner_w * 0.82]
const NODE_W = 13
// LEFT: nine equal boxes. Provenance has no spectra count in the catalog and
// inventing one would make the ribbons lie, so these are equal height and
// their ribbons are drawn thin and dashed to say "this is lineage, not flow".
const left_h = (inner_h - PAD * (provenance.length - 1)) / provenance.length
const left = provenance.map((a, i) => ({
id: a.accession, label: a.part ?? a.accession, url: a.url,
y0: i * (left_h + PAD), y1: i * (left_h + PAD) + left_h
}))
// MIDDLE: the single curated deposit.
const mid = { id: "MSV000081382", label: "MSV000081382",
sub: "DeepNovo's curated deposit",
y0: inner_h * 0.18, y1: inner_h * 0.82 }
// RIGHT: one box per version plus the unstated bucket, heights proportional
// to the papers that name them. A version nobody named gets a floor so it is
// still visible and still labelled -- it exists whether or not it is cited.
const right_items = versions.map(v => ({
id: v.version, label: v.version,
papers: papers_of.get(v.version) ?? 0,
n_spectra: v.n_spectra
}))
if ((papers_of.get("__none__") ?? 0) > 0) {
right_items.push({ id: "__none__", label: "version not stated",
papers: papers_of.get("__none__"), unstated: true })
}
const FLOOR = 16
const weighted = d3.sum(right_items, d => d.papers)
const avail = inner_h - PAD * (right_items.length - 1)
const floor_total = right_items.filter(d => d.papers === 0).length * FLOOR
let cursor = 0
for (const r of right_items) {
r.h = r.papers === 0 ? FLOOR
: ((r.papers / weighted) * (avail - floor_total))
r.y0 = cursor; r.y1 = cursor + r.h; cursor += r.h + PAD
}
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width:100%; height:auto; background:#fbfbfd; font:11px sans-serif;")
const g = svg.append("g").attr("transform", `translate(${margin.left},${margin.top})`)
const ribbon = (x0, y0a, y0b, x1, y1a, y1b) => {
const xm = (x0 + x1) / 2
return `M${x0},${y0a}C${xm},${y0a} ${xm},${y1a} ${x1},${y1a}`
+ `L${x1},${y1b}C${xm},${y1b} ${xm},${y0b} ${x0},${y0b}Z`
}
// provenance ribbons: thin, dashed, uncounted
for (const l of left) {
const mid_y = mid.y0 + (mid.y1 - mid.y0) * ((l.y0 + l.y1) / 2) / inner_h
g.append("path")
.attr("d", ribbon(col_x[0] + NODE_W, l.y0 + 1, l.y1 - 1,
col_x[1], mid_y - 3, mid_y + 3))
.attr("fill", "#94a3b8").attr("opacity", 0.26)
}
// version ribbons: weighted by papers
let mid_off = 0
const mid_h = mid.y1 - mid.y0
for (const r of right_items) {
const share = weighted ? (r.papers / weighted) : 0
const h = Math.max(2, share * mid_h)
g.append("path")
.attr("d", ribbon(col_x[1] + NODE_W, mid.y0 + mid_off, mid.y0 + mid_off + h,
col_x[2], r.y0, r.y1))
.attr("fill", r.unstated ? "#d97706" : "#0f766e")
.attr("opacity", r.unstated ? 0.42 : 0.3)
mid_off += h
}
const box = (x, y0, y1, fill) => g.append("rect")
.attr("x", x).attr("y", y0).attr("width", NODE_W)
.attr("height", Math.max(2, y1 - y0)).attr("fill", fill).attr("rx", 2)
for (const l of left) {
box(col_x[0], l.y0, l.y1, "#94a3b8")
g.append("a").attr("href", l.url).attr("target", "_blank")
.append("text").attr("x", col_x[0] - 8).attr("y", (l.y0 + l.y1) / 2 + 3)
.attr("text-anchor", "end").attr("fill", "#24292f")
.text(`${l.label} · ${l.id}`)
}
box(col_x[1], mid.y0, mid.y1, "#475569")
g.append("text").attr("x", col_x[1] + NODE_W / 2).attr("y", mid.y0 - 10)
.attr("text-anchor", "middle").attr("font-weight", 600).attr("fill", "#24292f")
.text(mid.label)
g.append("text").attr("x", col_x[1] + NODE_W / 2).attr("y", mid.y1 + 16)
.attr("text-anchor", "middle").attr("fill", "#57606a").text(mid.sub)
for (const r of right_items) {
box(col_x[2], r.y0, r.y1, r.unstated ? "#d97706" : "#0f766e")
const t = g.append("text").attr("x", col_x[2] + NODE_W + 8)
.attr("y", (r.y0 + r.y1) / 2 - (r.n_spectra ? 3 : 0) + 3)
.attr("fill", "#24292f").attr("font-weight", r.unstated ? 600 : 500)
t.text(`${r.label} · ${r.papers} paper${r.papers === 1 ? "" : "s"}`)
if (r.n_spectra)
g.append("text").attr("x", col_x[2] + NODE_W + 8)
.attr("y", (r.y0 + r.y1) / 2 + 12).attr("fill", "#57606a")
.text(`${d3.format(",")(r.n_spectra)} spectra`)
}
for (const [x, label] of [[col_x[0], "nine third-party submissions"],
[col_x[1], "one curated deposit"],
[col_x[2], "five versions in use"]]) {
g.append("text").attr("x", x).attr("y", -26).attr("fill", "#57606a")
.attr("font-weight", 600).text(label)
}
return svg.node()
}The left-hand ribbons carry no count, and are drawn pale for that reason. The nine source submissions are unrelated proteomics studies, of honeybees and tomatoes and bottlenose dolphins, whose spectra were re-curated into one deposit; the catalog stores no per-species spectrum count, and sizing those ribbons would have invented one. The right-hand ribbons are weighted by papers. A version with no papers still gets a visible box, because it exists whether or not anyone cited it.
Deposits that travel together
Six of these have no shared provenance at all. They are unrelated studies, of dolphin tissue and bronchoalveolar lavage fluid and diabetic beta-cells, that nobody ever merged into one object. What they share is that the same papers keep reaching for all of them at once, which makes them a benchmark suite in practice and nothing in name. The catalog keeps them as six deposits for that reason, and this chart is where the convention becomes visible.
couse_graph = {
const by_dataset = d3.group(dataset_usage_t.filter(u => u.kind === "deposit"),
u => u.dataset)
const deposits = []
for (const [name, rows] of by_dataset) {
const papers = Array.from(new Set(rows.map(r => r.pub_id))).sort(d3.ascending)
if (papers.length >= couse_min_papers) deposits.push({ name, papers })
}
// A SIGNATURE is the exact set of papers that cite a deposit. Deposits with
// an identical signature are used by exactly the same work, which is the
// evidence that they form a suite; the largest such group is highlighted.
// Grouping on the signature rather than on pairwise overlap keeps the claim
// falsifiable: these are not "similar", they are identical.
for (const d of deposits) d.sig = d.papers.join(",")
const sig_groups = d3.rollups(deposits, v => v.length, d => d.sig)
.sort((a, b) => b[1] - a[1])
const block_sig = sig_groups.length && sig_groups[0][1] > 1 ? sig_groups[0][0] : null
// A deposit one paper short of the block is still part of the practice, so
// near-members are marked too: the same signature plus any one extra paper.
const block_papers = block_sig ? new Set(block_sig.split(",")) : new Set()
for (const d of deposits) {
d.in_block = d.sig === block_sig
d.near_block = !d.in_block && block_sig
&& block_papers.size > 0
&& [...block_papers].every(p => d.papers.map(String).includes(p))
}
const acc_of = new Map()
for (const a of dataset_addresses_t) if (!acc_of.has(a.dataset)) acc_of.set(a.dataset, a.accession)
const title_of = new Map(pubs_t.map(p => [p.id, p]))
const paper_deg = new Map()
for (const d of deposits) for (const p of d.papers) paper_deg.set(p, (paper_deg.get(p) ?? 0) + 1)
const papers = Array.from(paper_deg.entries())
.map(([id, deg]) => {
const p = title_of.get(id)
const first_model = (p?.models_described ?? "").split(",")[0].trim()
const raw = first_model || (p?.title ?? `publication ${id}`)
return { id, deg, title: p?.title ?? "", year: p?.year,
// 40 characters keeps the longest label inside the right margin
// at 11 px; measured, three ran past the frame at no limit.
label: raw.length > 40 ? raw.slice(0, 39) + "\u2026" : raw }
})
.sort((a, b) => b.deg - a.deg || d3.ascending(a.label, b.label))
// Two papers can share a leading method name -- a paper describing both
// PandaNovo and pi-HelixNovo leads with the same string as the pi-HelixNovo
// paper -- and two identical labels in one column are unreadable. Add the
// year only where it is needed to tell them apart.
{
const seen = d3.rollup(papers, v => v.length, d => d.label)
for (const p of papers) {
if (seen.get(p.label) > 1 && p.year) p.label = `${p.label} (${p.year})`
}
}
deposits.sort((a, b) =>
(b.in_block | 0) - (a.in_block | 0)
|| (b.near_block | 0) - (a.near_block | 0)
|| b.papers.length - a.papers.length
|| d3.ascending(a.name, b.name))
return { deposits, papers, acc_of, block_size: deposits.filter(d => d.in_block).length }
}html`<p style="margin:0.4rem 0 0.9rem; color:#57606a;">
<b>${couse_graph.deposits.length}</b> deposits are reused by
${couse_min_papers} or more papers. <b>${couse_graph.block_size}</b> of them are
cited by the <i>identical</i> set of
<b>${couse_graph.deposits.find(d => d.in_block)?.papers.length ?? 0}</b> papers,
drawn in teal below, with near-members in a lighter shade: not merely similar
usage, the same set exactly.</p>`couse_flow = {
const { deposits, papers, acc_of } = couse_graph
const width = 1100
const PAD = 7
const NODE_W = 13
const height = Math.max(520, 56 + papers.length * 25)
// right 330: the longest paper label, truncated to 40 characters plus its
// degree suffix, ran 3 labels past the frame at 250.
const margin = { top: 42, right: 330, bottom: 26, left: 300 }
const inner_w = width - margin.left - margin.right
const inner_h = height - margin.top - margin.bottom
// Both columns are weighted by edge count, and every edge weighs 1: a paper
// either used a deposit or did not, and there is no spectrum count here to
// scale by. So node height is degree, which is the honest encoding.
const total_edges = d3.sum(deposits, d => d.papers.length)
const lay = (nodes, degree) => {
const avail = inner_h - PAD * (nodes.length - 1)
let y = 0
for (const n of nodes) {
n.h = (degree(n) / total_edges) * avail
n.y0 = y; n.y1 = y + n.h; y += n.h + PAD
}
}
lay(deposits, d => d.papers.length)
lay(papers, p => p.deg)
const COL = { block: "#0f766e", near: "#5eada6", other: "#94a3b8" }
const colour = d => d.in_block ? COL.block : d.near_block ? COL.near : COL.other
const paper_by_id = new Map(papers.map(p => [p.id, p]))
const l_off = new Map(), r_off = new Map()
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width:100%; height:auto; background:#fbfbfd; font:11px sans-serif;")
const g = svg.append("g").attr("transform", `translate(${margin.left},${margin.top})`)
const ribbon = (x0, y0a, y0b, x1, y1a, y1b) => {
const xm = (x0 + x1) / 2
return `M${x0},${y0a}C${xm},${y0a} ${xm},${y1a} ${x1},${y1a}`
+ `L${x1},${y1b}C${xm},${y1b} ${xm},${y0b} ${x0},${y0b}Z`
}
for (const d of deposits) {
for (const pid of d.papers) {
const p = paper_by_id.get(pid)
if (!p) continue
const lh = d.h / d.papers.length
const rh = p.h / p.deg
const lo = l_off.get(d.name) ?? 0
const ro = r_off.get(p.id) ?? 0
g.append("path")
.attr("d", ribbon(NODE_W, d.y0 + lo, d.y0 + lo + lh,
inner_w, p.y0 + ro, p.y0 + ro + rh))
.attr("fill", colour(d))
.attr("opacity", d.in_block ? 0.34 : 0.17)
l_off.set(d.name, lo + lh)
r_off.set(p.id, ro + rh)
}
}
for (const d of deposits) {
g.append("rect").attr("x", 0).attr("y", d.y0).attr("width", NODE_W)
.attr("height", Math.max(2, d.h)).attr("fill", colour(d)).attr("rx", 2)
const acc = acc_of.get(d.name) ?? ""
const t = g.append("text").attr("x", -8).attr("y", (d.y0 + d.y1) / 2 - 2)
.attr("text-anchor", "end").attr("fill", "#24292f")
.attr("font-weight", d.in_block ? 600 : 400)
t.text(acc)
g.append("text").attr("x", -8).attr("y", (d.y0 + d.y1) / 2 + 11)
.attr("text-anchor", "end").attr("fill", "#57606a")
// 44 characters keeps the longest repository title inside the 300 px
// left margin at 11 px; the full name is on the paper pages.
.text(d.name.length > 44 ? d.name.slice(0, 43) + "…" : d.name)
}
for (const p of papers) {
g.append("rect").attr("x", inner_w).attr("y", p.y0).attr("width", NODE_W)
.attr("height", Math.max(2, p.h)).attr("fill", "#475569").attr("rx", 2)
g.append("text").attr("x", inner_w + NODE_W + 8).attr("y", (p.y0 + p.y1) / 2 + 3)
.attr("fill", "#24292f").attr("font-weight", p.deg >= 5 ? 600 : 400)
.text(`${p.label}${p.deg > 1 ? ` · ${p.deg}` : ""}`)
}
for (const [x, anchor, label] of [[0, "start", "deposit"],
[inner_w + NODE_W, "end", "papers reusing it"]]) {
g.append("text").attr("x", x).attr("y", -24).attr("text-anchor", anchor)
.attr("fill", "#57606a").attr("font-weight", 600).text(label)
}
return svg.node()
}Every ribbon here weighs the same, deliberately. A paper either used a deposit or it did not, and unlike the nine-species flow above there is no spectrum count to scale by, so node height is simply degree. The teal block is computed from an identical citing-paper signature rather than from pairwise similarity, which is what makes it a claim worth printing: these deposits are not used by overlapping sets of papers, they are used by the same set. The lighter shade is a deposit matching that signature plus one extra paper.
The dataset network
md`A dataset is in this graph once **two or more** papers link to it: a deposit
cited only by the study that made it says nothing about reuse, and there are
${dataset_versions_t.length - 0 > 0 ? (all_dataset_rows.filter(d => d.papers < 2).length) : 0}
of those. Diamonds are datasets, sized by how many papers use them and coloured
by kind; circles are papers, sized by how many of these datasets they draw on.`dataset_paper_network = {
// 1250, not the 1400 the other two force charts use. This graph has 176 nodes
// against their 173 and 360, but its links are dense enough that it settles
// into a tighter ball: at 1400 the nodes filled 69% of the frame width with
// nothing within 179 px of the right edge. Narrowing the frame to the content
// beats pushing the content at the frame, which is how nodes end up parked on
// the border.
const width = 1250
// 800, between the co-authorship chart's 650 for 173 nodes and the
// bipartite's 1000 for 360. This holds 176, and the density that pinned
// nodes to the frame in those two scales with node count, not with ambition.
const PLOT_H = 800
const TITLE_PX = 34
const LEGEND_H = 56
// Edges are unweighted on purpose: a paper either used a dataset or it did
// not, and there is no spectrum count to scale by. "Weighted by reuse" is
// therefore carried by NODE SIZE (degree), not by edge width.
const papers_of = d3.group(dataset_usage_t, u => u.dataset)
const ds_nodes = []
for (const [name, rows] of papers_of) {
const pubs = Array.from(new Set(rows.map(r => r.pub_id)))
if (pubs.length < 2) continue
ds_nodes.push({ id: `ds:${name}`, type: "dataset", name,
kind: rows[0].kind, pubs, degree: pubs.length })
}
const kept = new Set(ds_nodes.flatMap(d => d.pubs))
const pub_by_id = new Map(pubs_t.map(p => [p.id, p]))
const paper_nodes = Array.from(kept).map(id => {
const p = pub_by_id.get(id)
const model = (p?.models_described ?? "").split(",")[0].trim()
const raw = model || (p?.title ?? `publication ${id}`)
return { id: `p:${id}`, type: "paper", pub_id: id,
name: raw.length > 32 ? raw.slice(0, 31) + "…" : raw,
full: p?.title ?? "", degree: 0 }
})
const paper_by_key = new Map(paper_nodes.map(n => [n.id, n]))
const links = []
for (const d of ds_nodes) for (const pid of d.pubs) {
const t = paper_by_key.get(`p:${pid}`)
if (!t) continue
links.push({ source: d.id, target: t.id })
t.degree += 1
}
const nodes = [...ds_nodes, ...paper_nodes]
const max_deg = d3.max(nodes, d => d.degree) ?? 1
const NODE_R = d => d.type === "dataset"
? 5 + 11 * Math.sqrt(d.degree / max_deg)
: 3.5 + 7 * Math.sqrt(d.degree / max_deg)
const KIND_COLOUR = { benchmark: "#0f766e", training: "#b45309", deposit: "#64748b" }
const colour = d => d.type === "paper" ? "#cbd5e1" : (KIND_COLOUR[d.kind] ?? "#94a3b8")
// Cluster targets on a phyllotaxis disc, by kind for a dataset and by its
// most-used dataset's kind for a paper, so every node has a target. A RING of
// targets is what emptied the middle of the two older force charts; this is
// area-uniform instead. Strength is low (0.06) because the point is to
// separate the three populations loosely, not to draw three blobs.
const GOLDEN_ANGLE = Math.PI * (3 - Math.sqrt(5))
const kinds = Array.from(new Set(ds_nodes.map(d => d.kind))).sort()
// Recentred on their own centroid. A phyllotaxis walk of only three points
// is not symmetric about the origin, so the three targets drifted the whole
// network off-centre: measured, nodes spanned 36-902 px of a 1099 px frame,
// leaving 197 px empty on the right against 36 on the left. Subtracting the
// centroid costs nothing and keeps the disc's area-uniform spacing.
const raw_centers = kinds.map((k, i) => {
const r = Math.sqrt((i + 0.5) / kinds.length)
const a = i * GOLDEN_ANGLE
return [Math.cos(a) * r * width * 0.34, Math.sin(a) * r * PLOT_H * 0.26]
})
// kind_of FIRST: the weighted centroid below reads it, and `const` is not
// hoisted, so computing the centroid above this threw a ReferenceError.
const kind_of = new Map()
for (const d of ds_nodes) kind_of.set(d.id, d.kind)
for (const p of paper_nodes) {
const ks = ds_nodes.filter(d => d.pubs.includes(p.pub_id)).map(d => d.kind)
kind_of.set(p.id, d3.mode(ks) ?? kinds[0])
}
// WEIGHTED by group size, not a plain mean. The unweighted centroid moved the
// targets but not the network: 72 of the 83 datasets are deposits, so the
// layout's centre of mass sits at THAT group's target wherever the three
// targets happen to be, and the right-hand gap stayed at 191 px of 1099.
// Weighting puts the heaviest group near the middle, which is what centring
// a lopsided graph actually means.
const kind_weight = new Map(kinds.map(k =>
[k, nodes.filter(n => kind_of.get(n.id) === k).length || 1]))
const w_total = d3.sum(kinds, k => kind_weight.get(k))
const cx0 = d3.sum(kinds, (k, i) => raw_centers[i][0] * kind_weight.get(k)) / w_total
const cy0 = d3.sum(kinds, (k, i) => raw_centers[i][1] * kind_weight.get(k)) / w_total
const centers = new Map()
kinds.forEach((k, i) => {
centers.set(k, [width / 2 + raw_centers[i][0] - cx0,
(TITLE_PX + PLOT_H) / 2 + raw_centers[i][1] - cy0])
})
const target = d => centers.get(kind_of.get(d.id)) ?? [width / 2, PLOT_H / 2]
const sim = d3.forceSimulation(nodes)
// distance 78 and charge -240, not 58 and -170. At the tighter settings the
// layout CONVERGED but did not fill: 176 nodes occupied 663 px of a 1099 px
// frame, because the deposit cluster holds about 150 of them and a charge
// capped at 190 px cannot separate a ball that dense. Measured node spread
// is the number to tune against here, exactly as wall_probe.js was for the
// other two -- the label audit passes happily on a compact knot.
.force("link", d3.forceLink(links).id(d => d.id).distance(78).strength(0.2))
// Still capped, for the reason recorded on the two charts below: an
// unbounded forceManyBody grows its outward pressure with the square of the
// node count and the layout never converges inside the frame.
.force("charge", d3.forceManyBody().strength(-240).distanceMax(210))
.force("collide", d3.forceCollide().radius(d => NODE_R(d) + 3))
.force("center", d3.forceCenter(width / 2, (TITLE_PX + PLOT_H) / 2))
.force("kx", d3.forceX(d => target(d)[0]).strength(0.06))
.force("ky", d3.forceY(d => target(d)[1]).strength(0.06))
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, PLOT_H + LEGEND_H])
.attr("style", "max-width:100%; height:auto; font:11px sans-serif; cursor:grab;")
svg.append("text").attr("x", width / 2).attr("y", 22).attr("text-anchor", "middle")
.attr("font-size", 16).attr("font-weight", "bold").attr("fill", "#24292f")
.text(`Datasets and the papers reusing them · ${ds_nodes.length} datasets, `
+ `${paper_nodes.length} papers, ${links.length} links`)
const g = svg.append("g")
svg.call(d3.zoom().on("zoom", ev => g.attr("transform", ev.transform)))
const link = g.append("g").attr("stroke", "#94a3b8").attr("stroke-opacity", 0.3)
.selectAll("line").data(links).join("line").attr("stroke-width", 1)
const node_g = g.append("g").selectAll("g").data(nodes).join("g")
.call(d3.drag()
.on("start", (ev, d) => { if (!ev.active) sim.alphaTarget(0.25).restart(); d.fx = d.x; d.fy = d.y })
.on("drag", (ev, d) => { d.fx = ev.x; d.fy = ev.y })
.on("end", (ev, d) => { if (!ev.active) sim.alphaTarget(0); d.fx = null; d.fy = null }))
node_g.filter(d => d.type === "dataset").append("rect")
.attr("width", d => NODE_R(d) * 1.7).attr("height", d => NODE_R(d) * 1.7)
.attr("x", d => -NODE_R(d) * 0.85).attr("y", d => -NODE_R(d) * 0.85)
.attr("transform", "rotate(45)")
.attr("fill", colour).attr("stroke", "white").attr("stroke-width", 1)
node_g.filter(d => d.type === "paper").append("circle")
.attr("r", NODE_R).attr("fill", colour).attr("stroke", "white").attr("stroke-width", 1)
node_g.append("title").text(d => d.type === "dataset"
? `${d.name}\n${d.kind} · ${d.degree} papers`
: `${d.full || d.name}\n${d.degree} of these datasets`)
const label = g.append("g").selectAll("text").data(nodes).join("text")
.attr("font-size", 10)
.attr("font-weight", d => d.type === "dataset" ? 600 : 400)
.attr("fill", d => d.type === "dataset" ? "#24292f" : "#475569")
.attr("paint-order", "stroke").attr("stroke", "white").attr("stroke-width", 2.5)
.text(d => d.name)
// Same greedy pass as the other two force charts: six positions per name in
// descending degree, hidden if none is free. Every node still answers a
// hover, so the hidden ones are the least connected rather than lost.
const LABEL_CHAR_PX = 5.8
const SIDES = [[1, 0], [-1, 0], [1, -11], [-1, -11], [1, 11], [-1, 11]]
const label_order = nodes.slice().sort((a, b) => d3.descending(a.degree, b.degree))
const place_labels = () => {
const shown = []
for (const d of label_order) {
const w = d.name.length * LABEL_CHAR_PX
d.label_on = false
for (const [side, dy] of SIDES) {
const x0 = side > 0 ? d.x + NODE_R(d) + 3 : d.x - NODE_R(d) - 3 - w
const y = d.y + dy
const box = { x0, x1: x0 + w, y0: y - 6, y1: y + 6 }
if (box.x0 < 0 || box.x1 > width || box.y0 < TITLE_PX || box.y1 > PLOT_H) continue
const clash =
shown.some(o => box.x0 < o.x1 + 2 && box.x1 + 2 > o.x0 && box.y0 < o.y1 && box.y1 > o.y0) ||
nodes.some(n => box.x0 < n.x + NODE_R(n) && box.x1 > n.x - NODE_R(n)
&& box.y0 < n.y + NODE_R(n) && box.y1 > n.y - NODE_R(n))
if (!clash) { d.label_on = true; d.label_side = side; d.label_dy = dy; shown.push(box); break }
}
}
}
let tick_n = 0
sim.on("tick", () => {
for (const d of nodes) {
const r = NODE_R(d) + 2
d.x = Math.max(r, Math.min(width - r, d.x))
d.y = Math.max(TITLE_PX + r, Math.min(PLOT_H - r, d.y))
}
link.attr("x1", d => d.source.x).attr("y1", d => d.source.y)
.attr("x2", d => d.target.x).attr("y2", d => d.target.y)
node_g.attr("transform", d => `translate(${d.x},${d.y})`)
if ((tick_n++ & 3) === 0) place_labels()
label.attr("x", d => d.x).attr("y", d => d.y + (d.label_dy ?? 0))
.attr("display", d => d.label_on ? null : "none")
.attr("dx", d => (d.label_side ?? 1) > 0 ? NODE_R(d) + 3 : -(NODE_R(d) + 3))
.attr("text-anchor", d => (d.label_side ?? 1) > 0 ? "start" : "end")
})
const lg = svg.append("g").attr("transform", `translate(40,${PLOT_H + 20})`)
let lx = 0
for (const k of kinds) {
lg.append("rect").attr("x", lx).attr("y", -6).attr("width", 11).attr("height", 11)
.attr("transform", `rotate(45,${lx + 5.5},-0.5)`).attr("fill", KIND_COLOUR[k] ?? "#94a3b8")
lg.append("text").attr("x", lx + 16).attr("y", 3).attr("fill", "#24292f").text(k)
lx += 20 + k.length * 7
}
lg.append("circle").attr("cx", lx + 6).attr("cy", -1).attr("r", 5).attr("fill", "#cbd5e1")
lg.append("text").attr("x", lx + 16).attr("y", 3).attr("fill", "#24292f").text("paper")
return svg.node()
}Every dataset
all_dataset_rows = {
const papers_of = d3.group(dataset_usage_t, u => u.dataset)
const versions_of = d3.group(dataset_versions_t, v => v.dataset)
const addrs_of = d3.group(dataset_addresses_t, a => a.dataset)
const rows = []
for (const v of versions_of.keys()) {
const vs = versions_of.get(v) ?? []
const us = papers_of.get(v) ?? []
const ad = addrs_of.get(v) ?? []
rows.push({
Dataset: v,
Kind: vs[0]?.kind ?? "",
Versions: vs.length,
papers: new Set(us.map(u => u.pub_id)).size,
Deposited: us.filter(u => u.role === "introduces").length > 0 ? "yes" : "",
Spectra: d3.max(vs, x => x.n_spectra) ?? null,
// Every accession, so the search box reaches the long tail: typing
// PXD004424 or "zenodo" finds its dataset, which is the whole reason
// this table exists alongside the charts.
Accessions: ad.map(a => a.accession).join(" "),
Repositories: Array.from(new Set(ad.map(a => a.repository))).sort().join(", "),
url: ad.find(a => a.url)?.url ?? null
})
}
return rows.sort((a, b) => b.papers - a.papers || d3.ascending(a.Dataset, b.Dataset))
}Inputs.table(dataset_search, {
columns: ["Dataset", "Kind", "Versions", "papers", "Deposited", "Spectra",
"Repositories", "Accessions"],
header: { papers: "Papers", Accessions: "Accession(s)" },
sort: "papers", reverse: true,
rows: 18,
format: {
Dataset: (d, i, rows) => {
const row = (rows ?? [])[i]
const u = row?.url
return u ? html`<a href="${u}" target="_blank" title="${d}">${
d.length > 60 ? d.slice(0, 59) + "…" : d}</a>`
: (d.length > 60 ? d.slice(0, 59) + "…" : d)
},
Spectra: d => d == null ? "" : d3.format(",")(d),
Deposited: d => d ? "✓" : ""
},
width: { Dataset: 330, Kind: 80, Versions: 70, papers: 62, Deposited: 80,
Spectra: 100, Repositories: 120 }
})html`<p style="color:#57606a; margin-top:0.4rem;">
<b>${all_dataset_rows.length}</b> datasets,
<b>${d3.sum(all_dataset_rows, d => d.Versions)}</b> versions and
<b>${all_dataset_rows.filter(d => d.papers >= 2).length}</b> reused by two or
more papers. A tick under <i>Deposited</i> means a catalogued paper produced
this data rather than merely running on it.</p>`Application areas
De novo peptide sequencing gets picked up by a handful of distinct scientific communities, each with its own workflow conventions. Each dot below is one application-focused method or workflow, placed at its first publication date and stacked into a lane by sub-domain.
subdomain_order = [
"general-proteomics", "glycoproteomics", "plant-pathogen", "metaproteomics",
"wastewater-metaproteomics",
"immunopeptidomics", "antibodyomics", "venomics", "neuropeptidomics",
"bioactive-peptides", "wildlife-proteomics",
"food-authentication", "pathogen-identification", "forensics", "palaeoproteomics",
"astrobiology", "toxin-identification"
]
subdomain_label = new Map(transpose(subdomains).map(d => [d.name, d.label]))
// Lane colours must stay apart in CIE Lab, and ADJACENCY IN subdomain_order IS
// WHAT MATTERS: two similar colours far apart in the stack are tolerable, two
// similar colours in neighbouring lanes are not. Every adjacent pair here is now
// at least deltaE 20 (worst: neuropeptidomics against bioactive-peptides at
// 21.9). Two pairs used to fail that, both same-hue and side by side:
// antibodyomics against venomics at 16.0, two reds, and plant-pathogen against
// metaproteomics at 11.1, two greens. The global minimum is 12.7, between
// immunopeptidomics and food-authentication, which sit six lanes apart.
subdomain_color = ({
"general-proteomics": "#8a96a0",
"glycoproteomics": "#701a75",
"plant-pathogen": "#6f4e37",
"metaproteomics": "#1a7f37",
"wastewater-metaproteomics":"#218380",
"immunopeptidomics": "#bf8700",
"antibodyomics": "#bf3989",
"venomics": "#a40e26",
"neuropeptidomics": "#116329",
"bioactive-peptides": "#4e7a1a",
"wildlife-proteomics": "#0969da",
"food-authentication": "#d4a72c",
"pathogen-identification": "#0a3069",
"forensics": "#953800",
"palaeoproteomics": "#8250df",
"astrobiology": "#155e75",
"toxin-identification": "#4338ca"
})application_timeline = {
// Same swim-lane pattern as the architectures chart, keyed on subdomain
// instead of family. Lane order is THEMATIC, not chronological: general and
// environmental proteomics at the top, then immune and organismal domains,
// then the identification-driven ones, ending with the three where no
// reference proteome exists at all (forensics, palaeoproteomics,
// astrobiology). Lane heights scale with how crowded a lane is.
// Metadata (order/label/color/height) lives in the shared subdomain_* cells
// above so this chart and the Sankey diagram below stay in sync.
const rows = algorithms_t
.filter(a => a.kind === "downstream-application" && a.subdomain && a.first_pub)
.slice()
.sort((a, b) => a.first_pub - b.first_pub)
// A subdomain present in the data but absent from subdomain_order is
// APPENDED rather than dropped, so a new one always shows up somewhere even
// before anyone registers it. Dropping it would hide real rows silently,
// and dereferencing an unregistered lane is what used to crash this cell.
const extra = Array.from(new Set(rows.map(r => r.subdomain)))
.filter(sd => !subdomain_order.includes(sd))
.sort()
// Lane order is CHRONOLOGICAL by each subdomain's first appearance, the order
// "The long view" and the architectures swim lane both use, so all three
// timelines can be read against one another. subdomain_order survives as the
// registry of labels, colours and lane heights, and as the test for an
// unregistered subdomain; it is no longer the display order.
//
// Reordering was free on the constraint that matters: the CIE Lab rule below
// cares about ADJACENT pairs, and chronological order improves the worst
// adjacent distance from 21.9 to 38.4, comfortably past the 20 threshold.
//
// Computed over ALL entries rather than the filtered ones so a lane keeps its
// place as filters change, and latest-first because lanes stack upward from
// y = 0, which puts the earliest subdomain at the top to match the long view.
// A registered subdomain with no dated entry sorts as "latest", at the bottom.
const sd_first = d3.rollup(
algorithms_t.filter(a => a.subdomain && a.first_pub),
v => d3.min(v, a => a.first_pub),
a => a.subdomain
)
const lane_order = subdomain_order.concat(extra)
.sort((a, b) => d3.descending(sd_first.get(a) ?? "9999",
sd_first.get(b) ?? "9999")
|| d3.descending(a, b))
const lane_label = subdomain_label
const lane_color = subdomain_color
// Lane geometry is measured in PIXELS, the same first-fit row packing the
// architectures chart above uses and for the same reason: fixed per-lane
// heights plus a date-gap tier list collided, 7 label-label and 9 label-dot
// overlaps measured in the DOM. A lane is now exactly as tall as the rows its
// labels need, which is also why there is no subdomain_lane_height registry
// any more.
const ROW_PX = 43 // 24 px for up to two label lines, 3 gap, 16 dot
const DOT_IN_ROW_PX = 35 // dot centre from the row top, label in the 24 above
const CHAR_PX = 5.7 // mean advance of 10 px bold Source Sans Pro
const LABEL_GAP_PX = 8 // clear air between two labels sharing a row
const PLOT_WIDTH_PX = 925 // width 1200 minus marginLeft 215 and marginRight 60
// Wrap the display label onto two lines at a word boundary when it exceeds
// the max width (Plot renders \n as separate tspans). Falls back to the raw
// string if the label already fits or has no whitespace to break on.
const wrap_label = (s, max_chars = 22) => {
if (!s || s.length <= max_chars) return s
const words = s.split(' ')
if (words.length === 1) return s
let line1 = ''
let i = 0
while (i < words.length && (line1 ? line1.length + 1 + words[i].length : words[i].length) <= max_chars) {
line1 = line1 ? `${line1} ${words[i]}` : words[i]
i++
}
if (!line1) { line1 = words[0]; i = 1 }
const line2 = words.slice(i).join(' ')
return line2 ? `${line1}\n${line2}` : line1
}
const half_px = label => Math.max(9,
d3.max(String(label).split('\n'), l => l.length) * CHAR_PX / 2)
const labelled = rows.map(r => ({ ...r, label: wrap_label(r.model) }))
const max_half = d3.max(labelled, r => half_px(r.label)) ?? 9
// The x domain stays pinned to whole years, so the gridlines land on years and
// a quiet stretch reads as a real gap. How MANY years of padding is chosen by
// measurement rather than fixed at one: the widest label needs 63 px of half
// width and a year is only ~35 px wide here, so one year clipped it.
const y_lo = d3.min(rows, d => d.first_pub).getUTCFullYear()
const y_hi = d3.max(rows, d => d.first_pub).getUTCFullYear()
let pad_years = 1
const px_per_day_for = p => PLOT_WIDTH_PX /
((Date.UTC(y_hi + p, 0, 1) - Date.UTC(y_lo - p, 0, 1)) / 86400000)
while (pad_years < 5 && max_half > pad_years * 365 * px_per_day_for(pad_years)) pad_years++
const x_pad_min = new Date(Date.UTC(y_lo - pad_years, 0, 1))
const x_pad_max = new Date(Date.UTC(y_hi + pad_years, 0, 1))
const px_per_day = px_per_day_for(pad_years)
const x_px = d => (d - x_pad_min) / 86400000 * px_per_day
// Greedy first-fit row packing per lane: a row is reused only when the two
// labels cannot touch. Filled in date order, so an early paper is never
// pushed down by a later one.
const row_right = new Map()
const rows_used = new Map()
const packed = []
for (const row of labelled) {
const cx = x_px(row.first_pub)
const half = half_px(row.label)
let r = 0
for (;; r++) {
const key = `${row.subdomain}|${r}`
const right = row_right.get(key)
if (right === undefined || cx - half >= right + LABEL_GAP_PX) {
row_right.set(key, cx + half)
break
}
}
rows_used.set(row.subdomain, Math.max(rows_used.get(row.subdomain) ?? 0, r + 1))
packed.push({ ...row, row: r })
}
// lane_order is latest-first and y grows upward, so walking it in order puts
// the earliest subdomain at the top, matching "The long view".
let y_cursor = 0
const lanes = lane_order.map(sd => {
const h = Math.max(1, rows_used.get(sd) ?? 1) * ROW_PX
const lane = { subdomain: sd, y0: y_cursor, y1: y_cursor + h, center: y_cursor + h/2, height: h }
y_cursor += h
return lane
})
const total_height = y_cursor
const lane_index = new Map(lanes.map(l => [l.subdomain, l]))
// Row 0 is the TOP row of its lane (y grows upward), so within a lane the rows
// also read earliest-first.
const items = packed.map(d => {
const lane = lane_index.get(d.subdomain)
return { ...d, y: lane.y1 - d.row * ROW_PX - DOT_IN_ROW_PX }
})
const chart = Plot.plot({
width: 1200,
// total_height is in pixels, so 1 y unit = 1 px and the row packing above
// survives the scale unchanged.
height: Math.max(240, total_height + 60),
marginLeft: 215, // measured 2 px over
marginRight: 60, // measured 15 px over
marginTop: 20,
marginBottom: 40,
x: { type: "time", label: "First publication →", domain: [x_pad_min, x_pad_max], grid: true },
y: { domain: [0, total_height], axis: null },
color: { domain: lane_order, range: lane_order.map(s => lane_color[s] ?? "#888"), legend: false },
marks: [
Plot.rect(lanes, {
x1: () => x_pad_min, x2: () => x_pad_max,
y1: "y0", y2: "y1",
fill: "subdomain",
fillOpacity: 0.09
}),
Plot.ruleY(lanes.flatMap(l => [l.y0, l.y1]), { stroke: "#ddd", strokeWidth: 0.5 }),
Plot.text(lanes, {
x: () => x_pad_min, y: "center",
text: d => lane_label.get(d.subdomain) ?? d.subdomain,
textAnchor: "end", dx: -8,
fontSize: 12, fontWeight: "bold",
fill: "subdomain",
// Each lane label is now a link to that area's own page, which is the
// destination this axis never had.
href: d => page_href("subdomains", d.subdomain), target: "_self"
}),
Plot.dot(items, {
x: "first_pub", y: "y",
fill: "subdomain",
r: 7, stroke: "white", strokeWidth: 1.5,
href: d => algo_href(d.base_model ?? d.model),
target: "_self"
}),
// One label per dot, always directly above it: the row packing gives each
// its own 24 px, so the old above-or-below split is no longer needed to
// buy vertical space. Names over ~22 chars wrap onto two lines; the full
// name stays in the tooltip.
Plot.text(items, {
x: "first_pub", y: "y", text: "label",
textAnchor: "middle", lineAnchor: "bottom", dy: -11,
fontSize: 10, fontWeight: "bold", fill: "subdomain"
})
]
})
// Rich .map-tooltip on the Plot.dot circles (matches the shared style).
// Plot renders dots in the order of `items`, so bind by DOM order.
const app_tooltip = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_app_tip = ev => {
const pad = 12
const box = app_tooltip.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
app_tooltip.style.left = `${Math.max(pad, left)}px`
app_tooltip.style.top = `${Math.max(pad, top)}px`
}
// Same selector fix as the architectures chart — Plot uses aria-label='dot'.
const app_dot_circles = chart.querySelectorAll('g[aria-label="dot"] circle')
app_dot_circles.forEach((circle, i) => {
const d = items[i]
if (!d) return
circle.style.cursor = "pointer"
circle.setAttribute("aria-label", `${d.model}. ${lane_label.get(d.subdomain) ?? d.subdomain}. First publication ${d.first_pub.toISOString().slice(0,10)}.`)
circle.addEventListener("mouseenter", ev => {
app_tooltip.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.model}</div>
<div class="map-tooltip-meta">${d.first_pub.toISOString().slice(0,10)}</div>
<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${lane_color[d.subdomain] ?? "#999"}`}></span>
<span>${lane_label.get(d.subdomain) ?? d.subdomain}</span>
</div>
${d.description ? html`<div class="map-tooltip-meta" style="margin-top:6px; max-width:280px; white-space:normal;">${d.description}</div>` : ""}
</div>`)
app_tooltip.classList.add("visible")
app_tooltip.setAttribute("aria-hidden", "false")
move_app_tip(ev)
})
circle.addEventListener("mousemove", move_app_tip)
circle.addEventListener("mouseleave", () => {
app_tooltip.classList.remove("visible")
app_tooltip.setAttribute("aria-hidden", "true")
})
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${chart}</div>
${app_tooltip}
</div>`
}Application → sequencer flow
Which sequencing tools does each application community actually reach for? Two signals can drive the Sankey and they’re different strengths:
- Uses: the paper has a curated
publication_algorithmlink to the tool — usually because the methods section names it explicitly. These are the links whoseroleisuses, plus the rare case of an application paper that also introduces a tool. - Cites: the paper’s Crossref reference list contains a publication describing the tool — a weaker signal (the tool might just be mentioned as prior art). Rebuilt monthly by
build_citations.py.
Use the toggle to switch between them, or view the union of both. An edge from metaproteomics to PEAKS means at least one metaproteomics-tagged paper satisfies the selected signal. Edge thickness is the paper count.
application_sankey = {
// Which sequencer(s) does each downstream-application paper cite?
const alg_by_name = new Map(algorithms_t.map(a => [a.model, a]))
const pub_by_id = new Map(pubs_t.map(p => [p.id, p]))
// For each pub, gather its algorithm rows (parsed from the comma-joined
// 'models' string).
const pub_algs = new Map()
// And separately the methods each paper DESCRIBES, for the citation signal
// below: citing a venomics paper that happened to run PEAKS is not evidence
// that the citing paper uses PEAKS, but citing the PEAKS paper is.
const pub_described = new Map()
for (const p of pubs_t) {
const resolve = field => (field ?? "").split(",").map(s => s.trim())
.filter(Boolean).map(n => alg_by_name.get(n)).filter(Boolean)
pub_algs.set(p.id, resolve(p.models))
pub_described.set(p.id, resolve(p.models_described))
}
// For each downstream-application publication, resolve its subdomain via
// any linked kind='downstream-application' algorithm that has a subdomain.
const pub_subdomain = new Map()
for (const p of pubs_t) {
const app_alg = (pub_algs.get(p.id) || []).find(a => a.kind === "downstream-application" && a.subdomain)
if (app_alg) pub_subdomain.set(p.id, app_alg.subdomain)
}
// Aggregate: (subdomain, sequencer) → number of downstream-app papers.
// TWO signals are unioned so a paper counts once per sequencer regardless
// of which surfaced the edge:
// 1. Direct publication_algorithm links (curator-asserted "uses tool X").
// 2. Citation edges to papers describing tool X (via publication_citation).
// The seen_edges dedup key is (paper, subdomain, tool), so a paper that both
// uses and cites the same tool contributes just once to the flow count.
const flow = new Map()
const sd_totals = new Map()
const seq_totals = new Map()
const seen_edges = new Set()
const record = (paper_id, sd, a) => {
if (!(a.kind === "algorithm" || a.kind === "adjacent") || !a.family) return
const edge_key = `${paper_id}|${sd}|${a.model}`
if (seen_edges.has(edge_key)) return
seen_edges.add(edge_key)
const key = `${sd}|${a.model}`
flow.set(key, (flow.get(key) ?? 0) + 1)
sd_totals.set(sd, (sd_totals.get(sd) ?? 0) + 1)
seq_totals.set(a.model,(seq_totals.get(a.model)?? 0) + 1)
}
// Signal 1: curated tool-usage. Included when the radio is 'Uses' or 'Either'.
if (sankey_signal !== "Cites") {
for (const p of pubs_t) {
const sd = pub_subdomain.get(p.id)
if (!sd) continue
for (const a of pub_algs.get(p.id) || []) record(p.id, sd, a)
}
}
// Signal 2: intra-catalog citation graph. Included when the radio is 'Cites' or 'Either'.
if (sankey_signal !== "Uses") {
for (const e of citations_t) {
const sd = pub_subdomain.get(e.citing_id)
if (!sd) continue
for (const a of pub_described.get(e.cited_id) || []) record(e.citing_id, sd, a)
}
}
if (flow.size === 0) {
return html`<p style="color:#57606a; font-style:italic; padding:1rem;">
No subdomain→sequencer flows for signal <b>${sankey_signal}</b>. Try
switching the radio above to <i>Either</i> to combine curated tool-usage
with the citation graph.
</p>`
}
// Layout parameters.
//
// HEIGHT IS DERIVED, NOT FIXED. Node heights are proportional to traffic, so
// the many single-paper sequencers on the right collapse to slivers: at the
// old fixed 480px the six sequencers with one paper each got about 13px
// against an 11px font, so their labels all but touched. Sizing from the
// busier side keeps the proportions honest -- a minimum-node-height floor
// would lie about the flows -- and simply gives them room. MIN_NODE_PX is the
// vertical budget per node; 34 leaves an 11px label clear of its neighbours.
// The chart now grows as sequencers are added instead of squeezing them.
// HEIGHT SCALES WITH THE NUMBER OF NODES ON THE BUSIER SIDE. The default
// signal is "Either", which unions curated tool usage with the citation
// graph and puts 36 sequencers on the right; at the old fixed 480px their
// labels shared about 430px of inner height, roughly 12px each including
// padding, which is where the right-hand crowding came from. "Uses" alone has
// only 10 and was never the problem.
//
// Sizing by each node's SHARE of the flow was tried and is wrong: shares are
// so skewed (the smallest is 1/109 of the total under "Either") that
// guaranteeing a label for the thinnest ribbon would demand a 2700px chart.
// Node count is the constraint that actually binds, and the thinnest ribbons
// stay thin, which is honest about how little traffic they carry.
const PAD_Y = 8
const ROW_PX = 24
const n_rows = Math.max(sd_totals.size, seq_totals.size)
const width = 1100
const height = Math.min(1200, Math.max(480, 50 + n_rows * ROW_PX))
// right 285: the longest sequencer name, 'PSD MALDI high-sensitivity de novo
// sequencing', ran 6 px past the frame at 260.
const margin = { top: 30, right: 285, bottom: 20, left: 260 }
const inner_w = width - margin.left - margin.right
const inner_h = height - margin.top - margin.bottom
// Node lists sorted by traffic.
const left_nodes = Array.from(sd_totals.entries())
.sort((a, b) => b[1] - a[1])
.map(([sd, total]) => ({ id: sd, kind: "subdomain", total }))
const right_nodes = Array.from(seq_totals.entries())
.sort((a, b) => b[1] - a[1])
.map(([seq, total]) => ({ id: seq, kind: "sequencer", total }))
// Palette + labels come from the shared subdomain_color / subdomain_label
// cells defined above (next to the timeline chart) so the two charts can't
// drift out of sync.
// Assign Y positions (stacked, proportional to traffic + a small padding).
const pad_y = PAD_Y
const total_left = d3.sum(left_nodes, d => d.total)
const total_right = d3.sum(right_nodes, d => d.total)
const avail_left = inner_h - pad_y * (left_nodes.length - 1)
const avail_right = inner_h - pad_y * (right_nodes.length - 1)
let y_cursor = 0
for (const n of left_nodes) {
n.h = (n.total / total_left) * avail_left
n.y0 = y_cursor
n.y1 = y_cursor + n.h
y_cursor += n.h + pad_y
}
y_cursor = 0
for (const n of right_nodes) {
n.h = (n.total / total_right) * avail_right
n.y0 = y_cursor
n.y1 = y_cursor + n.h
y_cursor += n.h + pad_y
}
const left_by_id = new Map(left_nodes.map(n => [n.id, n]))
const right_by_id = new Map(right_nodes.map(n => [n.id, n]))
// Build edge list with y-positions inside each node (proportional stacks).
const edges = []
const left_offset = new Map(left_nodes.map(n => [n.id, 0]))
const right_offset = new Map(right_nodes.map(n => [n.id, 0]))
for (const [key, count] of Array.from(flow.entries())
.sort((a, b) => b[1] - a[1])) {
const [sd, seq] = key.split("|")
const L = left_by_id.get(sd), R = right_by_id.get(seq)
const l_h = (count / L.total) * L.h
const r_h = (count / R.total) * R.h
const l_off = left_offset.get(sd)
const r_off = right_offset.get(seq)
edges.push({
subdomain: sd, sequencer: seq, count,
l_y0: L.y0 + l_off, l_y1: L.y0 + l_off + l_h,
r_y0: R.y0 + r_off, r_y1: R.y0 + r_off + r_h,
color: subdomain_color[sd] ?? "#888"
})
left_offset.set(sd, l_off + l_h)
right_offset.set(seq, r_off + r_h)
}
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width:100%; height:auto; background:#fbfbfd; font:11px sans-serif;")
const g = svg.append("g").attr("transform", `translate(${margin.left}, ${margin.top})`)
// Ribbon paths (S-curves between left and right node segments).
const ribbon = d => {
const x0 = 0, x1 = inner_w
const mid = (x0 + x1) / 2
return `M ${x0} ${d.l_y0}
C ${mid} ${d.l_y0}, ${mid} ${d.r_y0}, ${x1} ${d.r_y0}
L ${x1} ${d.r_y1}
C ${mid} ${d.r_y1}, ${mid} ${d.l_y1}, ${x0} ${d.l_y1} Z`
}
// Styled tooltip for the ribbons (matches world-map / co-auth / bipartite).
// aria-label carries the same info for keyboard / screen-reader users.
const sankey_tooltip = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_sankey_tip = ev => {
const pad = 12
const box = sankey_tooltip.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
sankey_tooltip.style.left = `${Math.max(pad, left)}px`
sankey_tooltip.style.top = `${Math.max(pad, top)}px`
}
const ribbon_paths = g.append("g")
.selectAll("path")
.data(edges)
.join("path")
.attr("d", ribbon)
.attr("fill", d => d.color)
.attr("fill-opacity", 0.55)
.attr("stroke", "none")
.attr("cursor", "pointer")
.attr("aria-label", d => `${subdomain_label.get(d.subdomain) ?? d.subdomain} to ${d.sequencer}. ${d.count} paper${d.count === 1 ? "" : "s"}.`)
.on("mouseenter", (ev, d) => {
sankey_tooltip.replaceChildren(html`<div>
<div class="map-tooltip-title">${subdomain_label.get(d.subdomain) ?? d.subdomain} → ${d.sequencer}</div>
<div class="map-tooltip-meta">${d.count} paper${d.count === 1 ? "" : "s"}</div>
<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${d.color}`}></span>
<span>${subdomain_label.get(d.subdomain) ?? d.subdomain}</span>
</div>
</div>`)
sankey_tooltip.classList.add("visible")
sankey_tooltip.setAttribute("aria-hidden", "false")
move_sankey_tip(ev)
})
.on("mousemove", move_sankey_tip)
.on("mouseleave", () => {
sankey_tooltip.classList.remove("visible")
sankey_tooltip.setAttribute("aria-hidden", "true")
})
// Left node bars + labels.
const left_rects = g.append("g")
.selectAll("rect")
.data(left_nodes)
.join("rect")
.attr("x", -8).attr("y", d => d.y0)
.attr("width", 8).attr("height", d => d.y1 - d.y0)
.attr("fill", d => subdomain_color[d.id] ?? "#888")
const left_labels = g.append("g")
.selectAll("text")
.data(left_nodes)
.join("text")
.attr("x", -14).attr("y", d => (d.y0 + d.y1) / 2 + 4)
.attr("text-anchor", "end")
.attr("font-weight", "bold")
.attr("fill", d => subdomain_color[d.id] ?? "#333")
.text(d => `${subdomain_label.get(d.id) ?? d.id} (${d.total})`)
// Right node bars + labels.
const right_rects = g.append("g")
.selectAll("rect")
.data(right_nodes)
.join("rect")
.attr("x", inner_w).attr("y", d => d.y0)
.attr("width", 8).attr("height", d => d.y1 - d.y0)
.attr("fill", "#1f6feb")
const right_labels = g.append("g")
.selectAll("text")
.data(right_nodes)
.join("text")
.attr("x", inner_w + 14).attr("y", d => (d.y0 + d.y1) / 2 + 4)
.attr("text-anchor", "start")
.attr("fill", "#1d2330")
.text(d => `${d.id} (${d.total})`)
// Both ends link now. The right-hand nodes are sequencers and go to their
// algorithm page; the left-hand nodes are application areas and go to theirs,
// which they did not have when this chart was written.
svg_link_wrap(right_rects, d => algo_href(d.id))
svg_link_wrap(right_labels, d => algo_href(d.id))
svg_link_wrap(left_rects, d => page_href("subdomains", d.id))
svg_link_wrap(left_labels, d => page_href("subdomains", d.id))
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${svg.node()}</div>
${sankey_tooltip}
</div>`
}Where the work happens
Authors and the organizations behind them span countries. Pan and zoom the map. Zoomed out, nearby work collapses to country-level pies; zoom in past about 2× and those aggregates split into city-level pies. Circle area scales with the selected impact metric; pie colours group organizations by type. Organization type is inferred from institution names and should be read as a practical display category.
where_the_work_map = {
// ===== Aggregate from the currently-filtered pubs at organization level =====
const pub_ids = new Set(pubs_filtered.map(p => p.id))
const org_type_color = {
"Academic": "#2563eb",
"Industry": "#ea580c",
"Research Institute": "#7c3aed",
"Government": "#059669",
"Nonprofit": "#be123c",
"Lab": "#0f766e"
}
const classify_org = name => {
const n = (name ?? "").toLowerCase()
if (/\b(inc|ltd|llc|gmbh|corp|corporation|company|co\.|biosystems|biotech|pharmaceuticals|pfizer|bruker|procter|protein metrics|instadeep|novor|sciex|bioinformatics solutions|dp technology|deepmind|microsoft|nvidia|google)\b/.test(n)) return "Industry"
if (/\b(national laboratory|government|ministry|academy of sciences|nih|national institute|cnrs|csiro|bam|federal institute|nist|institute of metrology|los alamos)\b/.test(n)) return "Government"
if (/\b(nonprofit|foundation|initiative|biohub|vib|wellcome|sanger)\b/.test(n)) return "Nonprofit"
if (/\b(laboratory|lab)\b/.test(n) && !/\b(university|institute|hospital|center|centre)\b/.test(n)) return "Lab"
if (/\b(institute|institut|instituto|center|centre|laboratory|hospital|clinic|medical center|research)\b/.test(n) && !/\b(university|college|school|faculty)\b/.test(n)) return "Research Institute"
return "Academic"
}
const org_rows = new Map()
for (const r of pub_authorship_t) {
if (!pub_ids.has(r.pub_id) || !r.affiliation || !r.country || r.lat == null || r.lng == null) continue
const k = r.affiliation
if (!org_rows.has(k)) {
org_rows.set(k, {
organization: r.affiliation,
type: classify_org(r.affiliation),
city: r.city,
country: r.country,
lat_sum: 0,
lng_sum: 0,
loc_w: 0,
authors: new Set(),
papers: new Set(),
affiliations: new Set()
})
}
const org = org_rows.get(k)
org.authors.add(r.author_id)
org.papers.add(r.pub_id)
org.affiliations.add(r.affiliation_id)
org.lat_sum += r.lat
org.lng_sum += r.lng
org.loc_w += 1
}
const org_data = Array.from(org_rows.values()).map(r => ({
organization: r.organization,
type: r.type,
city: r.city,
country: r.country,
lat: r.lat_sum / r.loc_w,
lng: r.lng_sum / r.loc_w,
authors: r.authors.size,
papers: r.papers.size,
affiliations: r.affiliations.size,
citations: Array.from(r.papers).reduce((sum, pid) => sum + (publication_impact_by_pub.get(pid)?.cited_by_count ?? 0), 0)
})).filter(r => org_type_filter.includes(r.type))
const type_counts = org_types_present.map(type => ({
type,
n: org_data.filter(d => d.type === type).length
}))
const make_bucket = fields => ({
...fields,
lat_sum: 0,
lng_sum: 0,
loc_w: 0,
authors: new Set(),
papers: new Set(),
affiliations: new Set(),
organizations: new Set(),
types: new Map()
})
const add_to_bucket = (bucket, r, type) => {
bucket.authors.add(r.author_id)
bucket.papers.add(r.pub_id)
bucket.affiliations.add(r.affiliation_id)
bucket.organizations.add(r.affiliation)
bucket.lat_sum += r.lat
bucket.lng_sum += r.lng
bucket.loc_w += 1
if (!bucket.types.has(type)) {
bucket.types.set(type, {
type,
authors: new Set(),
papers: new Set(),
affiliations: new Set(),
organizations: new Set()
})
}
const by_type = bucket.types.get(type)
by_type.authors.add(r.author_id)
by_type.papers.add(r.pub_id)
by_type.affiliations.add(r.affiliation_id)
by_type.organizations.add(r.affiliation)
}
const city_rows = new Map()
const country_rows = new Map()
for (const r of pub_authorship_t) {
if (!pub_ids.has(r.pub_id) || !r.affiliation || !r.country || r.lat == null || r.lng == null) continue
const type = classify_org(r.affiliation)
if (!org_type_filter.includes(type)) continue
const country_key = r.country
if (!country_rows.has(country_key)) {
country_rows.set(country_key, make_bucket({
level: "country",
label: r.country,
country: r.country
}))
}
add_to_bucket(country_rows.get(country_key), r, type)
if (!r.city_id) continue
const city_key = String(r.city_id)
if (!city_rows.has(city_key)) {
city_rows.set(city_key, make_bucket({
level: "city",
label: `${r.city}, ${r.country}`,
city: r.city,
country: r.country,
lat: r.lat,
lng: r.lng
}))
}
add_to_bucket(city_rows.get(city_key), r, type)
}
const metric_value = (sets, metric) => {
if (!sets) return 0
if (metric === "authors") return sets.authors.size
if (metric === "papers") return sets.papers.size
return Array.from(sets.papers).reduce((sum, pid) => sum + (publication_impact_by_pub.get(pid)?.cited_by_count ?? 0), 0)
}
const finalize_bucket = bucket => {
const lat = bucket.lat ?? (bucket.lat_sum / bucket.loc_w)
const lng = bucket.lng ?? (bucket.lng_sum / bucket.loc_w)
const type_values = org_types_present.map(type => ({
type,
value: metric_value(bucket.types.get(type), org_metric)
})).filter(d => d.value > 0)
return {
level: bucket.level,
label: bucket.label,
city: bucket.city,
country: bucket.country,
lat,
lng,
authors: bucket.authors.size,
papers: bucket.papers.size,
affiliations: bucket.affiliations.size,
organizations: bucket.organizations.size,
citations: Array.from(bucket.papers).reduce((sum, pid) => sum + (publication_impact_by_pub.get(pid)?.cited_by_count ?? 0), 0),
type_values
}
}
const city_data = Array.from(city_rows.values())
.filter(d => d.loc_w > 0)
.map(finalize_bucket)
.filter(d => Number.isFinite(d.lat) && Number.isFinite(d.lng))
const country_data = Array.from(country_rows.values())
.filter(d => d.loc_w > 0)
.map(finalize_bucket)
.filter(d => Number.isFinite(d.lat) && Number.isFinite(d.lng))
// ===== Set up scales =====
const width = 1100, height = 620
const projection = d3.geoNaturalEarth1()
.fitExtent([[10, 10], [width - 10, height - 10]],
topojson.feature(world_atlas, world_atlas.objects.countries))
const path = d3.geoPath(projection)
// Size: sqrt(metric) so circle area is linear in the selected metric.
const metric_label = { citations: "global citations", authors: "authors", papers: "papers" }[org_metric]
const max_city = d3.max(city_data, d => d[org_metric]) || 1
const max_country = d3.max(country_data, d => d[org_metric]) || 1
const cityR = v => 2.5 + 13 * Math.sqrt((v || 0) / max_city)
const countryR = v => 7 + 25 * Math.sqrt((v || 0) / max_country)
const sorted_city_data = city_data.slice().sort((a, b) => d3.descending(a[org_metric], b[org_metric]))
const sorted_country_data = country_data.slice().sort((a, b) => d3.descending(a[org_metric], b[org_metric]))
// ===== Build SVG =====
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width:100%; height:auto; background:#f6f8fa; font:11px sans-serif; cursor: grab;")
svg.append("rect")
.attr("width", width)
.attr("height", height)
.attr("fill", "transparent")
.attr("pointer-events", "all")
const map_layer = svg.append("g")
// Country fills (greyed base map).
const countries_feature = topojson.feature(world_atlas, world_atlas.objects.countries)
const base_layer = map_layer.append("g").attr("class", "base-map")
base_layer
.selectAll("path")
.data(countries_feature.features)
.join("path")
.attr("d", path)
.attr("fill", "#e9ecef")
.attr("stroke", "#cfd4d9")
.attr("stroke-width", 0.5)
const type_breakdown = d => d.type_values
.slice()
.sort((a, b) => d3.descending(a.value, b.value))
// Drop the metric label from per-row breakdowns: the chart's metric
// selector already says which metric is being displayed, and the value
// stands alone.
.map(v => `${v.type}: ${v.value}`)
.join("\n")
const tooltip = d => `${d.label}
${d.authors} authors · ${d.papers} papers · ${d.citations} global citations
${d.organizations} organizations · ${d.affiliations} affiliation rows
${type_breakdown(d)}`
const tooltip_el = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const fmt = v => Number(v ?? 0).toLocaleString()
const move_tooltip = ev => {
const pad = 16
const { innerWidth, innerHeight } = window
const box = tooltip_el.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, innerHeight - box.height - pad)
tooltip_el.style.left = `${Math.max(pad, left)}px`
tooltip_el.style.top = `${Math.max(pad, top)}px`
}
const show_tooltip = (ev, d) => {
const rows = d.type_values
.slice()
.sort((a, b) => d3.descending(a.value, b.value))
// 'global citations' used to repeat in (a) the headline metric line,
// (b) the canonical triple, and (c) every per-type breakdown row. Fold
// (a) and (c) so the phrase only appears once in (b).
tooltip_el.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.label}</div>
<div class="map-tooltip-meta">
${fmt(d.authors)} authors · ${fmt(d.papers)} papers · ${fmt(d.citations)} global citations<br>
${fmt(d.organizations)} organizations · ${fmt(d.affiliations)} affiliation rows
</div>
${rows.map(row => html`<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${org_type_color[row.type]}`}></span>
<span>${row.type}: ${fmt(row.value)}</span>
</div>`)}
</div>`)
tooltip_el.classList.add("visible")
tooltip_el.setAttribute("aria-hidden", "false")
move_tooltip(ev)
}
const hide_tooltip = () => {
tooltip_el.classList.remove("visible")
tooltip_el.setAttribute("aria-hidden", "true")
}
const marker_radius = (radius_fn, d, k) => radius_fn(d[org_metric]) / Math.sqrt(k)
const draw_markers = (layer, data, radius_fn, k = 1) => {
const markers = layer.selectAll("g.marker")
.data(data, d => d.label)
.join("g")
.attr("class", "marker")
.attr("transform", d => {
const [x, y] = projection([d.lng, d.lat])
return `translate(${x},${y})`
})
.style("cursor", "help")
.on("mouseenter", show_tooltip)
.on("mousemove", move_tooltip)
.on("mouseleave", hide_tooltip)
.on("focus", function(ev, d) {
const rect = this.getBoundingClientRect()
show_tooltip({ clientX: rect.left + rect.width / 2, clientY: rect.top + rect.height / 2 }, d)
})
.on("blur", hide_tooltip)
markers.each(function(d) {
const g = d3.select(this)
const r = marker_radius(radius_fn, d, k)
const slices = d.type_values.length ? d.type_values : [{ type: "Academic", value: 1 }]
const pie = d3.pie().sort(null).value(s => s.value)(slices)
const arc = d3.arc().innerRadius(0).outerRadius(r)
g.selectAll("path.slice")
.data(pie, p => p.data.type)
.join("path")
.attr("class", "slice")
.attr("d", arc)
.attr("fill", p => org_type_color[p.data.type])
.attr("fill-opacity", 0.82)
g.selectAll("circle.outline")
.data([d])
.join("circle")
.attr("class", "outline")
.attr("r", r)
.attr("fill", "none")
.attr("stroke", "#1f2328")
.attr("stroke-width", 0.65 / Math.sqrt(k))
g.selectAll("title").remove()
g.attr("aria-label", tooltip(d))
})
}
const country_layer = map_layer.append("g").attr("class", "country-layer")
const city_layer = map_layer.append("g").attr("class", "city-layer").attr("opacity", 0).attr("pointer-events", "none")
draw_markers(country_layer, sorted_country_data, countryR)
draw_markers(city_layer, sorted_city_data, cityR)
const SPLIT_K = 2.0
const zoom = d3.zoom()
.scaleExtent([1, 8])
.translateExtent([[0, 0], [width, height]])
.extent([[0, 0], [width, height]])
.on("start", () => svg.style("cursor", "grabbing"))
.on("zoom", ev => {
const k = ev.transform.k
const show_cities = k >= SPLIT_K
map_layer.attr("transform", ev.transform)
base_layer.selectAll("path").attr("stroke-width", 0.5 / Math.sqrt(k))
country_layer.attr("opacity", show_cities ? 0 : 1).attr("pointer-events", show_cities ? "none" : "auto")
city_layer.attr("opacity", show_cities ? 1 : 0).attr("pointer-events", show_cities ? "auto" : "none")
draw_markers(country_layer, sorted_country_data, countryR, k)
draw_markers(city_layer, sorted_city_data, cityR, k)
})
.on("end", () => svg.style("cursor", "grab"))
svg.call(zoom)
// ===== Legends =====
const legend = svg.append("g").attr("transform", `translate(20, ${height - 74})`)
legend.append("text").attr("x", 0).attr("y", -10)
.attr("font-weight", "bold").attr("fill", "#2a3140")
.text(`${org_data.length} organizations · ${city_data.length} cities · circle area ∝ ${metric_label}`)
const visible_types = org_types_present.filter(t => type_counts.find(d => d.type === t)?.n)
const leg = legend.selectAll("g.type")
.data(visible_types)
.join("g")
.attr("transform", (_, i) => `translate(${(i % 3) * 210}, ${Math.floor(i / 3) * 24})`)
leg.append("circle")
.attr("r", 7)
.attr("cx", 7)
.attr("cy", 7)
.attr("fill", d => org_type_color[d])
.attr("fill-opacity", 0.78)
.attr("stroke", "#1f2328")
.attr("stroke-width", 0.5)
leg.append("text")
.attr("x", 20)
.attr("y", 11)
.attr("fill", "#2a3140")
.text(d => `${d} (${type_counts.find(x => x.type === d)?.n ?? 0})`)
const max_value = max_country
const max_r = countryR(max_value)
// width - 210, not - 170: the longest legend line is '11229 global citations'
// at ~121 px from its x of 70, which ran 6 px past the frame.
const size_legend = svg.append("g").attr("transform", `translate(${width - 210}, ${height - 86})`)
size_legend.append("circle").attr("cx", 28).attr("cy", 36).attr("r", max_r).attr("fill", "none").attr("stroke", "#57606a")
size_legend.append("circle").attr("cx", 28).attr("cy", 36 + max_r - countryR(max_value / 4)).attr("r", countryR(max_value / 4)).attr("fill", "none").attr("stroke", "#57606a")
size_legend.append("text").attr("x", 70).attr("y", 24).attr("fill", "#57606a").text(`${max_value} ${metric_label}`)
size_legend.append("text").attr("x", 70).attr("y", 48).attr("fill", "#57606a").text(`${Math.round(max_value / 4)} ${metric_label}`)
const zoom_in = html`<button class="map-zoom-btn" title="Zoom in" aria-label="Zoom in">+</button>`
zoom_in.onclick = () => svg.transition().duration(180).call(zoom.scaleBy, 1.45)
const zoom_out = html`<button class="map-zoom-btn" title="Zoom out" aria-label="Zoom out">-</button>`
zoom_out.onclick = () => svg.transition().duration(180).call(zoom.scaleBy, 1 / 1.45)
const zoom_reset = html`<button class="map-zoom-btn" title="Reset map zoom" aria-label="Reset map zoom">↺</button>`
zoom_reset.onclick = () => svg.transition().duration(220).call(zoom.transform, d3.zoomIdentity)
return html`<div class="chart-wrap">
<div class="map-zoom-controls">${zoom_in}${zoom_out}${zoom_reset}</div>
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${svg.node()}</div>
${tooltip_el}
</div>`
}Top institutions
{
// Top-15 institutions by distinct authors, recomputed from filtered pubs so
// the chart tracks the global Kind / Approach / Acquisition filters.
const pub_ids = new Set(pubs_filtered.map(p => p.id))
const authors_by_inst = new Map() // affiliation → Set(author_id)
const country_of = new Map() // affiliation → country (for color)
for (const r of pub_authorship_t) {
if (!pub_ids.has(r.pub_id) || !r.affiliation) continue
if (!authors_by_inst.has(r.affiliation)) authors_by_inst.set(r.affiliation, new Set())
authors_by_inst.get(r.affiliation).add(r.author_id)
if (r.country) country_of.set(r.affiliation, r.country)
}
const rows = Array.from(authors_by_inst, ([institution, authors]) => ({
institution,
country: country_of.get(institution) ?? "",
authors: authors.size
}))
.sort((a, b) => b.authors - a.authors)
.slice(0, 15)
// Built with the SAME expression the y channel uses, so the tick label text
// matches exactly and the linkifier can find it.
const label_of = d => `${d.institution} · ${d.country}`
const href_by_label = new Map(
rows.map(d => [label_of(d), page_href("institutions", d.institution)])
.filter(([, href]) => href)
)
const chart = Plot.plot({
marginLeft: 280,
width: 1100,
height: Math.max(260, rows.length * 26),
marginBottom: 48,
// See the note on the wave chart: the label needs its own line.
x: { label: `Distinct authors (${pubs_filtered.length} papers)`, grid: true, labelOffset: 40 },
y: { label: null },
marks: [
Plot.barX(rows, {
x: "authors",
y: label_of,
// Shared scale (see country_color) so these colours match the
// most-published-authors chart below.
fill: d => country_color(d.country),
sort: { y: "x", reverse: true },
tip: plot_tip_style,
href: d => page_href("institutions", d.institution),
target: "_self"
}),
Plot.ruleX([0])
]
})
return linkify_tick_labels(chart, href_by_label)
}Who’s driving it
The chart below shows the twenty most-published authors.
{
// Re-aggregate top authors from the currently-filtered pubs (so the chart
// reacts to the global Kind / Approach / Acquisition filters).
const counts = new Map()
for (const p of pubs_filtered) {
for (const a of (p.authors ?? "").split(", ").filter(Boolean)) {
counts.set(a, (counts.get(a) ?? 0) + 1)
}
}
// Country per author. Most authors sit in exactly one country, but a
// meaningful minority hold posts in several (Ming Li: Canada + China;
// Bittremieux: Belgium + USA; Pevzner: USA + Germany), so a single fill has
// to pick one. We colour by the author's *dominant* country — the one backing
// most of their affiliation rows, ties broken alphabetically for stability —
// and list every country in the tooltip so a multi-country author is never
// silently reduced to one.
const ctry_tally = new Map() // author -> Map(country -> count)
for (const r of pub_authorship_t) {
if (!r.author || !r.country) continue
if (!ctry_tally.has(r.author)) ctry_tally.set(r.author, new Map())
const m = ctry_tally.get(r.author)
m.set(r.country, (m.get(r.country) ?? 0) + 1)
}
const countries_of = name => {
const m = ctry_tally.get(name)
if (!m) return []
return Array.from(m).sort((a, b) => d3.descending(a[1], b[1]) || d3.ascending(a[0], b[0]))
}
const top = Array.from(counts, ([name, publications]) => {
const cs = countries_of(name)
return {
name,
publications,
country: cs.length ? cs[0][0] : "", // dominant
all_countries: cs.map(c => c[0]),
}
})
.sort((a, b) => b.publications - a.publications)
.slice(0, 20)
// Same expression as the y channel, so the tick label text matches exactly.
const label_of = d =>
`${d.name} · ${d.all_countries.length ? d.all_countries.join(" / ") : "unknown"}`
const href_by_label = new Map(
top.map(d => [label_of(d), page_href("authors", d.name)])
.filter(([, href]) => href)
)
const chart = Plot.plot({
// Wider left margin than the plain-name version: labels now carry the
// country too, and a few authors list several.
marginLeft: 300,
height: Math.max(260, top.length * 22),
marginBottom: 48,
x: { label: `Papers in current filter (${pubs_filtered.length} total)`, grid: true, labelOffset: 40 },
y: { label: null },
marks: [
Plot.barX(top, {
x: "publications",
// '<author> · <country>' to match the Top-institutions axis format.
// Multi-country authors get all of them, so the label never implies a
// single country the colour merely happened to pick.
y: label_of,
// Same shared scale as the Top-institutions chart above.
fill: d => country_color(d.country),
sort: { y: "x", reverse: true },
href: d => page_href("authors", d.name),
target: "_self",
tip: plot_tip_style,
title: d => `${d.name}\n${d.publications} paper${d.publications === 1 ? "" : "s"}\n` +
(d.all_countries.length > 1
? `countries: ${d.all_countries.join(", ")} (coloured by ${d.country})`
: `country: ${d.country || "unknown"}`)
}),
Plot.ruleX([0])
]
})
return linkify_tick_labels(chart, href_by_label)
}The collaboration network
The network below shows how authors with ≥ 3 papers are connected through co-authorship; drag a node to reshape the layout, or hover to highlight a neighborhood.
Edge thickness is Newman fractional collaboration strength: a pair sharing an n-author paper earns 1/(n−1), summed over every paper they share. So an intimate two-author collaboration scores a full 1.0 while each pair in a 53-author consortium scores ~0.02. Without this correction one community benchmark paper alone contributes a 33-node clique that swamps the whole graph. Use min strength to peel away weak ties and expose the dense research groups, and max authors/paper to drop mega-author papers outright.
network_chart = {
const width = 1400
// Legend height grows so all N affiliations render — was clipping at LEGEND_H=110
// (25-item limit) while all_affs runs to 70+. PLOT_H stays at 650 for the network.
const PLOT_H = 650
const cols_per_row = 6 // was 5; slightly denser
const col_w = 220 // 6 × 220 = 1320 px, fits in 1400 chart
const row_h = 18
const LEGEND_TITLE_PX = 20 // room for the "Affiliation (N)" header
// All affiliations per author (an author can have multiple).
const affs_by_author = d3.rollup(
author_affs_t,
v => Array.from(new Set(v.map(d => d.affiliation))).sort(),
d => d.author
)
// Newman fractional strengths under the reactive max-authors cutoff, then
// filtered to edges at or above the min-strength threshold. Nodes are derived
// from the SURVIVING edges, so authors whose only ties were weak consortium
// co-signatures drop off the canvas instead of floating as isolates.
const links = coauth_strength(coauth_max_authors)
.filter(e => e.weight >= coauth_min_strength)
if (!links.length) {
return html`<p style="color:#57606a; font-style:italic; padding:1rem;">
No collaborations at min strength ≥ ${coauth_min_strength.toFixed(2)}
with ≤ ${coauth_max_authors} authors/paper. Lower the threshold to see edges.
</p>`
}
const node_set = new Set()
for (const e of links) { node_set.add(e.source); node_set.add(e.target) }
const nodes = Array.from(node_set, name => ({
id: name,
affiliations: affs_by_author.get(name) ?? ["Unknown"],
degree: 0
}))
const node_by_name = new Map(nodes.map(n => [n.id, n]))
for (const l of links) {
node_by_name.get(l.source).degree += l.weight
node_by_name.get(l.target).degree += l.weight
}
const max_deg = d3.max(nodes, n => n.degree) || 1
// Edge widths scale against the strongest SURVIVING tie, so the thin/thick
// contrast stays readable at any min-strength threshold.
const max_link_w = d3.max(links, l => l.weight) || 1
// Unique affiliations sorted for legend + a 22-slot palette so we don't run out of colors.
const all_affs = Array.from(new Set(nodes.flatMap(n => n.affiliations))).sort()
const palette = d3.schemeTableau10.concat(d3.schemeSet3)
const color = d3.scaleOrdinal().domain(all_affs).range(all_affs.map((_, i) => palette[i % palette.length]))
// Legend height grown to accommodate every affiliation (was hard-coded 110).
const legend_rows = Math.ceil(all_affs.length / cols_per_row)
const LEGEND_H = LEGEND_TITLE_PX + legend_rows * row_h + 12 // + padding
const height = PLOT_H + LEGEND_H
// Primary affiliation = the one shared with most neighbors (for edge coloring).
const adj = new Map(nodes.map(n => [n.id, []]))
for (const l of links) { adj.get(l.source).push(l.target); adj.get(l.target).push(l.source) }
const primary_aff = new Map()
for (const n of nodes) {
if (n.affiliations.length === 1) { primary_aff.set(n.id, n.affiliations[0]); continue }
let best = n.affiliations[0], best_count = -1
for (const aff of n.affiliations) {
const c = adj.get(n.id).reduce((acc, nb) => acc + (node_by_name.get(nb).affiliations.includes(aff) ? 1 : 0), 0)
if (c > best_count) { best = aff; best_count = c }
}
primary_aff.set(n.id, best)
}
// Within-component clustering force: pull nodes of the same primary-aff
// together, toward a target laid out across the INTERIOR of the frame.
//
// These targets used to sit on a single ellipse, cos/sin of i/n at 0.3 width
// and 0.35 height, which put every one of the 60-odd affiliations on one ring
// and left the middle of the chart empty. Combined with the global repulsion
// below it pressed the whole network outward against the frame: measured, 39
// of 173 nodes within 4 px of an edge, as a visible ring of circles with an
// empty middle.
//
// A phyllotaxis disc instead: radius sqrt((i + 0.5) / n) is area-uniform, so
// the targets fill the ellipse rather than tracing it, and the golden angle
// keeps consecutive ones from lining up.
const GOLDEN_ANGLE = Math.PI * (3 - Math.sqrt(5))
const aff_centers = new Map()
all_affs.forEach((aff, i) => {
const r = Math.sqrt((i + 0.5) / all_affs.length)
const a = i * GOLDEN_ANGLE
aff_centers.set(aff, [
width / 2 + Math.cos(a) * r * width * 0.48,
PLOT_H / 2 + Math.sin(a) * r * PLOT_H * 0.30
])
})
const NODE_R = d => 8 + 10 * Math.sqrt(d.degree / max_deg)
const sim = d3.forceSimulation(nodes)
.force("link", d3.forceLink(links).id(d => d.id).distance(70).strength(d => 0.15 + 0.25 * Math.min(1, d.weight)))
// distanceMax is what makes this layout CONVERGE inside the frame. An
// unbounded forceManyBody repels every pair at any distance, so the outward
// pressure grows with the square of the node count while the positional
// forces below stay constant: with 173 nodes the equilibrium was wider than
// the chart, and clamping positions in the tick handler turned that into a
// ring of nodes parked on the border rather than a layout. Capped at 170 px
// the repulsion only separates neighbours, which is all it is for.
.force("charge", d3.forceManyBody().strength(-220).distanceMax(170))
.force("collide", d3.forceCollide().radius(d => NODE_R(d) + 4))
.force("center", d3.forceCenter(width / 2, PLOT_H / 2))
.force("aff_x", d3.forceX(d => aff_centers.get(primary_aff.get(d.id))[0]).strength(0.10))
.force("aff_y", d3.forceY(d => aff_centers.get(primary_aff.get(d.id))[1]).strength(0.10))
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width: 100%; height: auto; font: 11px sans-serif; cursor: grab;")
// Title
svg.append("text")
.attr("x", width / 2)
.attr("y", 22)
.attr("text-anchor", "middle")
.attr("font-size", 16)
.attr("font-weight", "bold")
.attr("fill", "#24292f")
.text(`Co-authorship network · ${nodes.length} authors, ${links.length} ties (Newman strength ≥ ${coauth_min_strength.toFixed(2)}${coauth_max_authors < 60 ? `, ≤ ${coauth_max_authors} authors/paper` : ""})`)
const g = svg.append("g")
svg.call(d3.zoom().on("zoom", ev => g.attr("transform", ev.transform)))
// Edges: tint same-primary-affiliation edges with the affiliation color; others gray.
const link = g.append("g")
.attr("stroke-opacity", 0.45)
.selectAll("line")
.data(links)
.join("line")
.attr("stroke", d => {
const a = primary_aff.get(d.source.id ?? d.source)
const b = primary_aff.get(d.target.id ?? d.target)
return a && a === b ? color(a) : "#bbb"
})
.attr("stroke-width", d => 0.8 + 3.2 * Math.sqrt(d.weight / max_link_w))
// Node groups (one <g> per author; contains either a circle or pie slices).
const node_g = g.append("g")
.selectAll("g.node")
.data(nodes)
.join("g")
.attr("class", "node")
.call(drag(sim))
// Rich HTML tooltip that follows the cursor, matching the world map's .map-tooltip.
// Screen-reader accessibility uses aria-label (not <title>) to avoid stacking a
// native black browser tooltip on top of the styled HTML one.
node_g.attr("aria-label", d =>
`${d.id}. ${d.affiliations.join(", ")}. Collaboration strength ${d.degree.toFixed(2)}.`
)
// Node id is the author display name, which is the slug key. The drag guard
// inside svg_link_wrap keeps a drag from ending in navigation.
svg_link_wrap(node_g, d => page_href("authors", d.id))
const tooltip_el = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_tip = ev => {
const pad = 12
const box = tooltip_el.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
tooltip_el.style.left = `${Math.max(pad, left)}px`
tooltip_el.style.top = `${Math.max(pad, top)}px`
}
const show_tip = (ev, d) => {
tooltip_el.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.id}</div>
<div class="map-tooltip-meta">collaboration strength ${d.degree.toFixed(2)} · ${adj.get(d.id).length} co-author${adj.get(d.id).length === 1 ? "" : "s"}</div>
${d.affiliations.map(aff => html`<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${color(aff)}`}></span>
<span>${aff}</span>
</div>`)}
</div>`)
tooltip_el.classList.add("visible")
tooltip_el.setAttribute("aria-hidden", "false")
move_tip(ev)
}
const hide_tip = () => {
tooltip_el.classList.remove("visible")
tooltip_el.setAttribute("aria-hidden", "true")
}
node_g
.style("cursor", "pointer")
.on("mouseenter", show_tip)
.on("mousemove", move_tip)
.on("mouseleave", hide_tip)
// Render each node: single circle if 1 affiliation, pie wedges if multiple.
node_g.each(function (d) {
const r = NODE_R(d)
const sel = d3.select(this)
if (d.affiliations.length === 1) {
sel.append("circle")
.attr("r", r)
.attr("fill", color(d.affiliations[0]))
.attr("stroke", "#fff")
.attr("stroke-width", 1.5)
} else {
const arc = d3.arc().innerRadius(0).outerRadius(r)
const pie = d3.pie().value(1).sort(null)(d.affiliations.map(a => ({ aff: a })))
sel.selectAll("path")
.data(pie)
.join("path")
.attr("d", arc)
.attr("fill", p => color(p.data.aff))
.attr("stroke", "#fff")
.attr("stroke-width", 1)
}
})
const label = g.append("g")
.attr("pointer-events", "none")
.selectAll("text")
.data(nodes)
.join("text")
.text(d => d.id)
.attr("font-size", 10)
.attr("font-weight", 500)
.attr("fill", "#1d2330")
// A white halo under the glyphs, so a label that sits over an edge or a
// pie-slice node still reads.
.attr("paint-order", "stroke")
.attr("stroke", "#fff")
.attr("stroke-width", 2.5)
.attr("dx", d => NODE_R(d) + 3)
.attr("dy", 3)
// HIDE COLLIDING LABELS. 201 names cannot all be shown in one 1400x790 frame:
// measured 123 overlapping label pairs and 109 labels lying over a node. A
// greedy pass in descending degree keeps the most connected name wherever two
// clash and hides the rest, which is the same rule the citation-flow chart
// uses. Nothing is lost: every node still answers a hover with its full name,
// and the hidden ones are the least connected.
const LABEL_CHAR_PX = 5.8 // 10 px, weight 500, measured against the widest names
const label_order = nodes.slice().sort((a, b) => d3.descending(a.degree, b.degree))
// Six candidate positions per name, nearest first: right of the dot, left of
// it, then each of those nudged a line up or down. Trying only the right side
// showed 52 of 173 names; with all six it is 96, for the same zero overlaps.
const SIDES = [[1, 0], [-1, 0], [1, -11], [-1, -11], [1, 11], [-1, 11]]
const place_labels = () => {
const shown = []
for (const d of label_order) {
const w = d.id.length * LABEL_CHAR_PX
d.label_on = false
for (const [side, dy] of SIDES) {
const x0 = side > 0 ? d.x + NODE_R(d) + 3 : d.x - NODE_R(d) - 3 - w
const y = d.y + dy
const box = { x0, x1: x0 + w, y0: y - 6, y1: y + 6 }
if (box.x0 < 0 || box.x1 > width || box.y0 < TITLE_PX || box.y1 > PLOT_H) continue
const clash =
shown.some(o => box.x0 < o.x1 + 2 && box.x1 + 2 > o.x0 && box.y0 < o.y1 && box.y1 > o.y0) ||
nodes.some(n => box.x0 < n.x + NODE_R(n) && box.x1 > n.x - NODE_R(n)
&& box.y0 < n.y + NODE_R(n) && box.y1 > n.y - NODE_R(n))
if (!clash) {
d.label_on = true; d.label_side = side; d.label_dy = dy
shown.push(box)
break
}
}
}
}
// CONTAIN THE SIMULATION. Neither force chart constrained node positions, and
// charge (-220 here, -350 for algorithm nodes in the bipartite) overwhelms
// centring forces of strength 0.05, so nodes drift outward and KEEP drifting.
// Measured in headless Chrome against the SVG box: this chart had 8 of 201
// nodes outside after 60 s of simulated time and 19 after 180 s, worst 198 px
// out; the bipartite went from 34 of 179 to 47, worst 488 px. It diverges
// rather than settling, and since the simulation runs until the cell is
// invalidated, the longer a reader leaves the page open the more of the
// network silently leaves the frame.
//
// Clamping positions in the tick handler is the GUARD, not the fix, and
// taking it for the fix was a mistake worth recording: a clamp turns
// divergence into a ring of nodes parked on the frame with an empty middle,
// which looks worse than the drift it prevents and hides the cause. The cause
// was unbounded repulsion plus cluster targets on a ring, both addressed
// above; with those right the clamp almost never binds (measured: 0 of 173
// nodes within 4 px of an edge, against 39 before). Labels are kept inside by
// place_labels above, which rejects any box that leaves the frame.
const TITLE_PX = 34
let tick_n = 0
sim.on("tick", () => {
for (const d of nodes) {
const r = NODE_R(d) + 2
d.x = Math.max(r, Math.min(width - r, d.x))
// TITLE_PX keeps the network clear of the chart title, which is drawn at
// y = 22 outside the zoom group: clamping only to the node radius let a
// top-row node and its label sit under the title text, 14 px deep.
d.y = Math.max(TITLE_PX + r, Math.min(PLOT_H - r, d.y))
}
link
.attr("x1", d => d.source.x).attr("y1", d => d.source.y)
.attr("x2", d => d.target.x).attr("y2", d => d.target.y)
node_g.attr("transform", d => `translate(${d.x}, ${d.y})`)
// The collision pass is O(n^2) over 201 nodes, so run it every fourth tick
// rather than on all of them. Positions move by a pixel or two per tick, so
// the placement stays current and the flicker of a label appearing and
// disappearing is damped rather than amplified.
if ((tick_n++ & 3) === 0) place_labels()
label
.attr("x", d => d.x).attr("y", d => d.y + (d.label_dy ?? 0))
.attr("display", d => d.label_on ? null : "none")
.attr("dx", d => (d.label_side ?? 1) > 0 ? NODE_R(d) + 3 : -(NODE_R(d) + 3))
.attr("text-anchor", d => (d.label_side ?? 1) > 0 ? "start" : "end")
})
// Affiliation legend at the bottom, now sized to show every affiliation.
const legend_g = svg.append("g")
.attr("transform", `translate(40, ${PLOT_H + 10})`)
legend_g.append("text")
.attr("x", 0).attr("y", 0)
.attr("font-weight", "bold")
.attr("font-size", 12)
.attr("fill", "#57606a")
.text(`Affiliation (${all_affs.length})`)
legend_g.selectAll("g.leg-item")
.data(all_affs)
.join("g")
.attr("class", "leg-item")
.attr("transform", (_, i) => `translate(${(i % cols_per_row) * col_w}, ${LEGEND_TITLE_PX + Math.floor(i / cols_per_row) * row_h})`)
.call(s => {
s.append("circle").attr("r", 6).attr("cx", 6).attr("cy", -4).attr("fill", color)
s.append("text")
.attr("x", 18).attr("y", 0)
.attr("font-size", 11)
.attr("fill", "#24292f")
.append("title").text(d => d) // native tooltip shows the untruncated name
// Truncate long names to keep the row width bounded (column is 220 px wide).
s.select("text").text(d => d.length > 26 ? d.slice(0, 24) + "…" : d)
})
invalidation.then(() => sim.stop())
function drag(simulation) {
return d3.drag()
.on("start", (ev, d) => { if (!ev.active) simulation.alphaTarget(0.3).restart(); d.fx = d.x; d.fy = d.y })
.on("drag", (ev, d) => { d.fx = ev.x; d.fy = ev.y })
.on("end", (ev, d) => { if (!ev.active) simulation.alphaTarget(0); d.fx = null; d.fy = null })
}
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${svg.node()}</div>
${tooltip_el}
</div>`
}Models and the authors behind them
This second view rewires the same network as a bipartite graph: every prolific author (≥ 3 papers) is linked to the models they helped publish. Algorithm nodes are diamonds colored by their architecture family, so you can see which research groups own which slice of the architectural landscape.
author_algo_network = {
const width = 1400
// Legend height computed dynamically after families_seen is known (below);
// hardcoded 90 clipped anything past the first 3 rows × 5 cols = 15 items.
// 1000, not the 650 the co-authorship chart uses: this graph holds 360 nodes
// against 173, and the layout is only as separable as the room it has. At 650
// the density alone pushed 19 nodes onto the top and bottom edges however the
// forces were tuned; at 820, once the link forces were strong enough to pull
// groups into knots, 29 were pinned there. 1400 x 1000 gives each node about
// 3900 px2 and the knots somewhere to be.
const PLOT_H = 1000
const legend_cols_per_row = 5
const legend_col_w = 260
const legend_row_h = 22
const LEGEND_TITLE_PX = 22
// Same prolific-author filter as the co-auth network (authors with ≥ 3 papers).
// Filter-independent: the co-auth network above has reactive strength/size
// filters, but this chart should always show every prolific author.
const prolific = coauth_all_authors
// Author → algorithm edges, weighted by number of papers connecting them.
// When a publication carries a version tag (Casanovo v1/v2/v5), the edge
// points at the version-suffixed label so each release becomes its own node.
const edge_w = new Map() // key: `${author}|${algo}`
const algo_to_base = new Map() // versioned label → base algorithm name (for family lookup)
// models_described, not models: an edge here says "helped publish this
// model", so a paper that merely ran PEAKS must not make its authors PEAKS
// authors. That edge set is what the algorithm pages show as a method's
// authors, and it took PEAKS from 107 of them to 6.
for (const p of pubs_t) {
if (!p.authors || !p.models_described) continue
const authors = String(p.authors).split(",").map(s => s.trim()).filter(Boolean)
const models = String(p.models_described).split(",").map(s => s.trim()).filter(Boolean)
for (const a of authors) {
if (!prolific.has(a)) continue
for (const m of models) {
const label = p.version ? `${m} ${p.version}` : m
algo_to_base.set(label, m)
const key = a + "<<|>>" + label
edge_w.set(key, (edge_w.get(key) ?? 0) + 1)
}
}
}
const algo_meta = new Map(algorithms_t.map(a => [a.model, a]))
const links = []
const author_set = new Set()
const algo_set = new Set()
for (const [k, w] of edge_w) {
const [a, m] = k.split("<<|>>")
author_set.add(a); algo_set.add(m)
links.push({ source: a, target: m, weight: w })
}
// Resolve a display "family" for each algorithm. Row-level algorithm_family
// is preferred, but when it's null (reviews / benchmarks / downstream-app
// workflow rows / metas / uncategorised adjacent tools) fall back to a
// kind-based pseudo-family so every diamond gets a distinct color instead
// of lumping into a single "Unknown" gray bucket.
const resolveFamily = meta => {
if (meta?.family) return meta.family
switch (meta?.kind) {
case "review": return "Reviews"
case "benchmark": return "Benchmarks"
case "meta": return "Meta / catalogs"
case "post-processor": return "Post-processors (misc)"
case "adjacent": return "Adjacent tools (misc)"
case "downstream-application": return `Application: ${meta.subdomain ?? "misc"}`
default: return "Unknown"
}
}
const nodes = [
...Array.from(author_set, name => ({ id: name, kind: "author", degree: 0 })),
...Array.from(algo_set, name => {
const base = algo_to_base.get(name) ?? name
const meta = algo_meta.get(base)
return { id: name, kind: "algo", family: resolveFamily(meta), degree: 0 }
})
]
const node_by_name = new Map(nodes.map(n => [n.id, n]))
for (const l of links) {
node_by_name.get(l.source).degree += l.weight
node_by_name.get(l.target).degree += l.weight
}
// Color palette: gray for authors, family color for algorithms.
// Family colour comes from the shared scale, so a family is the same colour
// here, in the architectures swim lane and in the code-activity scatter. This
// block used to hold its own copy of the 13 canonical hexes plus a 20-colour
// fallback palette, which drifted from the others as families were added.
const family_color = f => f === "Unknown" ? "#999999" : family_color_scale(f)
const nodeFill = d => d.kind === "author" ? "#cdd6e0" : (family_color(d.family))
const nodeStroke = d => d.kind === "author" ? "#8a94a3" : "#222"
const max_algo_deg = d3.max(nodes.filter(n => n.kind === "algo"), n => n.degree) || 1
const max_auth_deg = d3.max(nodes.filter(n => n.kind === "author"), n => n.degree) || 1
const nodeR = d => d.kind === "author"
? 5 + 8 * Math.sqrt(d.degree / max_auth_deg)
: 9 + 14 * Math.sqrt(d.degree / max_algo_deg)
// Pull algorithm nodes toward a central ring so they cluster by family.
// families_seen is only the families that ACTUALLY have a diamond on the
// chart — a family present in the DB but with no algorithm linked to a
// prolific author (≥3 papers) won't be here.
const families_seen = Array.from(new Set(nodes.filter(n => n.kind === "algo").map(n => n.family))).sort()
// Legend height + SVG total height derived from actual family count so the
// legend never clips (was hard-coded to LEGEND_H=90 which only fit ~15
// items; we now have ~25 pseudo-families thanks to resolveFamily fallback).
const legend_row_count = Math.ceil((families_seen.length + 1) / legend_cols_per_row) // +1 for the author swatch
const LEGEND_H = LEGEND_TITLE_PX + legend_row_count * legend_row_h + 20
const height = PLOT_H + LEGEND_H
// Family targets on a phyllotaxis disc, for the reason spelled out on the
// co-authorship chart above: on a ring they left the middle empty and, with
// unbounded repulsion, pushed the network onto the frame.
const GOLDEN_ANGLE = Math.PI * (3 - Math.sqrt(5))
const family_target = new Map()
families_seen.forEach((f, i) => {
const r = Math.sqrt((i + 0.5) / families_seen.length)
const a = i * GOLDEN_ANGLE
family_target.set(f, [
width / 2 + Math.cos(a) * r * width * 0.50,
PLOT_H / 2 + Math.sin(a) * r * PLOT_H * 0.30
])
})
// EVERY NODE GETS ITS OWN TARGET, authors included. Pulling all 180 authors
// toward one point, `forceX(width / 2)`, is what turned this chart into a
// single blob once the repulsion was capped: with nothing pushing back at
// range, a shared target is a shared destination, and the authors dragged
// their models in after them. An author's target is instead the
// weight-weighted centroid of the family targets of the models it links, so a
// group sits with its own models and an author bridging two families sits
// between them, which is the structure the chart exists to show.
const links_of = new Map(nodes.map(n => [n.id, []]))
for (const l of links) {
links_of.get(l.source).push([l.target, l.weight])
links_of.get(l.target).push([l.source, l.weight])
}
const author_target = new Map()
for (const n of nodes) {
if (n.kind !== "author") continue
let sx = 0, sy = 0, sw = 0
for (const [other, w] of links_of.get(n.id)) {
const t = family_target.get(node_by_name.get(other)?.family)
if (!t) continue
sx += t[0] * w; sy += t[1] * w; sw += w
}
author_target.set(n.id, sw ? [sx / sw, sy / sw] : [width / 2, PLOT_H / 2])
}
const target_of = d => d.kind === "algo"
? family_target.get(d.family) : author_target.get(d.id)
const sim = d3.forceSimulation(nodes)
// Link distance and strength both scale with the tie: a co-published model
// pulls its authors in tight, a single shared paper barely pulls at all.
// A constant 0.35 for every edge is what made the layout uniform -- every
// tie equally important is no structure at all.
.force("link", d3.forceLink(links).id(d => d.id)
.distance(d => 42 + 28 / d.weight)
.strength(d => 0.12 + 0.48 * Math.min(1, d.weight / 2)))
// Capped repulsion, as on the co-authorship chart. This one was worse: an
// author node was held in the frame only by forceX/forceY of strength 0.01,
// against 360 nodes repelling each other at any range, so 94 of them ended
// up within 4 px of an edge, lining the top and bottom. 260 px rather than
// that fix's 200: with per-node targets holding the layout together,
// neighbouring clusters need enough range to push each other apart.
.force("charge", d3.forceManyBody()
.strength(d => d.kind === "algo" ? -300 : -110).distanceMax(300))
.force("collide", d3.forceCollide().radius(d => nodeR(d) + 5))
.force("center", d3.forceCenter(width / 2, PLOT_H / 2))
// The family target is now a NUDGE, not the layout: it decides roughly
// where a group lands so the picture is stable between renders, and the
// links decide the shape. Pulled hard (0.14) it fought the link forces and
// flattened everything into an even mesh.
.force("fam_x", d3.forceX(d => target_of(d)[0])
.strength(d => d.kind === "algo" ? 0.045 : 0.02))
.force("fam_y", d3.forceY(d => target_of(d)[1])
.strength(d => d.kind === "algo" ? 0.045 : 0.02))
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width: 100%; height: auto; font: 11px sans-serif; cursor: grab;")
svg.append("text")
.attr("x", width / 2).attr("y", 22)
.attr("text-anchor", "middle")
.attr("font-size", 16)
.attr("font-weight", "bold")
.attr("fill", "#24292f")
.text("Authors ↔ models · diamonds are algorithms colored by family, circles are authors")
const g = svg.append("g")
svg.call(d3.zoom().on("zoom", ev => g.attr("transform", ev.transform)))
const link = g.append("g")
.attr("stroke", "#aaa")
.attr("stroke-opacity", 0.45)
.selectAll("line")
.data(links)
.join("line")
.attr("stroke-width", d => Math.max(0.8, Math.sqrt(d.weight) * 0.7))
const node_g = g.append("g")
.selectAll("g.node")
.data(nodes)
.join("g")
.attr("class", "node")
.call(drag(sim))
// Rich HTML tooltip that follows the cursor (matches world-map + co-auth style).
// aria-label handles screen-reader accessibility so we don't also stack a
// native black SVG <title> tooltip on top.
node_g.attr("aria-label", d => d.kind === "algo"
? `${d.id} algorithm. Family: ${d.family}. ${d.degree} author links.`
: `${d.id} author. ${d.degree} model links.`
)
// One selection, two entity types: `kind` picks which page to target. Algorithm
// nodes carry a version suffix (Casanovo v2), so fall back to base_model.
svg_link_wrap(node_g, d => d.kind === "algo"
? algo_href(d.base_model ?? d.id)
: page_href("authors", d.id))
const tooltip_el = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_tip = ev => {
const pad = 12
const box = tooltip_el.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
tooltip_el.style.left = `${Math.max(pad, left)}px`
tooltip_el.style.top = `${Math.max(pad, top)}px`
}
const show_tip = (ev, d) => {
if (d.kind === "algo") {
tooltip_el.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.id}</div>
<div class="map-tooltip-meta">algorithm · ${d.degree} author link${d.degree === 1 ? "" : "s"}</div>
<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${family_color(d.family)}`}></span>
<span>${d.family}</span>
</div>
</div>`)
} else {
tooltip_el.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.id}</div>
<div class="map-tooltip-meta">author · ${d.degree} model link${d.degree === 1 ? "" : "s"}</div>
</div>`)
}
tooltip_el.classList.add("visible")
tooltip_el.setAttribute("aria-hidden", "false")
move_tip(ev)
}
const hide_tip = () => {
tooltip_el.classList.remove("visible")
tooltip_el.setAttribute("aria-hidden", "true")
}
node_g
.style("cursor", "pointer")
.on("mouseenter", show_tip)
.on("mousemove", move_tip)
.on("mouseleave", hide_tip)
// Authors as circles; algorithms as diamonds (rotated squares).
node_g.each(function (d) {
const sel = d3.select(this)
const r = nodeR(d)
if (d.kind === "author") {
sel.append("circle")
.attr("r", r)
.attr("fill", nodeFill(d))
.attr("stroke", nodeStroke(d))
.attr("stroke-width", 1.2)
} else {
sel.append("rect")
.attr("x", -r).attr("y", -r)
.attr("width", r * 2).attr("height", r * 2)
.attr("transform", "rotate(45)")
.attr("fill", nodeFill(d))
.attr("stroke", nodeStroke(d))
.attr("stroke-width", 1.5)
}
})
const label = g.append("g")
.attr("pointer-events", "none")
.selectAll("text")
.data(nodes)
.join("text")
.text(d => d.id)
.attr("font-size", d => d.kind === "algo" ? 12 : 9)
.attr("font-weight", d => d.kind === "algo" ? "bold" : 500)
.attr("fill", d => d.kind === "algo" ? "#0b0d12" : "#3b4150")
.attr("paint-order", "stroke")
.attr("stroke", "#fff")
.attr("stroke-width", 2.5)
.attr("dx", d => nodeR(d) + 3)
.attr("dy", 3)
// Same greedy label placement as the co-authorship chart above, and needed
// more: this frame held 457 overlapping label pairs and 282 labels over a
// node. Model names outrank author names whatever the degree, because the
// models are what the chart is arranged around; ties go to the more connected
// node. Every hidden name is still one hover away.
const label_order = nodes.slice().sort((a, b) =>
d3.descending(a.kind === "algo", b.kind === "algo") || d3.descending(a.degree, b.degree))
// Plus a 3 px pad: an underestimated box is an invisible licence to overlap,
// and at 7.0 exactly "PEAKS DB" still landed 4 px over a neighbour's diamond.
const label_px = d => d.kind === "algo" ? 7.0 : 5.2
const label_pad = 3
const SIDES = [[1, 0], [-1, 0], [1, -11], [-1, -11], [1, 11], [-1, 11]]
const place_labels = () => {
const shown = []
for (const d of label_order) {
const w = d.id.length * label_px(d) + 2 * label_pad
const h = d.kind === "algo" ? 8 : 6.5
d.label_on = false
for (const [side, dy] of SIDES) {
const x0 = side > 0 ? d.x + nodeR(d) + 3 : d.x - nodeR(d) - 3 - w
const y = d.y + dy
const box = { x0, x1: x0 + w, y0: y - h, y1: y + h }
if (box.x0 < 0 || box.x1 > width || box.y0 < TITLE_PX || box.y1 > PLOT_H) continue
const clash =
shown.some(o => box.x0 < o.x1 + 2 && box.x1 + 2 > o.x0 && box.y0 < o.y1 && box.y1 > o.y0) ||
nodes.some(n => box.x0 < n.x + nodeR(n) && box.x1 > n.x - nodeR(n)
&& box.y0 < n.y + nodeR(n) && box.y1 > n.y - nodeR(n))
if (!clash) {
d.label_on = true; d.label_side = side; d.label_dy = dy
shown.push(box)
break
}
}
}
}
// Same containment as the co-authorship chart above; see the note there.
const TITLE_PX = 34
let tick_n = 0
sim.on("tick", () => {
for (const d of nodes) {
const r = nodeR(d) + 2
d.x = Math.max(r, Math.min(width - r, d.x))
// Clear of the title, as in the co-authorship chart above.
d.y = Math.max(TITLE_PX + r, Math.min(PLOT_H - r, d.y))
}
link
.attr("x1", d => d.source.x).attr("y1", d => d.source.y)
.attr("x2", d => d.target.x).attr("y2", d => d.target.y)
node_g.attr("transform", d => `translate(${d.x}, ${d.y})`)
if ((tick_n++ & 3) === 0) place_labels()
label
.attr("x", d => d.x).attr("y", d => d.y + (d.label_dy ?? 0))
.attr("display", d => d.label_on ? null : "none")
.attr("dx", d => (d.label_side ?? 1) > 0 ? nodeR(d) + 3 : -(nodeR(d) + 3))
.attr("text-anchor", d => (d.label_side ?? 1) > 0 ? "start" : "end")
})
// Legend: algorithm-family colors + the author swatch.
const legend_g = svg.append("g").attr("transform", `translate(40, ${PLOT_H + 15})`)
legend_g.append("text")
.attr("x", 0).attr("y", 0)
.attr("font-weight", "bold")
.attr("font-size", 12)
.attr("fill", "#57606a")
.text("Algorithm family")
const legend_items = families_seen.map(f => ({ label: f, fill: family_color(f), shape: "diamond" }))
.concat([{ label: "Author (size = # model links)", fill: "#cdd6e0", shape: "circle" }])
legend_g.selectAll("g.leg-item")
.data(legend_items)
.join("g")
.attr("class", "leg-item")
.attr("transform", (_, i) => `translate(${(i % legend_cols_per_row) * legend_col_w}, ${LEGEND_TITLE_PX + Math.floor(i / legend_cols_per_row) * legend_row_h})`)
.call(s => {
s.append(d => d.shape === "diamond"
? document.createElementNS("http://www.w3.org/2000/svg", "rect")
: document.createElementNS("http://www.w3.org/2000/svg", "circle"))
.attr("transform", d => d.shape === "diamond" ? "rotate(45) translate(0,0)" : null)
.each(function (d) {
const el = d3.select(this)
if (d.shape === "diamond") el.attr("x", -7).attr("y", -7).attr("width", 14).attr("height", 14)
else el.attr("r", 7).attr("cx", 0).attr("cy", 0)
el.attr("fill", d.fill).attr("stroke", d.fill === "#cdd6e0" ? "#8a94a3" : "#222").attr("stroke-width", 1.2)
})
s.append("text")
.attr("x", 14).attr("y", 4)
.attr("font-size", 11)
.attr("fill", "#24292f")
.text(d => d.label)
})
invalidation.then(() => sim.stop())
function drag(simulation) {
return d3.drag()
.on("start", (ev, d) => { if (!ev.active) simulation.alphaTarget(0.3).restart(); d.fx = d.x; d.fy = d.y })
.on("drag", (ev, d) => { d.fx = ev.x; d.fy = ev.y })
.on("end", (ev, d) => { if (!ev.active) simulation.alphaTarget(0); d.fx = null; d.fy = null })
}
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${svg.node()}</div>
${tooltip_el}
</div>`
}How the field cites itself
A chronological citation arc diagram. Papers are placed left-to-right by publication date and stratified vertically by kind; within each row, the most-cited papers float to the top. Each arc connects a citing paper (right end) to a paper it cites (left end), curving upward above the row. Hover any paper to highlight the citations into it (red) and out of it (blue), and dim everything else.
Edges resolved from Crossref (by DOI) and Semantic Scholar (by DOI or title-search fallback), matched back to publications via DOI-exact (and, for refs without a DOI, fuzzy-title with token-set ratio ≥ 92). Only intra-catalog citations are drawn; references to papers outside the catalog are filtered out. Every arrow runs citing → cited, so the arrowhead always lands on the older paper.
citation_arcs = {
if (!citations_t.length) {
return html`<p style="color:#57606a; font-style:italic; padding:1rem;">
No citation edges yet. Run <code>uv run python build_citations.py</code> to populate the graph.
</p>`
}
// ===== Layout parameters =====
const width = 1600
const height = 820
const marginTop = 120 // extra headroom for arc apexes near the top rows
const marginBottom = 60
const marginLeft = 130
const marginRight = 52 // 12 px over, then 4 px still over at 45
const innerWidth = width - marginLeft - marginRight
const innerHeight = height - marginTop - marginBottom
// ===== Per-publication node data =====
const pub_by_id = new Map(pubs_t.map(p => [p.id, p]))
const cite_count = new Map() // pub_id → in-degree (times cited by other catalog papers)
const cite_out = new Map() // pub_id → out-degree (papers in catalog it cites)
for (const e of citations_t) {
cite_count.set(e.cited_id, (cite_count.get(e.cited_id) ?? 0) + 1)
cite_out.set(e.citing_id, (cite_out.get(e.citing_id) ?? 0) + 1)
}
// Y-strata by kind. Meta types at the top, core algorithms at the bottom, so
// arcs (which always curve upward) have the most headroom for citations heading
// *into* heavily-cited algorithm-row papers.
const KIND_ORDER = [
{ kind: "meta", label: "Meta" },
{ kind: "benchmark", label: "Benchmarks" },
{ kind: "review", label: "Reviews / surveys" },
{ kind: "adjacent", label: "Adjacent" },
{ kind: "downstream-application", label: "Downstream apps" },
{ kind: "post-processor", label: "Post-processors" },
{ kind: "algorithm", label: "Algorithms" }
]
const kind_color = {
"algorithm": "#1f6feb",
"post-processor": "#bf8700",
"downstream-application": "#1a7f37",
"adjacent": "#a371f7",
"review": "#cf222e",
"benchmark": "#0969da",
"meta": "#8c959f",
"unknown": "#8c959f"
}
const row_h = innerHeight / KIND_ORDER.length
const row_index = new Map(KIND_ORDER.map((d, i) => [d.kind, i]))
// Only include papers that participate in ≥ 1 edge.
const involved = new Set()
for (const e of citations_t) { involved.add(e.citing_id); involved.add(e.cited_id) }
const nodes = Array.from(involved, id => {
const p = pub_by_id.get(id) ?? {}
const d = p.date instanceof Date ? p.date : new Date(p.date ?? Date.now())
return {
id,
title: p.title ?? `#${id}`,
year: p.year,
date: d,
kind: p.kind ?? "unknown",
models: p.models ?? "",
in_deg: cite_count.get(id) ?? 0,
out_deg: cite_out.get(id) ?? 0
}
}).filter(n => n.date instanceof Date && !isNaN(n.date))
// X scale: publication date.
const x_extent = d3.extent(nodes, n => n.date)
const x_pad = (x_extent[1] - x_extent[0]) * 0.02 || 1e9
const xScale = d3.scaleTime()
.domain([new Date(+x_extent[0] - x_pad), new Date(+x_extent[1] + x_pad)])
.range([0, innerWidth])
// Within-row jitter: rank papers in each row by date so they bucket into vertical lanes.
const max_in = d3.max(nodes, n => n.in_deg) || 1
for (const n of nodes) {
const ri = row_index.get(n.kind) ?? KIND_ORDER.length - 1
const top = ri * row_h + 16
const bot = (ri + 1) * row_h - 16
// Top of row = most-cited; spread less-cited papers downward.
const r = 1 - Math.sqrt(n.in_deg / max_in)
n.y = top + r * (bot - top)
n.x = xScale(n.date)
}
const node_by_id = new Map(nodes.map(n => [n.id, n]))
const links = citations_t
.map(e => ({ source: node_by_id.get(e.citing_id), target: node_by_id.get(e.cited_id), source_kind: e.source }))
.filter(l => l.source && l.target)
// ===== Draw =====
const svg = d3.create("svg")
.attr("viewBox", [0, 0, width, height])
.attr("style", "max-width: 100%; height: auto; font: 11px sans-serif;")
const root = svg.append("g").attr("transform", `translate(${marginLeft}, ${marginTop})`)
// Row backgrounds + labels
root.selectAll("g.row")
.data(KIND_ORDER)
.join("g")
.attr("class", "row")
.call(g => {
g.append("rect")
.attr("x", 0)
.attr("y", (_, i) => i * row_h)
.attr("width", innerWidth)
.attr("height", row_h)
.attr("fill", d => kind_color[d.kind] ?? "#999")
.attr("fill-opacity", 0.05)
g.append("text")
.attr("x", -8)
.attr("y", (_, i) => i * row_h + row_h / 2 + 4)
.attr("text-anchor", "end")
.attr("font-size", 12)
.attr("font-weight", "bold")
.attr("fill", d => kind_color[d.kind] ?? "#999")
.text(d => d.label)
})
// Time axis
root.append("g")
.attr("transform", `translate(0, ${innerHeight})`)
.call(d3.axisBottom(xScale).ticks(d3.timeYear.every(2)).tickFormat(d3.timeFormat("%Y")))
.selectAll("text")
.attr("font-size", 11)
// Title
svg.append("text")
.attr("x", width / 2).attr("y", 20)
.attr("text-anchor", "middle")
.attr("font-size", 15)
.attr("font-weight", "bold")
.attr("fill", "#24292f")
.text(`Citation flow · ${nodes.length} papers, ${links.length} intra-catalog citations (citing → cited)`)
// Citation arcs. Cubic-Bezier curving upward (above the baseline).
function arcPath(d) {
const x1 = d.source.x, y1 = d.source.y
const x2 = d.target.x, y2 = d.target.y
const mx = (x1 + x2) / 2
// Lift the control points by a fraction of the horizontal span.
const lift = Math.min(180, Math.abs(x1 - x2) * 0.6)
const cy1 = Math.min(y1, y2) - lift
const cy2 = cy1
return `M ${x1} ${y1} C ${mx} ${cy1}, ${mx} ${cy2}, ${x2} ${y2}`
}
// Arrowhead markers. We need one per arc color (default/red/blue) because
// SVG markers inherit their fill from the marker itself, not the path's
// stroke. Refs get swapped on hover so the marker color matches the arc.
const defs = svg.append("defs")
const mk_marker = (id, color) => defs.append("marker")
.attr("id", id).attr("viewBox", "0 0 10 10").attr("refX", 9).attr("refY", 5)
.attr("markerWidth", 6).attr("markerHeight", 6).attr("orient", "auto-start-reverse")
.append("path").attr("d", "M0,0 L10,5 L0,10 z").attr("fill", color)
mk_marker("arr-default", "#94a3b8")
mk_marker("arr-red", "#cf222e")
mk_marker("arr-blue", "#1f6feb")
const arc_layer = root.append("g").attr("class", "arcs").attr("fill", "none")
const arc = arc_layer.selectAll("path")
.data(links)
.join("path")
.attr("d", arcPath)
.attr("stroke", "#94a3b8")
.attr("stroke-opacity", 0.18)
.attr("stroke-width", d => d.source_kind === "both" ? 0.9 : 0.6)
.attr("marker-end", "url(#arr-default)")
// Nodes
const nodeR = d => 3 + 5 * Math.sqrt(d.in_deg / max_in)
const node_layer = root.append("g").attr("class", "nodes")
const node = node_layer.selectAll("circle")
.data(nodes)
.join("circle")
.attr("cx", d => d.x)
.attr("cy", d => d.y)
.attr("r", nodeR)
.attr("fill", d => kind_color[d.kind] ?? "#999")
.attr("stroke", "#fff")
.attr("stroke-width", 0.8)
.attr("cursor", "pointer")
// These nodes already carry the publication id, which makes this the most
// direct link site on the page. No drag here, so the guard is a no-op.
svg_link_wrap(node, d => page_href("publications", d.id))
// Styled tooltip that follows the cursor (matches world-map / co-auth /
// bipartite / Sankey). aria-label handles screen-reader accessibility.
node.attr("aria-label", d =>
`${d.models || d.title}. ${d.year ?? ""}. ${d.kind}. Cited by ${d.in_deg}, cites ${d.out_deg}.`
)
const arc_tooltip = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_arc_tip = ev => {
const pad = 12
const box = arc_tooltip.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
arc_tooltip.style.left = `${Math.max(pad, left)}px`
arc_tooltip.style.top = `${Math.max(pad, top)}px`
}
// Note: the actual mouseenter/mouseleave handlers are attached below,
// AFTER out_edges/in_edges are computed, so the tooltip + highlight logic
// can live in one handler. Attach mousemove here since it only needs the
// tooltip element and never fights with the highlight handler.
node.on("mousemove", move_arc_tip)
// Labels: only the top in-degree papers (one per row, so labels are legible).
const top_per_row = new Map()
for (const n of nodes.slice().sort((a, b) => b.in_deg - a.in_deg)) {
const r = n.kind
if (!top_per_row.has(r)) top_per_row.set(r, [])
if (top_per_row.get(r).length < 5 && n.in_deg >= 4) top_per_row.get(r).push(n)
}
const labelled = Array.from(top_per_row.values()).flat()
// Truncated at 26 characters: the box the placer reserves is estimated from
// the character count, and an unbounded model name made that estimate wrong
// by enough to land the label on a dot.
const labelText = d => {
const t = (d.models?.split(",")[0]?.trim()) || d.title
return t.length > 26 ? t.slice(0, 25) + "…" : t
}
// EVERY node is an obstacle, not just the five per row that carry a label: a
// label was landing on unlabelled dots, 18 of them, up to 7 px deep.
//
// Candidate positions are expressed RELATIVE TO THE DOT rather than as fixed
// offsets clamped into the row. Two things went wrong with the clamped form.
// A label near the top of its row got clamped back down onto its own dot, and
// its own dot was exempt from the collision test, so nothing objected: that is
// the last four of the original 18. And the box was centred on the text's
// ANCHOR, which is its baseline, not its middle -- glyphs sit ~11 px above the
// baseline and ~3 px below it, so the box was half a line out.
const ASC = 11, DESC = 3
const label_boxes = []
for (const group of d3.groups(labelled, d => d.kind).map(([, values]) => values)) {
const placed = []
const ordered = group.slice().sort((a, b) => d3.descending(a.in_deg, b.in_deg))
for (const d of ordered) {
const text = labelText(d)
const ri = row_index.get(d.kind) ?? KIND_ORDER.length - 1
const rowTop = ri * row_h + 6
const rowBottom = (ri + 1) * row_h - 6
const r = nodeR(d)
// No upper cap on the width: the estimate used to stop at 150 px while the
// text kept going, so a long model name overhung its own box by 67 px and
// the collision test never saw it.
const w = Math.max(32, text.length * 6.6)
// Baselines that clear the dot, nearest first: just above it, just below
// it, then in 14 px steps outward.
const cands = []
for (let k = 0; k < 4; k++) {
cands.push(d.y - r - DESC - 2 - k * 14)
cands.push(d.y + r + ASC + 2 + k * 14)
}
let chosen = null
for (const y of cands) {
const box = { x0: d.x - w / 2, x1: d.x + w / 2, y0: y - ASC, y1: y + DESC }
if (box.y0 < rowTop || box.y1 > rowBottom) continue
const hits = o => box.x0 < o.x1 && box.x1 > o.x0 && box.y0 < o.y1 && box.y1 > o.y0
const overlaps = placed.some(p =>
box.x0 < p.x1 + 6 && box.x1 + 6 > p.x0 && box.y0 < p.y1 + 3 && box.y1 + 3 > p.y0
) || nodes.some(n => hits({
x0: n.x - nodeR(n) - 2, x1: n.x + nodeR(n) + 2,
y0: n.y - nodeR(n) - 2, y1: n.y + nodeR(n) + 2
}))
if (!overlaps) {
chosen = { ...box, d, text, y }
break
}
}
if (chosen) {
placed.push(chosen)
label_boxes.push(chosen)
}
}
}
const label_layer = root.append("g").attr("class", "labels").attr("pointer-events", "none")
label_layer.selectAll("line")
.data(label_boxes.filter(l => Math.abs(l.y - (l.d.y - nodeR(l.d) - 5)) > 8))
.join("line")
.attr("x1", d => d.d.x)
.attr("y1", d => d.d.y - nodeR(d.d) - 1)
.attr("x2", d => d.d.x)
.attr("y2", d => d.y + 4)
.attr("stroke", "#8c959f")
.attr("stroke-opacity", 0.55)
.attr("stroke-width", 0.6)
label_layer.selectAll("text")
.data(label_boxes)
.join("text")
.text(d => d.text)
.attr("x", d => d.d.x)
.attr("y", d => d.y)
.attr("text-anchor", "middle")
.attr("font-size", 10)
.attr("font-weight", 600)
.attr("paint-order", "stroke")
.attr("stroke", "white")
.attr("stroke-width", 3)
.attr("fill", "#1d2330")
// ===== Hover behaviour =====
// Precompute neighborhood maps.
const out_edges = new Map()
const in_edges = new Map()
for (const l of links) {
if (!out_edges.has(l.source.id)) out_edges.set(l.source.id, [])
if (!in_edges.has (l.target.id)) in_edges.set (l.target.id, [])
out_edges.get(l.source.id).push(l)
in_edges.get (l.target.id).push(l)
}
// Combined mouseenter: show the styled tooltip AND highlight the citation
// neighbourhood in one pass — attaching two separate .on('mouseenter')
// handlers here would silently replace the first (D3 semantics).
node.on("mouseenter", function (ev, focus) {
// 1. Show styled tooltip
arc_tooltip.replaceChildren(html`<div>
<div class="map-tooltip-title">${focus.models || "(no model)"}</div>
<div class="map-tooltip-meta">${focus.title}</div>
<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${kind_color[focus.kind] ?? "#999"}`}></span>
<span>${focus.kind}${focus.year ? ` · ${focus.year}` : ""}</span>
</div>
<div class="map-tooltip-row"><span>cited by <b>${focus.in_deg}</b> · cites <b>${focus.out_deg}</b></span></div>
</div>`)
arc_tooltip.classList.add("visible")
arc_tooltip.setAttribute("aria-hidden", "false")
move_arc_tip(ev)
// 2. Highlight citation neighbourhood
const ins = new Set((in_edges.get(focus.id) ?? []).map(l => l.source.id))
const outs = new Set((out_edges.get(focus.id) ?? []).map(l => l.target.id))
arc
.attr("stroke", l => {
if (l.target.id === focus.id) return "#cf222e" // who cites this paper
if (l.source.id === focus.id) return "#1f6feb" // who this paper cites
return "#94a3b8"
})
.attr("marker-end", l => {
if (l.target.id === focus.id) return "url(#arr-red)"
if (l.source.id === focus.id) return "url(#arr-blue)"
return "url(#arr-default)"
})
.attr("stroke-opacity", l => (l.source.id === focus.id || l.target.id === focus.id) ? 0.85 : 0.04)
.attr("stroke-width", l => (l.source.id === focus.id || l.target.id === focus.id) ? 1.4 : 0.4)
node.attr("opacity", d => {
if (d.id === focus.id) return 1
if (ins.has(d.id) || outs.has(d.id)) return 1
return 0.22
})
})
node.on("mouseleave", function () {
// 1. Hide tooltip
arc_tooltip.classList.remove("visible")
arc_tooltip.setAttribute("aria-hidden", "true")
// 2. Reset arc + node styling
arc.attr("stroke", "#94a3b8").attr("stroke-opacity", 0.18)
.attr("stroke-width", d => d.source_kind === "both" ? 0.9 : 0.6)
.attr("marker-end", "url(#arr-default)")
node.attr("opacity", 1)
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${svg.node()}</div>
${arc_tooltip}
</div>`
}Academic impact by citation count
Publication-level global citation counts from OpenAlex cited_by_count. Counts are matched to catalog publications by DOI first, then by high-confidence title search when DOI lookup is unavailable.
impact_rows = {
const rows = pubs_filtered.map(p => ({
...p,
global_citations: publication_impact_by_pub.get(p.id)?.cited_by_count ?? null,
citation_year: publication_impact_by_pub.get(p.id)?.year_collected ?? null,
citation_match: publication_impact_by_pub.get(p.id)?.match_method ?? null
}))
return rows.sort((a, b) =>
d3.descending(a.global_citations ?? -1, b.global_citations ?? -1) ||
d3.ascending(a.year ?? Infinity, b.year ?? Infinity) ||
d3.ascending(a.title, b.title)
)
}{
const table = Inputs.table(impact_search, {
columns: ["global_citations", "year", "models", "kind", "title", "authors", "journal", "type", "citation_year"],
header: {
global_citations: "Global citations",
year: "Year",
models: "Method(s)",
kind: "Kind",
title: "Title",
authors: "Authors",
journal: "Venue",
type: "Type",
citation_year: "Collected"
},
format: {
global_citations: c => c == null ? "" : String(c),
year: y => y == null ? "" : String(y),
citation_year: y => y == null ? "" : String(y),
title: (t, i) => {
const row = impact_search[i]
const internal = page_href("publications", row?.id)
const external = row?.url || (row?.doi ? `https://doi.org/${row.doi}` : null)
const label = internal ? htl.html`<a href=${internal}>${t}</a>` : t
return external
? htl.html`${label} <a href="${external}" target="_blank" rel="noopener"
title="Publisher / DOI" style="text-decoration:none">↗</a>`
: label
},
models: m => page_links("algorithms", m, { max: 4 }),
journal: j => page_link("venues", j, j),
authors: a => page_links("authors", a, { max: 3 })
},
sort: "global_citations",
reverse: true,
rows: 30,
width: { global_citations: 120, year: 60, models: 170, kind: 130, type: 110, citation_year: 90 }
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${table}</div>
</div>`
}Code activity
Open-source uptake snapshot from the GitHub API: stars, fork count, open / closed issues + PRs, the most recent push, and the latest released tag (falling back to the latest plain tag if the project doesn’t formally release). Pulled offline by build_repo_metrics.py against algorithm_repository.url. Non-GitHub repos (PyPI, project home pages, anonymised review repos) aren’t counted here.
{
// Bar chart of the top 20 repos by the selected metric.
// repo_metrics_by_url is deduped so repos backing multiple algorithms
// (InstaNovo + InstaNovo-P, ...) surface as one bar with a combined label.
const top = repo_metrics_by_url
.filter(r => r[repo_metric] != null)
.sort((a, b) => (b[repo_metric] ?? 0) - (a[repo_metric] ?? 0))
.slice(0, 20)
.map(r => ({
...r,
label: r.model + (r.url.includes("/tree/") ? "*" : ""),
}))
const labelfmt = {
stars: "Stars (★)", forks: "Forks", open_issues: "Open issues",
open_prs: "Open PRs", closed_issues: "Closed issues", closed_prs: "Closed PRs"
}
// The label carries a trailing "*" for repos pointing at a subdirectory, so
// the lookup keys off the label exactly as rendered.
const href_by_label = new Map(
top.map(d => [d.label, algo_href(d.model)]).filter(([, href]) => href)
)
const chart = Plot.plot({
marginLeft: 235, width: 1100, // measured 50 px of repo-name overflow
height: Math.max(280, top.length * 24),
marginBottom: 48,
// labelOffset drops the axis label onto its own line; on the tick baseline
// it touched the '200' tick by 3 px.
x: { label: labelfmt[repo_metric], grid: true, labelOffset: 40 },
y: { label: null },
marks: [
Plot.barX(top, {
x: repo_metric,
y: "label",
fill: "#1f6feb",
sort: { y: "x", reverse: true },
href: d => algo_href(d.model),
target: "_self"
}),
Plot.ruleX([0])
]
})
// Tick labels too, so the method name is clickable and not just the bar. This
// touches only g[aria-label="y-axis tick label"], leaving the bar rects the
// tooltip binding below depends on untouched.
linkify_tick_labels(chart, href_by_label)
// Rich .map-tooltip on the bars. Plot renders barX rects in the sort order
// (already computed above) — we replicate that here to bind data by DOM
// index. Same helper pattern as the swim-lane charts.
const repo_bar_tooltip = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_repo_bar_tip = ev => {
const pad = 12
const box = repo_bar_tooltip.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
repo_bar_tooltip.style.left = `${Math.max(pad, left)}px`
repo_bar_tooltip.style.top = `${Math.max(pad, top)}px`
}
const sorted = top.slice().sort((a, b) => (b[repo_metric] ?? 0) - (a[repo_metric] ?? 0))
// Plot labels barX marks with aria-label='bar' (className was silently
// ignored). Selector scoped to this chart's DOM.
chart.querySelectorAll('g[aria-label="bar"] rect').forEach((rect, i) => {
const d = sorted[i]
if (!d) return
rect.style.cursor = "pointer"
rect.setAttribute("aria-label",
`${d.model}. ${d.stars} stars, ${d.forks} forks. ${d.open_issues} open issues, ${d.open_prs} open PRs.`)
rect.addEventListener("mouseenter", ev => {
repo_bar_tooltip.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.model}</div>
<div class="map-tooltip-meta">
<a href=${d.url} target="_blank" rel="noopener" style="color:inherit;">${d.url.replace(/^https?:\/\/(?:www\.)?github\.com\//, "")}</a>
</div>
<div class="map-tooltip-row"><span>★ <b>${d.stars ?? 0}</b> · forks <b>${d.forks ?? 0}</b></span></div>
<div class="map-tooltip-row"><span>issues open/closed: <b>${d.open_issues ?? 0}</b>/${d.closed_issues ?? 0}</span></div>
<div class="map-tooltip-row"><span>PRs open/closed: <b>${d.open_prs ?? 0}</b>/${d.closed_prs ?? 0}</span></div>
${d.latest_release ? html`<div class="map-tooltip-row"><span>release: ${d.latest_release}</span></div>` : ""}
${d.last_pushed ? html`<div class="map-tooltip-row"><span>last pushed: ${d.last_pushed.toISOString().slice(0,10)}</span></div>` : ""}
</div>`)
repo_bar_tooltip.classList.add("visible")
repo_bar_tooltip.setAttribute("aria-hidden", "false")
move_repo_bar_tip(ev)
})
rect.addEventListener("mousemove", move_repo_bar_tip)
rect.addEventListener("mouseleave", () => {
repo_bar_tooltip.classList.remove("visible")
repo_bar_tooltip.setAttribute("aria-hidden", "true")
})
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${chart}</div>
${repo_bar_tooltip}
</div>`
}Plot popularity (★, log scale) against staleness (months since last push). Vibrant repos cluster on the right at higher star counts; abandoned-but-historically-popular projects drift toward the left. Dot size scales with open issues + open PRs. Hover any dot for the repo name, url, and counts.
{
// Scatter: stars vs months since last push. Stale-but-popular repos sit in the
// upper right; actively maintained but small projects sit lower left.
// repo_metrics_by_url dedupes shared-repo algorithm pairs (e.g. InstaNovo +
// InstaNovo-P) so we don't stack two labels on the same coordinate.
const now = Date.now()
const points = repo_metrics_by_url
.filter(r => r.stars != null && r.last_pushed != null)
.map(r => ({
...r,
months_since_push: (now - r.last_pushed.getTime()) / (1000*60*60*24*30.44)
}))
// GUI wrappers, benchmark harnesses, meta resources and rescoring frameworks
// legitimately have no algorithm_family; they aren't sequencing
// architectures. Name that group so the legend shows a label rather than a
// blank swatch, which is what made its colour look like a bug.
const TOOLING = "Tooling (no single architecture)"
const fam_domain = Array.from(new Set(points.map(d => d.family ?? TOOLING))).sort()
const chart = Plot.plot({
width: 1100, height: 460,
marginLeft: 60, marginBottom: 50,
x: { label: "Months since last push →", grid: true, reverse: true },
y: { label: "Stars (★)", type: "log", grid: true },
// Explicit domain + range pulled from the shared family scale, so each
// family keeps its architectures-chart colour and adding/removing a
// category can't reshuffle the rest. Plot's default positional assignment
// is what previously turned Transformer (AR) brown.
color: {
domain: fam_domain,
range: fam_domain.map(family_color_scale),
legend: true,
label: "Family"
},
marks: [
// No always-on labels: even with dodgeY the ≥20-star subset still
// overlaps in dense clusters, and choosing WHICH stars threshold to
// label reads as arbitrary. Rely on the styled .map-tooltip on hover
// for identification — every dot gets a rich popup with the same info.
Plot.dot(points, {
x: "months_since_push", y: "stars",
fill: d => d.family ?? TOOLING,
r: d => 4 + Math.sqrt((d.open_issues ?? 0) + (d.open_prs ?? 0)),
href: d => algo_href(d.model),
target: "_self"
})
]
})
// Same tooltip pattern as the repo bars.
const repo_dot_tooltip = html`<div class="map-tooltip" role="tooltip" aria-hidden="true"></div>`
const move_repo_dot_tip = ev => {
const pad = 12
const box = repo_dot_tooltip.getBoundingClientRect()
const left = Math.min(ev.clientX + pad, window.innerWidth - box.width - pad)
const top = Math.min(ev.clientY + pad, window.innerHeight - box.height - pad)
repo_dot_tooltip.style.left = `${Math.max(pad, left)}px`
repo_dot_tooltip.style.top = `${Math.max(pad, top)}px`
}
// Plot re-sorts dot marks by DESCENDING RADIUS whenever r is a channel, so
// that small dots draw on top of large ones and stay clickable. DOM order is
// therefore NOT the order of `points`, and binding tooltips by DOM index
// attributed every repo's stats to some other repo's dot. Replicate the draw
// order instead (both Array.prototype.sort and Plot's sort are stable, so
// equal-radius ties keep input order).
const r_of = d => 4 + Math.sqrt((d.open_issues ?? 0) + (d.open_prs ?? 0))
const draw_order = points.slice().sort((a, b) => r_of(b) - r_of(a))
// Selector scoped to this Plot chart; aria-label='dot' is the marker Plot
// uses for dot marks (className mark option was silently ignored).
const dots = Array.from(chart.querySelectorAll('g[aria-label="dot"] circle'))
// Fail closed. The rendered radii are the r CHANNEL passed through Plot's r
// scale, so they can't be compared numerically against r_of(); what we can
// check is the assumption itself — that Plot emitted the dots in descending
// radius order and that the counts line up. If either stops holding, skip
// the tooltips entirely rather than confidently labelling dots with the
// wrong repository.
const dom_r = dots.map(c => +c.getAttribute("r"))
const order_holds = dom_r.length === draw_order.length
&& dom_r.every((v, i) => i === 0 || dom_r[i - 1] >= v - 1e-9)
if (order_holds) dots.forEach((circle, i) => {
const d = draw_order[i]
if (!d) return
// Pull the actual fill Plot picked for this dot — matches the legend swatch.
const family_swatch = circle.getAttribute("fill") ?? "#4C72B0"
circle.style.cursor = "pointer"
circle.setAttribute("aria-label",
`${d.model}. ${d.stars} stars. ${d.months_since_push.toFixed(1)} months since last push.`)
circle.addEventListener("mouseenter", ev => {
repo_dot_tooltip.replaceChildren(html`<div>
<div class="map-tooltip-title">${d.model}</div>
<div class="map-tooltip-meta">
<a href=${d.url} target="_blank" rel="noopener" style="color:inherit;">${d.url.replace(/^https?:\/\/(?:www\.)?github\.com\//, "")}</a>
</div>
<div class="map-tooltip-row">
<span class="map-tooltip-swatch" style=${`background:${family_swatch}`}></span>
<span>${d.family ?? TOOLING}</span>
</div>
<div class="map-tooltip-row"><span>★ <b>${d.stars}</b> · ${d.months_since_push.toFixed(1)} mo since push</span></div>
<div class="map-tooltip-row"><span>open issues <b>${d.open_issues ?? 0}</b> · open PRs <b>${d.open_prs ?? 0}</b></span></div>
</div>`)
repo_dot_tooltip.classList.add("visible")
repo_dot_tooltip.setAttribute("aria-hidden", "false")
move_repo_dot_tip(ev)
})
circle.addEventListener("mousemove", move_repo_dot_tip)
circle.addEventListener("mouseleave", () => {
repo_dot_tooltip.classList.remove("visible")
repo_dot_tooltip.setAttribute("aria-hidden", "true")
})
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${chart}</div>
${repo_dot_tooltip}
</div>`
}{
// Full sortable table of every repository with metrics.
const rows = repo_metrics_t.slice()
const table = Inputs.table(rows, {
columns: ["model", "family", "stars", "forks", "open_issues",
"closed_issues", "open_prs", "closed_prs",
"latest_release", "last_pushed", "url"],
header: {
model: "Method", family: "Family", stars: "★", forks: "Forks",
open_issues: "Issues (open)", closed_issues: "Issues (closed)",
open_prs: "PRs (open)", closed_prs: "PRs (closed)",
latest_release: "Release", last_pushed: "Last updated", url: "Repo"
},
format: {
stars: v => v == null ? "" : String(v),
last_pushed: d => d ? d.toISOString().slice(0, 10) : "",
latest_release: v => v ?? "",
// One repo can back several algorithms (instadeepai/instanovo serves both
// InstaNovo and InstaNovo+), in which case model is "A / B".
model: m => page_links("algorithms", m, { sep: " / ", max: 3 }),
url: u => u ? htl.html`<a href="${u}" target="_blank" rel="noopener">${u.replace(/^https?:\/\/(?:www\.)?github\.com\//, "")}</a>` : ""
},
sort: "stars", reverse: true,
rows: 25,
width: { model: 130, family: 130, stars: 70, forks: 70,
open_issues: 90, closed_issues: 100, open_prs: 80, closed_prs: 100,
latest_release: 100, last_pushed: 110, url: 240 }
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${table}</div>
</div>`
}Where it appears
Most papers in this space appear first on bioRxiv or arXiv. Toggle preprints vs. peer-reviewed to see how the venue distribution shifts.
{
// Re-aggregate venues from the globally-filtered pubs so the chart respects
// Kind / Approach / Acquisition, then apply the section's preprint/peer-reviewed radio.
const subset = pubs_filtered.filter(p => venue_type === "all" || p.type === venue_type)
const grouped = Array.from(
d3.rollup(subset.filter(p => p.journal), v => v.length, p => p.journal),
([venue, papers]) => ({ venue, papers })
).sort((a, b) => d3.descending(a.papers, b.papers)).slice(0, 15)
// The tick label IS the venue name here, so the lookup key needs no
// reconstruction, unlike the composite "name · country" axes elsewhere.
const href_by_label = new Map(
grouped.map(d => [d.venue, page_href("venues", d.venue)])
.filter(([, href]) => href)
)
const chart = Plot.plot({
marginLeft: 270, // measured 85 px of venue-name overflow
height: Math.max(260, grouped.length * 24),
// labelOffset + marginBottom give the axis label its own line. Without
// them Plot puts it level with the tick labels, where it overlapped the
// '40' tick by 3 px -- remedy 3 under 'No chart on the page has an
// overlapping label', applied here after check_chart_overlap.py caught it.
x: { label: "Papers", grid: true, labelOffset: 40 },
marginBottom: 48,
y: { label: null },
marks: [
Plot.barX(grouped, {
x: "papers",
y: "venue",
fill: "#6f42c1",
sort: { y: "x", reverse: true },
tip: plot_tip_style,
href: d => page_href("venues", d.venue),
target: "_self"
}),
Plot.ruleX([0])
]
})
return linkify_tick_labels(chart, href_by_label)
}Venue citedness (open-data analog of the Impact Factor)
Two-year mean citedness from OpenAlex (summary_stats.2yr_mean_citedness). Methodologically equivalent to the Clarivate Impact Factor formula (mean citations in year t to articles published in years t-1 and t-2), but computed over OpenAlex’s open Crossref-aggregated citation graph rather than the paywalled Web of Science one. Conferences and preprint servers are omitted (their non-rolling publication schedule makes the metric misleading). Built offline via build_journal_metrics.py; refresh annually.
{
// Decorate each venue row with the catalog's paper count for that venue,
// so users can spot where heavy curation overlaps with high-citedness venues.
const paper_counts = new Map()
for (const p of pubs_t) {
if (!p.journal) continue
paper_counts.set(p.journal, (paper_counts.get(p.journal) ?? 0) + 1)
}
const rows = venue_if_search.map(v => ({
...v,
papers_in_catalog: paper_counts.get(v.journal) ?? 0
}))
const table = Inputs.table(rows, {
columns: ["journal", "two_yr_citedness", "h_index", "papers_in_catalog", "year_collected"],
header: {
journal: "Venue",
two_yr_citedness: "IF₂ᵧᵣ (OpenAlex)",
h_index: "h-index",
papers_in_catalog: "Papers in catalog",
year_collected: "Year collected"
},
format: {
journal: j => page_link("venues", j, j),
two_yr_citedness: c => c == null ? "" : c.toFixed(2),
h_index: v => v == null ? "" : String(v),
year_collected: y => y == null ? "" : String(y)
},
sort: "two_yr_citedness",
reverse: true,
rows: 30,
width: { two_yr_citedness: 130, h_index: 90, papers_in_catalog: 140, year_collected: 130 }
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${table}</div>
</div>`
}Publication lifecycle
How a method goes from arXiv / bioRxiv preprint to a peer-reviewed publication. Each pairing is recorded explicitly rather than inferred from titles or dates, taken from bioRxiv’s own record of where a preprint was published, Crossref’s preprint relations, or hand-checked author overlap. Peer-reviewed, ML-conference and thesis publications all count as “post-preprint”. The Status column tells you whether each row is paired (lifecycle complete), preprint-only (still in flight), or peer-reviewed-only (published without a preprint we have on file).
lifecycle_rows = {
// Pairing comes from the explicit publication_version table, built offline
// by build_versions.py from bioRxiv's `published` field, Crossref
// is-preprint-of relations, exact normalised-title matches and
// human-reviewed fuzzy candidates.
//
// This used to be guessed here: publications were bucketed by (first
// algorithm name, version) and each preprint greedily claimed the earliest
// later peer-reviewed paper in its bucket. That read only
// `models.split(",")[0]`, so a paper linked to several algorithms was
// bucketed by whichever came first; a paper with no algorithm link could
// never pair at all; and date-order greed could hand one journal paper to
// the wrong preprint. The table also refuses that last case outright, via a
// UNIQUE index on published_id.
const POST_PREPRINT = new Set(["peer-reviewed", "ML conference", "thesis"])
// The method label the y-axis and the Method column display. Version is kept
// separate because Casanovo has versioned releases (v1 / v2 / v5) that each
// went through their own preprint-to-journal cycle.
const label_of = p => ({
method: (p.models ?? "").split(",")[0]?.trim()
|| (p.title ?? "").slice(0, 40)
|| `#${p.id}`,
version: p.version || null
})
const url_of = p => p.url || (p.doi ? `https://doi.org/${p.doi}` : null)
// Only pair when BOTH halves survive the global filters, so the chart never
// shows a pairing the reader has filtered half of away.
const visible = new Map(pubs_filtered.map(p => [p.id, p]))
const rows = []
const claimed = new Set()
for (const pp of pubs_filtered) {
if (pp.type !== "preprint") continue
const target_id = published_of.get(pp.id)
const target = target_id != null ? visible.get(target_id) : null
const { method, version } = label_of(pp)
if (target) {
claimed.add(target.id)
const gap_months = ((target.date ?? 0) - (pp.date ?? 0)) / 86400000 / 30.44
rows.push({
method, version, status: "paired",
preprint_id: pp.id, peer_id: target.id,
preprint_date: pp.date, preprint_title: pp.title, preprint_url: url_of(pp),
peer_date: target.date, peer_title: target.title, peer_url: url_of(target),
gap_months: Math.round(gap_months * 10) / 10
})
} else {
rows.push({
method, version, status: "preprint-only",
preprint_id: pp.id, peer_id: null,
preprint_date: pp.date, preprint_title: pp.title, preprint_url: url_of(pp),
peer_date: null, peer_title: null, peer_url: null, gap_months: null
})
}
}
for (const p of pubs_filtered) {
if (!POST_PREPRINT.has(p.type) || claimed.has(p.id)) continue
const { method, version } = label_of(p)
rows.push({
method, version, status: "peer-reviewed-only",
preprint_id: null, peer_id: p.id,
preprint_date: null, preprint_title: null, preprint_url: null,
peer_date: p.date, peer_title: p.title, peer_url: url_of(p), gap_months: null
})
}
// Order: paired first (shortest gap first), then preprint-only by preprint
// date (oldest first = the longest-outstanding follow-ups), then
// peer-reviewed-only by year.
const status_rank = { "paired": 0, "preprint-only": 1, "peer-reviewed-only": 2 }
return rows.sort((a, b) => {
const s = status_rank[a.status] - status_rank[b.status]
if (s !== 0) return s
if (a.status === "paired") return a.gap_months - b.gap_months
return (a.preprint_date ?? a.peer_date ?? 0) - (b.preprint_date ?? b.peer_date ?? 0)
})
}{
const table = Inputs.table(lifecycle_search, {
columns: ["method", "version", "status", "preprint_date", "peer_date", "gap_months"],
header: {
method: "Method",
version: "Version",
status: "Status",
preprint_date: "Preprint",
peer_date: "Peer-reviewed",
gap_months: "Gap (months)"
},
format: {
// Dates link to this catalog's page for that version; the publisher link
// lives on the page itself. The title stays as the hover text, which is
// what makes a bare "2025-06" cell legible.
method: m => page_links("algorithms", m, { max: 2 }),
preprint_date: (d, i) => {
const row = lifecycle_search[i]
if (!d) return ""
const label = d.toISOString().slice(0, 7)
const href = page_href("publications", row.preprint_id) ?? row.preprint_url
return href
? htl.html`<a href=${href} title=${row.preprint_title ?? ""}>${label}</a>`
: label
},
peer_date: (d, i) => {
const row = lifecycle_search[i]
if (!d) return ""
const label = d.toISOString().slice(0, 7)
const href = page_href("publications", row.peer_id) ?? row.peer_url
return href
? htl.html`<a href=${href} title=${row.peer_title ?? ""}>${label}</a>`
: label
},
gap_months: g => g == null ? "" : `${g.toFixed(1)} mo`
},
rows: 25,
width: { method: 200, version: 70, status: 160, preprint_date: 110, peer_date: 110, gap_months: 110 }
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${table}</div>
</div>`
}{
// Label = method + version (when versioned). Casanovo's three versions then
// get their own y-axis rows; everything else stays as just the method name.
const paired = lifecycle_rows
.filter(r => r.status === "paired")
.map(r => ({
...r,
label: r.version ? `${r.method} ${r.version}` : r.method
}))
if (!paired.length) {
return html`<p style="color:#57606a; font-style:italic;">No paired methods in the current filter.</p>`
}
// Two destinations, because the two elements mean different things: the tick
// label is the METHOD, and the bar is the specific peer-reviewed paper whose
// arrival the gap measures. `method` falls back to a truncated title for a
// publication with no algorithm link, in which case page_href returns null
// and the label simply stays plain text.
const href_by_label = new Map(
paired.map(r => [r.label, page_href("algorithms", r.method)])
.filter(([, href]) => href)
)
const chart = Plot.plot({
marginLeft: 270, // measured 64 px of method-label overflow
height: Math.max(220, paired.length * 20),
marginBottom: 48,
x: { label: "Preprint → peer-reviewed gap (months)", grid: true, labelOffset: 40 },
y: { label: null },
marks: [
Plot.barX(paired, {
x1: 0,
x2: "gap_months", // explicit x1/x2 bypasses Plot's auto stackX transform
y: "label",
fill: "#1f6feb",
sort: { y: "x2", reverse: true },
// Plot's tip labels the x channel with the SCALE's label and truncates
// each line at lineWidth. "Preprint -> peer-reviewed gap (months)" is 37
// characters, so the line overflowed and the ellipsis ate the number
// itself, leaving a tooltip that showed no gap at all. Give the tip its
// own short channel and hide the raw endpoints.
// y has no scale label, so Plot fell back to the field name and the tip
// read "label DiffuNovo". Name both channels explicitly.
channels: { "Method": "label", "Gap (months)": "gap_months" },
tip: { ...plot_tip_style, format: { x1: false, x2: false, y: false } },
href: d => page_href("publications", d.peer_id),
target: "_self"
}),
Plot.ruleX([0])
]
})
return linkify_tick_labels(chart, href_by_label)
}Browse all papers
Recently added
md`What changed since you last looked. The catalog stores no "date added"
column, but \`publication.id\` is handed out in insertion order, so ordering by
it descending IS the order things arrived. The
${recent_pubs.length} newest are below, and that same id is the sortable **#**
column in [the full table](#every-paper), which is the way to see further back.
The date beside each one is the paper's own publication date, not when it was
catalogued: ${recent_pubs.filter(d => d.date && d.date > new Date(Date.now() - 200 * 864e5)).length}
of these ${recent_pubs.length} appeared in the last six months, and the rest are
older work that turned up in a citation sweep.`// Three, not ten. This list sits directly under the 'Browse all papers'
// heading, so a reader following a link to the papers section used to land on a
// screenful of the newest arrivals rather than on the table they asked for.
// Every link that means "the table" now points at #every-paper, and this list
// is short enough not to be in the way. The count is read back out of this
// array by the prose above, so changing the 3 cannot leave the text lying.
recent_pubs = pubs_t.slice().sort((a, b) => d3.descending(+a.id, +b.id)).slice(0, 3)html`<ol class="recent-list" start="1">${recent_pubs.map(d => {
const href = page_href("publications", d.id)
const title = href ? htl.html`<a href=${href}>${d.title}</a>` : d.title
const doi = d.url || (d.doi ? `https://doi.org/${d.doi}` : null)
const models = String(d.models ?? "").split(",").map(x => x.trim()).filter(Boolean)
return htl.html`<li><span class="recent-num">#${d.id}</span> ${title}
${doi ? htl.html` <a href=${doi} target="_blank" rel="noopener"
title="Publisher / DOI" style="text-decoration:none">↗</a>` : ""}
<span class="recent-meta">${[fmt_day(d.date), d.journal, d.type]
.filter(Boolean).join(" · ")}</span>
${models.length ? htl.html`<span class="recent-meta"> — ${models.map((m, i) =>
htl.html`${i ? ", " : ""}${page_link("algorithms", m, m)}`)}</span>` : ""}
</li>`
})}</ol>`Every paper
viewof selected_pubs = {
// Decorate each row with the venue's OpenAlex 2-year citedness (or null).
// Also rewrite the comma-separated `models` string to append known aliases
// (e.g. π-HelixNovo → "π-HelixNovo (aka PandaNovo)") so renamed methods stay
// searchable by their historical name in the table's filter box.
const aliases_by_model = new Map(
algorithms_t.filter(a => a.aliases).map(a => [a.model, a.aliases])
)
const annotate_models = s => (s ?? "")
.split(",").map(x => x.trim()).filter(Boolean)
.map(m => aliases_by_model.has(m) ? `${m} (aka ${aliases_by_model.get(m)})` : m)
.join(", ")
const search_with_if = search.map(r => ({
...r,
models: annotate_models(r.models),
// The full date as an ISO STRING, not the Date object and not the year. It
// is what the table sorts on, and a string keeps three things working at
// once: ISO dates sort lexicographically, the search box can match
// "2026-09", and the cell needs no formatter.
date_str: fmt_day(r.date),
citedness: journal_impact_by_name.get(r.journal)?.two_yr_citedness ?? null
}))
const table = Inputs.table(search_with_if, {
// `id` is the insertion order, so sorting on it descending answers "what
// was added last", which the Date column cannot: an older paper added today
// sorts by its own date, wherever that falls.
columns: ["id", "date_str", "models", "version", "kind", "is_dl", "acquisition", "title", "authors", "journal", "citedness", "type", "repo"],
header: {
id: "#",
date_str: "Date", models: "Method(s)", version: "Ver.", kind: "Kind", is_dl: "DL?", acquisition: "Acq.",
title: "Title", authors: "Authors", journal: "Venue", citedness: "IF₂ᵧᵣ", type: "Type", repo: "Code"
},
format: {
is_dl: v => v === 1 || v === true ? "DL" : (v === 0 || v === false ? "classical" : ""),
citedness: c => c == null ? "" : c.toFixed(1),
// The title now goes to this catalog's own page for the paper, with the
// publisher/DOI kept as a separate chip so both destinations stay
// reachable. The row index lookup is load-bearing: 19 titles are shared
// by a preprint and its journal version, so the title string alone cannot
// identify the row.
title: (t, i) => {
const row = search_with_if[i]
const internal = page_href("publications", row?.id)
const external = row?.url || (row?.doi ? `https://doi.org/${row.doi}` : null)
const label = internal ? htl.html`<a href=${internal}>${t}</a>` : t
return external
? htl.html`${label} <a href="${external}" target="_blank" rel="noopener"
title="Publisher / DOI" style="text-decoration:none">↗</a>`
: label
},
// annotate_models() appended "(aka ...)" to renamed tools. Keep that
// visible but link only the name part, which is the slug key.
models: m => {
const parts = String(m ?? "").split(",").map(x => x.trim()).filter(Boolean)
const out = []
parts.slice(0, 4).forEach((part, idx) => {
if (idx) out.push(", ")
const alias = part.match(/^(.*?)(\s*\(aka [^)]*\))$/)
const name = alias ? alias[1] : part
out.push(page_link("algorithms", name, name))
if (alias) out.push(alias[2])
})
if (parts.length > 4) out.push(` +${parts.length - 4}`)
return htl.html`${out}`
},
// Three names then a count: a 53-author consortium paper would otherwise
// fill the row with links.
authors: a => page_links("authors", a, { max: 3 }),
journal: j => page_link("venues", j, j),
repo: r => {
if (!r) return ""
// The repository column holds either a single URL or two whitespace-separated URLs
// (a few RNovA-style entries). Render up to three short link chips.
const urls = String(r).split(/\s+/).filter(s => /^https?:\/\//.test(s)).slice(0, 3)
if (!urls.length) return ""
return htl.html`${urls.map(u => htl.html`<a href="${u}" target="_blank" rel="noopener" title="${u}" style="margin-right:4px">↗</a>`)}`
}
},
// On the DATE, not the year. Sorting on a year descending leaves the rows
// within each year in the array's own order, which is the SQL's
// `ORDER BY p.publication_date` -- oldest first. So the table opened on
// January 2026, and a paper added today landed sixty rows below it in a
// column that said 2026 all the way down.
sort: "date_str",
reverse: true,
rows: 25,
width: { id: 50, date_str: 96, type: 100, models: 130, version: 60, kind: 130, is_dl: 70, acquisition: 70, citedness: 70, repo: 60 }
})
// Reflect the inner Inputs.table's value/input on the fullscreen wrapper so
// `viewof selected_pubs` exposes the rows the user has ticked, used by the
// BibTeX download button below.
const wrapper = html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${table}</div>
</div>`
Object.defineProperty(wrapper, "value", { get: () => table.value })
table.addEventListener("input", e => wrapper.dispatchEvent(new CustomEvent("input", { bubbles: false })))
return wrapper
}{
// Format one publication as a BibTeX entry. Entry type picked from
// publication_type; key is a stable first-author-lastname + year + title-word slug.
const slug = s => (s ?? "").toString().toLowerCase().replace(/[^a-z0-9]+/g, "")
const escape = s => (s ?? "").toString().replace(/[{}\\]/g, "\\$&")
const entry_type_of = t => ({
"ML conference": "inproceedings",
"thesis": "phdthesis",
"preprint": "misc",
"postprint": "misc",
"commentary": "misc",
// A published meeting or showcase abstract is not an article: no full text
// exists behind it, so @misc plus an explicit note is the honest form.
"abstract": "misc",
// A talk, poster or recording. Same reasoning as 'abstract': there is no
// article, and @misc with a note says what the thing is instead of dressing
// a conference poster up as a journal paper.
"presentation": "misc"
})[t] ?? "article"
const to_bibtex = p => {
const author_list = (p.authors ?? "").split(",").map(s => s.trim()).filter(Boolean)
const first_last = author_list[0]?.split(/\s+/).pop() ?? "anon"
const title_word = (p.title ?? "ref").split(/\s+/).find(w => w.length > 3) ?? "ref"
const key = `${slug(first_last)}${p.year ?? ""}${slug(title_word)}`
const fields = []
if (p.title) fields.push(` title = {${escape(p.title)}}`)
if (author_list.length) fields.push(` author = {${author_list.map(escape).join(" and ")}}`)
if (p.year) fields.push(` year = {${p.year}}`)
if (p.journal) fields.push(` journal = {${escape(p.journal)}}`)
if (p.doi) fields.push(` doi = {${escape(p.doi)}}`)
if (p.url) fields.push(` url = {${escape(p.url)}}`)
if (p.type === "preprint" || p.type === "postprint" || p.type === "abstract"
|| p.type === "presentation")
fields.push(` note = {${p.type}}`)
return `@${entry_type_of(p.type)}{${key},\n${fields.join(",\n")}\n}`
}
const handle_download = () => {
if (!selected_pubs.length) return
const text = selected_pubs.map(to_bibtex).join("\n\n") + "\n"
const blob = new Blob([text], { type: "application/x-bibtex;charset=utf-8" })
const a = document.createElement("a")
a.href = URL.createObjectURL(blob)
a.download = `de-novo-papers-${new Date().toISOString().slice(0,10)}.bib`
document.body.appendChild(a); a.click(); a.remove()
URL.revokeObjectURL(a.href)
}
const disabled = selected_pubs.length === 0
const btn = html`<button class="download-btn" ?disabled=${disabled}>
⬇ Download ${selected_pubs.length || "0"} selected as BibTeX
</button>`
btn.disabled = disabled
btn.onclick = handle_download
return html`<div style="margin: 0.5rem 0 1rem; display: flex; gap: 0.75rem; align-items: center;">
${btn}
<span style="color:#57606a; font-size: 0.9em;">Tick rows in the table above to enable.</span>
</div>`
}Browse all authors
Aggregated from the currently-filtered set of papers. Searching is case-insensitive across every column (name, affiliation, country, methods).
authors_table_rows = {
// Roll up authors across the filtered pubs, tagging each with the
// model names and kinds they touched in that subset.
const detail = new Map(author_details_t.map(d => [d.name, d]))
const counts = new Map() // name → { papers, models:Set, kinds:Set }
for (const p of pubs_filtered) {
const ms = (p.models ?? "").split(", ").filter(Boolean)
for (const a of (p.authors ?? "").split(", ").filter(Boolean)) {
const slot = counts.get(a) ?? { papers: 0, models: new Set(), kinds: new Set() }
slot.papers += 1
for (const m of ms) slot.models.add(m)
if (p.kind) slot.kinds.add(p.kind)
counts.set(a, slot)
}
}
return Array.from(counts, ([name, v]) => ({
name,
papers: v.papers,
methods: Array.from(v.models).sort().join(", "),
kinds: Array.from(v.kinds).sort().join(", "),
affiliations: detail.get(name)?.affiliations ?? "",
countries: detail.get(name)?.countries ?? ""
})).sort((a, b) => b.papers - a.papers || a.name.localeCompare(b.name))
}{
const table = Inputs.table(author_search, {
columns: ["name", "papers", "methods", "kinds", "affiliations", "countries"],
header: {
name: "Author", papers: "Papers", methods: "Method(s)",
kinds: "Kind(s)", affiliations: "Affiliation(s)", countries: "Country / countries"
},
format: {
papers: n => n == null ? "" : String(n),
name: n => page_link("authors", n, n),
// affiliations is a "Institution / Department" list joined by "; ".
// Only the institution half has a page, so link that and keep the
// department as plain text.
affiliations: s => {
if (!s) return ""
const parts = String(s).split(";").map(x => x.trim()).filter(Boolean)
const out = []
parts.slice(0, 2).forEach((part, idx) => {
if (idx) out.push("; ")
const cut = part.indexOf(" / ")
const inst = cut === -1 ? part : part.slice(0, cut)
out.push(page_link("institutions", inst, inst))
if (cut !== -1) out.push(part.slice(cut))
})
if (parts.length > 2) out.push(` +${parts.length - 2}`)
return htl.html`${out}`
},
methods: s => page_links("algorithms", s, { max: 4 })
},
sort: "papers",
reverse: true,
rows: 25,
width: { name: 180, papers: 70, methods: 220, kinds: 130, affiliations: 320, countries: 120 }
})
return html`<div class="chart-wrap">
<button class="fs-btn" onclick="
const el = this.parentElement;
if (document.fullscreenElement) document.exitFullscreen();
else el.requestFullscreen();
">⛶ Fullscreen</button>
<div class="chart-scroll">${table}</div>
</div>`
}Contributing
Easiest path: open a GitHub issue with a link to the paper (DOI / arXiv / bioRxiv / OpenReview / …) and I’ll wire it into the database. Corrections are equally welcome: wrong author lists, missing affiliations, mis-classified kind / DL / acquisition, broken hyperlinks, anything that looks off.
Advanced: edit the database directly
The site is generated from denovo.db (SQLite, the source of truth). If you’re comfortable with SQL:
Edit
denovo.dbwith any SQLite tool (sqlite3CLI, DB Browser for SQLite, DataGrip, …). A new paper typically needs rows inpublication,publication_author, andpublication_algorithm(setroletousesfor a tool the paper runs rather than introduces, which keeps it out of that tool’s authors and papers); a new model also needs a row inalgorithm(setkind,is_deep_learning,acquisition_mode). Affiliations cascade throughcountry → city → affiliationand link to authors viaauthor_affiliation.Regenerate the human-readable SQL dump so the diff is reviewable:
Open a PR with both
denovo.dbanddenovo.sql. The GitHub Action rebuilds the site and publishes togh-pageson merge; typically live within ~3 minutes.
Cite this catalog
If you use this catalog, please cite it as:
Van Goey, J. Awesome De Novo Peptide Sequencing. Zenodo. https://doi.org/10.5281/zenodo.20825737
Machine-readable metadata is in CITATION.cff; GitHub’s “Cite this repository” button exports BibTeX / APA.
This page is a comprehensive map of de novo peptide sequencing covering algorithms, post-processors, downstream applications and adjacent tools, deep-learning and classical alike. Source data and code: GitHub, rebuilt automatically on every push to main.