repo-roadmap — benchmark iterations

A Claude Code skill that turns a maintained GitHub repo into an NKS-built roadmap. Benchmark repo: karakeep-app/karakeep (iters 1–8); generalization on navidrome/navidrome and Kareadita/Kavita. Feature arcs: iter-5 model the product (+product-mastery axis), iter-6 leverage the graph (kartas, figure-on-ground, tensions; +methodology-leverage), iter-7 close the in-moment advantages + add a read-only graph audit, iter-8 fix the two defects the audit caught — converging at 5.0 / 4.71 / 4.6 with the audit finding no theater. Two Claude judges + a Codex judge (decorrelation anchor) score each roadmap 1–5.

Scores by iteration

IterJudgesOutcome
jitsi · multi-repo · real-time-media · third-party5.0 / 4.86 · 4.14
multirepo_coherence 5/5/5 · m_leverage 5/5/5 · graph-audit 5/5
Second third-party generalization, a DIFFERENT shape. A real-time WebRTC product — Jitsi Meet — across 4 repos (client + client-lib + signaling focus + SFU media bridge), with NO milestones and a signaling+media seam. Generalization held AND the PR-state fix landed: every cited PR re-verified individually (no open-as-merged, no rebase-a-mergeable). multirepo_coherence 5/5/5 · methodology_leverage 5/5/5 · graph-audit 5/5 · multi-repo · real-time-media · no-milestones · third-party
appflowy · multi-repo · third-party4.0 / 4.71 · 3.7
multirepo_coherence 5/5/5 · m_leverage 5/5/5 · graph-audit 5/5
Third-party generalization. The now-native multi-repo skill on someone else's PUBLIC product — AppFlowy (Notion alternative) across 4 repos (Flutter client + Rust cloud + TS web + CRDT lib) with a REAL backlog (~1000 issues). Native multi-repo + graph leverage held (multirepo_coherence 5/5/5 · methodology_leverage 5/5/5 · graph-audit 5/5); the rich backlog exposed a PR-state-verification defect. multi-repo · real-backlog · third-party · graph-view.
knesset · multi-repo5.0 / 5.0 / 4.0
multirepo_coherence 5/5/4 · m_leverage 5/5/5 · audit 5/5
First MULTI-REPO run. ONE product (KnessetVotes — Open Knesset) across 4 private repos (data-pipelines → backend → frontend → ops) with a genuinely EMPTY backlog — 0 issues / PRs / milestones / releases. The roadmap is carried entirely by the product-model ground + the 107-PR merged trajectory; all 5 directions are cross-repo. The graph audit finds no theater (5/5). multi-repo · empty-backlog · blind.
14.3 / 4.67Strong, grounded — but omitted the committed milestone 0.33.0 (the one real defect).
2Failed: transport drop on a large agent return → no artifact. Drove the harness fix (disk hand-offs, short returns).
34.4 / 4.7Milestone 0.33.0 closed as the lead committed tier. New defect: only 5/15 milestone issues + over-claimed completeness.
44.6 / 4.7Plateau: 15/15 milestone issues, no over-claim, maintainer items present.
navidrome4.67 / 4.83 / 4.17Generalization (fresh repo, thematic milestones) + Codex judge. Codex caught a real state error both Claude judges missed → state-reverify fix.
kavita5.0 / 5.0 / 3.833rd repo (C#, issue-heavy/PR-light). state-reverify validated. Claude ceiling at 5.0; Codex (3.83) still finds gaps.
iter-5 · product-model4.71 / 4.43 / 3.6"Model the product." A verified vartamana ground under the backlog; leads with "what this product is today" + an estafeta trace. product_mastery 5/5/4.
iter-6 · leverage the graph5.0 / 4.71 / 3.71
m_leverage 5 / 5 / 3
"Leverage the graph." Kartas, an ahara figure-on-ground table, a tensions → structural-risks section. Both Claude judges: "could not have been a flat list" — but judged from text alone, which over-credited the leverage (iter-7's audit shows why).
iter-7 · in-moment + graph-audit4.9 / 4.43 / 4.14
m_leverage 4 / 4 / 4
Closes the in-moment advantages (anga, ahara/utpatti estafeta, in-moment modus, signal-audit table, author rigor) and adds a read-only graph audit that separates real leverage from prose "theater". Codex hits 4.14; the audit catches one real piece of theater (the maintainer "karta" is a zero-edge orphan) and de-inflates iter-6's 5/5 to an honest 4/4/4.
iter-8 · wired kartas + audit-clean5.0 / 4.71 / 4.6
m_leverage 5 / 5 / 5 · audit 5/5
Fixes the two defects the iter-7 audit caught: the maintainer "karta" is now genuinely wired (22 actor edges to the work it drives, not an orphan), and figure-on-ground uses graph-legal links (carrier-kriya → upadhi/specifies → capability, since a literal ahara from a bianhua is graph-invalid). Plus mergeable=null handling and full Backlog enumeration. The graph audit, which caught real theater in iter-7, now finds none and scores leverage 5/5 — methodology_leverage 5/5/5 is honest, not inflated. Codex reaches 4.6 — the highest karakeep score across the whole project — and the spread (5.0/4.71/4.6) is the tightest: the convergence point. Remaining items are diverging nitpicks (PR readiness columns, a rank-15 tie-break).

Roadmap artifacts

Jitsi Meet — video conferencing (multi-repo · third-party)

5.0 / 4.86 · 4.14
Second third-party generalization, a DIFFERENT shape — a real-time WebRTC product across 4 repos (client + client-lib + signaling focus + SFU media bridge), NO milestones, signaling+media seam. Generalization held AND the PR-state fix landed: every cited PR re-verified individually (no open-as-merged, no rebase-a-mergeable). multirepo_coherence 5/5/5 · methodology_leverage 5/5/5 · graph-audit 5/5 · multi-repo · real-time-media · no-milestones · third-partylatest

AppFlowy — Notion alternative (multi-repo · third-party)

4.0 / 4.71 · 3.7
Third-party generalization — the now-native multi-repo skill on someone else's PUBLIC product (4 repos: Flutter client + Rust cloud + TS web + CRDT lib) with a REAL backlog (~1000 issues). Native multi-repo + graph leverage held; the rich backlog exposed a PR-state-verification defect. multirepo_coherence 5/5/5 · methodology_leverage 5/5/5 · graph-audit 5/5 · multi-repo · real-backlog · third-party · graph-view

KnessetVotes — Open Knesset (multi-repo · iter-1)

5.0 / 5.0 · 4.0
First MULTI-REPO run — ONE product across 4 private repos (data-pipelines → backend → frontend → ops) with a genuinely EMPTY backlog (0 issues/PRs/milestones/releases); the roadmap is carried entirely by the product-model ground + 107-PR merged trajectory. 5 directions, ALL cross-repo. multirepo_coherence 5/5/4 · methodology_leverage 5/5/5 · graph-audit 5/5 (no theater) · multi-repo · empty-backlog · blind

Iteration 1

4.3 / 4.67
milestone omitted

Iteration 2

failed
no artifact (transport)

Iteration 3

4.4 / 4.7
milestone closed

Iteration 4

4.6 / 4.7
karakeep plateau

Navidrome

4.67 / 4.83 / 4.17
generalization · +Codex

Kavita

5.0 / 5.0 / 3.83
3rd repo · state-fix validated

iter-5 · product model

4.71 / 4.43 / 3.6
figure-on-ground · +product_mastery

iter-6 · leverage the graph

5.0 / 4.71 / 3.71
kartas + tensions · m_leverage 5/5/3

iter-7 · in-moment + audit

4.9 / 4.43 / 4.14
graph-audit catches theater · m_leverage 4/4/4

iter-8 · wired + audit-clean

5.0 / 4.71 / 4.6
karakeep · kartas wired · graph-legal figure-on-ground · audit: no theater · m_leverage 5/5/5

Roadmap template · live

5.0 / 4.71 / 4.6
the iter-8 roadmap rendered through the data-driven HTML template the skill now ships — working issue/PR links, status + PR-readiness badges, collapsible cards, status filter, light/dark

Skill (current) skills/repo-roadmap/SKILL.md

loading skill.md…

Skill iteration log

loading iterations.md…

Handoff (next session)

loading handoff…