A Claude Code skill that turns a maintained GitHub repo into an NKS-built roadmap. Benchmark repo: karakeep-app/karakeep (iters 1–8); generalization on navidrome/navidrome and Kareadita/Kavita. Feature arcs: iter-5 model the product (+product-mastery axis), iter-6 leverage the graph (kartas, figure-on-ground, tensions; +methodology-leverage), iter-7 close the in-moment advantages + add a read-only graph audit, iter-8 fix the two defects the audit caught — converging at 5.0 / 4.71 / 4.6 with the audit finding no theater. Two Claude judges + a Codex judge (decorrelation anchor) score each roadmap 1–5.
| Iter | Judges | Outcome |
|---|---|---|
| jitsi · multi-repo · real-time-media · third-party | 5.0 / 4.86 · 4.14 multirepo_coherence 5/5/5 · m_leverage 5/5/5 · graph-audit 5/5 | Second third-party generalization, a DIFFERENT shape. A real-time WebRTC product — Jitsi Meet — across 4 repos (client + client-lib + signaling focus + SFU media bridge), with NO milestones and a signaling+media seam. Generalization held AND the PR-state fix landed: every cited PR re-verified individually (no open-as-merged, no rebase-a-mergeable). multirepo_coherence 5/5/5 · methodology_leverage 5/5/5 · graph-audit 5/5 · multi-repo · real-time-media · no-milestones · third-party |
| appflowy · multi-repo · third-party | 4.0 / 4.71 · 3.7 multirepo_coherence 5/5/5 · m_leverage 5/5/5 · graph-audit 5/5 | Third-party generalization. The now-native multi-repo skill on someone else's PUBLIC product — AppFlowy (Notion alternative) across 4 repos (Flutter client + Rust cloud + TS web + CRDT lib) with a REAL backlog (~1000 issues). Native multi-repo + graph leverage held (multirepo_coherence 5/5/5 · methodology_leverage 5/5/5 · graph-audit 5/5); the rich backlog exposed a PR-state-verification defect. multi-repo · real-backlog · third-party · graph-view. |
| knesset · multi-repo | 5.0 / 5.0 / 4.0 multirepo_coherence 5/5/4 · m_leverage 5/5/5 · audit 5/5 | First MULTI-REPO run. ONE product (KnessetVotes — Open Knesset) across 4 private repos (data-pipelines → backend → frontend → ops) with a genuinely EMPTY backlog — 0 issues / PRs / milestones / releases. The roadmap is carried entirely by the product-model ground + the 107-PR merged trajectory; all 5 directions are cross-repo. The graph audit finds no theater (5/5). multi-repo · empty-backlog · blind. |
| 1 | 4.3 / 4.67 | Strong, grounded — but omitted the committed milestone 0.33.0 (the one real defect). |
| 2 | — | Failed: transport drop on a large agent return → no artifact. Drove the harness fix (disk hand-offs, short returns). |
| 3 | 4.4 / 4.7 | Milestone 0.33.0 closed as the lead committed tier. New defect: only 5/15 milestone issues + over-claimed completeness. |
| 4 | 4.6 / 4.7 | Plateau: 15/15 milestone issues, no over-claim, maintainer items present. |
| navidrome | 4.67 / 4.83 / 4.17 | Generalization (fresh repo, thematic milestones) + Codex judge. Codex caught a real state error both Claude judges missed → state-reverify fix. |
| kavita | 5.0 / 5.0 / 3.83 | 3rd repo (C#, issue-heavy/PR-light). state-reverify validated. Claude ceiling at 5.0; Codex (3.83) still finds gaps. |
| iter-5 · product-model | 4.71 / 4.43 / 3.6 | "Model the product." A verified vartamana ground under the backlog; leads with "what this product is today" + an estafeta trace. product_mastery 5/5/4. |
| iter-6 · leverage the graph | 5.0 / 4.71 / 3.71 m_leverage 5 / 5 / 3 | "Leverage the graph." Kartas, an ahara figure-on-ground table, a tensions → structural-risks section. Both Claude judges: "could not have been a flat list" — but judged from text alone, which over-credited the leverage (iter-7's audit shows why). |
| iter-7 · in-moment + graph-audit | 4.9 / 4.43 / 4.14 m_leverage 4 / 4 / 4 | Closes the in-moment advantages (anga, ahara/utpatti estafeta, in-moment modus, signal-audit table, author rigor) and adds a read-only graph audit that separates real leverage from prose "theater". Codex hits 4.14; the audit catches one real piece of theater (the maintainer "karta" is a zero-edge orphan) and de-inflates iter-6's 5/5 to an honest 4/4/4. |
| iter-8 · wired kartas + audit-clean | 5.0 / 4.71 / 4.6 m_leverage 5 / 5 / 5 · audit 5/5 | Fixes the two defects the iter-7 audit caught: the maintainer "karta" is now genuinely wired (22 actor edges to the work it drives, not an orphan), and figure-on-ground uses graph-legal links (carrier-kriya → upadhi/specifies → capability, since a literal ahara from a bianhua is graph-invalid). Plus mergeable=null handling and full Backlog enumeration. The graph audit, which caught real theater in iter-7, now finds none and scores leverage 5/5 — methodology_leverage 5/5/5 is honest, not inflated. Codex reaches 4.6 — the highest karakeep score across the whole project — and the spread (5.0/4.71/4.6) is the tightest: the convergence point. Remaining items are diverging nitpicks (PR readiness columns, a rank-15 tie-break). |
loading skill.md…
loading iterations.md…
loading handoff…