The audit ran for four months. Two coders, an eight-dimension rubric
derived from Terry et al. (2024) and Shen et al. (2024), sixteen benchmarks
picked from the most-cited evaluations in the alignment literature. The
instrument scores what each benchmark operationalises
as a scored property, not how well any model performs on it. Eight
dimensions, sixteen benchmarks, one hundred and twenty-eight cells, two
independent codings each. Final inter-rater κ of 0.87, near-perfect
agreement. The reconciled disagreement log is public.
Four findings emerged.
One. Verification support (D3, "does the model help the
user, not the benchmark's evaluator, verify the
answer?") gets a non-zero score on zero benchmarks. Not low scores. Not
scattered partial coverage. Zero. Across the most widely cited alignment
evaluations, the question of whether a model produces calibrated uncertainty
signals or evidence-pointing scaffolds that help an end-user judge an answer
is not a scored target anywhere. This is the canonical absent dimension.
It is also, arguably, the dimension that matters most for actual deployment.
Two. Process steerability (D2, "can the user redirect the model's
approach?") is nearly absent. tau-bench scores 1; almost everything else
scores 0. The benchmarks measure outputs, not the processes that produce
outputs, even when the deployment story turns on whether a user can intervene
mid-process.
Three. The four benchmarks nominally intended as interactional,
CURATe, MT-Bench, Common Ground, tau-bench, each spike on different
dimensions and overlap less than the field's casual citation pattern would
suggest. They are not a coherent suite. Combining their scores does not
produce a "covered" interactional tier; it produces a fragmented mosaic
with most dimensions still poorly scored.
Four, and most uncomfortable. The construction of the evaluation,
not the source of the data, determines what gets measured. WildBench
(Lin et al., 2024) draws on more than a million real user logs and still
scores zero on five of eight interactional dimensions. Real user data does
not automatically yield interactional measurement. What yields interactional
measurement is choosing to score interactional properties; the data source
is downstream of that choice. The opposite belief, that we can fix
evaluation by adding real-world data, has been a comfortable story in the
field for years. The audit refutes it directly.
The audit's instrument and full coding manual are now public at
github.com/varad-vishwarupe/alignment-benchmark-audit. The rubric is
separately released at github.com/varad-vishwarupe/interactional-benchmark-rubric
so other researchers can apply it to benchmarks we did not score. The
disagreement log is included with the resolution rationale for every
contested cell. If you want to replicate the audit, extend it, or argue
with it, the artefacts are there.
The closing claim: reproducibility for
evaluation work is not a virtue; it is a precondition for the work being
scientific at all. We released the audit's artefacts as T1 (fully public)
because anything less would have made the position paper's argument
self-undermining. The companion preprint on reproducibility standards for
frontier safety claims (Vishwarupe et al., 2026) generalises this principle:
evaluation work that cannot be reviewed at the artefact level cannot
establish the claims its authors want it to establish.
References (selected):
Terry, M., et al. (2024). AI alignment: a comprehensive survey;
Shen, T., et al. (2024). Towards bidirectional human-AI alignment;
Lin, B. Y., et al. (2024). WildBench;
Vishwarupe, V., et al. (2026, preprint). Deployment-relevant alignment
cannot be inferred from model-level audits.