Agent Evaluation Reliability: More Tasks Won't Always Fix a Leaderboard
Leaderboard numbers look crisp: model A beats B by three points, a system card says “leading,” and the same table shows up in policy memos. The hard question is what those numbers actually support. Are they saying the deployed system is more reliable, that the underlying model is stronger, or that an absolute score can serve as a capability threshold? Three claims often share one board and need different measurement goals. A score does not announce which claim it endorses.
Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, and Sanmi Koyejo (Stanford and UIUC) unpack that ambiguity in arXiv:2610.00651 (Agent Evaluation Reliability: More Tasks Won’t (Always) Fix An Agent Leaderboard, preprint 2026-09-30). Reliability is claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models or produce reliable absolute scores. They are not asking whether a fixed agent is stable across two runs. They ask which conclusions current agent evaluations can support, and what additional evaluation would actually strengthen those conclusions.[1]
Methodologically, they build a Bayesian variance-decomposition framework from Generalizability theory (G-theory) for the sparse, imbalanced designs common in agent leaderboards, and apply it to 22 benchmarks from the Holistic Agent Leaderboard (HAL) and the Harbor Index (9 HAL + 13 Harbor). Four practical results are worth parking up front: fixed model–scaffold systems (the surrounding orchestration stack; the paper treats “scaffold,” “harness,” and sometimes even “agent” as near-synonyms) rank reliably at Eρ² = 0.935–0.994; underlying-model ranking reliability is substantially lower at 0.148–0.841; even infinitely many similarly constructed tasks improve model-ranking reliability by at most about 0.097 when uncertainty is dominated by limited scaffold coverage; and pooling diverse benchmarks raises projected reliability from about 0.44 to 0.75 at the same task budget, with projected cost cuts up to about 83%.[1]
On this site, adjacent evaluation threads already include CIR (treat harness refresh as a causal decision; separate rescue from harm), SAGE (statistical acceptance gates for self-evolving skills), Bad Genius (counterfactual signals for harness evolution), and a lighter crosslink to MoMHa (accuracy / safety / token search). This piece sits one level upstream: which claim a leaderboard number supports, and which uncertainty your budget should buy down—instead of assuming “more tasks” is always the fix.[2]
Claim dependence: system ranking ≠ model ranking ≠ absolute score
In agent evaluations, models rarely score alone. A scaffold manages state, exposes tools, and turns model outputs into actions; model plus scaffold is the deployable system. The paper’s “model” also includes a reasoning-effort configuration—e.g., gpt-5-high versus gpt-5-minimal count as distinct objects, matching the primary dataset. That creates an immediate measurement choice: are you evaluating the complete system as shipped, or isolating differences attributable to the underlying model? Change the answer, and what counts as signal versus evaluation error changes with it.[1]
Reliability also depends on the conclusion you want. A practitioner may care whether a model ranking stays stable after conditions change, whether an absolute score would look similar, or whether performance generalizes across different agentic tasks. Those are different goals and need not share one reliability number. Prior agent-reliability work mostly studies whether a fixed agent behaves consistently under repeated runs, prompt perturbations, tool configurations, or environmental failures—system robustness. Hardy et al. study the reliability of comparative inference: whether the reported ordering of models or systems would persist under new tasks, scaffolds, or benchmarks. The distinction matters. A system can be repeatable inside one harness while its rank is highly contingent on that harness.[1]
In their notation, the relative generalizability coefficient is
[ E\rho^2_o=\frac{\sigma_o^2}{\sigma_o^2+\sigma_\delta^2},\qquad \mathrm{SNR}o=\frac{\sigma_o^2}{\sigma\delta^2}=\frac{E\rho^2_o}{1-E\rho^2_o}. ]
Here (\sigma_o^2) is variance in scores that is expected to persist across the universe the claim cares about; (\sigma_\delta^2) collects variation that can still change relative standing but is irrelevant to that claim. Reliability rises when signal dominates that noise; under the usual setup it is also (\mathrm{Corr}^2(X_o,Z_o)) between the observed score and the universe score. If the object (o) is a model–scaffold pair, scaffold differences and compatibility can count as signal. If the object is the underlying model, the same differences become error—because ranks should not flip when you swap a reasonable scaffold if you are claiming “the model itself is stronger.” A footnote and appendix add a further warning: even when rankings are reliable, score reliability on the observed scale is substantially lower. Stable ranks do not license absolute scores as capability thresholds. When decisions need numeric cutoffs, use absolute reliability (dependability, (\Phi)) separately; do not treat ranking (E\rho^2) as a certificate for score levels.[1]
The practical read is blunt. On a HAL-style board with multiple scaffolds and benchmarks, ask first what the table claims to rank. If a system card quietly rewrites “model X leads” as “model X + harness Y leads,” the claim has already changed. Conversely, reporting one generic “reliability” without naming the object and generalization universe invites readers to stack three claims on one number. Designs that start as “prove the model is stronger” and end as “we scored higher in this scaffolding” often only support the second sentence.[1]
G-theory intuition: signal, noise, and the facets that move ranks
G-theory treats evaluation conditions as measurement facets and decomposes score variation into main effects and interactions: models, tasks, scaffolds, benchmarks, and their interactions. In plain terms, the framework separates signal—performance differences that persist across the conditions you claim to generalize over—from noise—variation irrelevant to the claim that can still change scores or rankings. Reliability improves when signal dominates that noise. The crucial first choice is the object of measurement: model–scaffold compatibility is signal when selecting a deployable system, and error when ranking models independently of scaffolding. Reliability therefore belongs to an object, a generalization universe, and an evaluation design—not to a benchmark alone.[1]
They fit Bayesian Bernoulli–logit mixed models directly to binary task outcomes, avoiding Gaussian approximations that mislead under floor or ceiling effects; variance components and reliability coefficients live on a shared latent log-odds scale. A benchmark-level decomposition separates persistent model differences, task sensitivity, scaffold differences, and model–scaffold compatibility. For relative rankings, a scaffold main effect that shifts every model equally does not change order, so it does not enter the error denominator. Model×scaffold interactions do: they encode which models benefit from which scaffolds. That is why system ranking and model ranking can diverge sharply—the former counts compatibility as signal; the latter requires differences that survive averaging over scaffolds. When the same (task, model, scaffold) cell is rarely repeated, the three-way interaction is not separately identifiable from latent residual variation; the terminal residual uses the logistic residual variance (\pi^2/3). That is an identifiability cost of sparse designs, not a footnote to ignore.[1]
To ask whether scaffold choice preserves model order, they define inter-scaffold reliability: whether two independently sampled scaffolds, evaluating the same models on the same tasks, induce similar rankings—analogous to inter-rater reliability, with scaffolds as alternative measurement procedures. A low value means that changing only the scaffold can change which model looks strongest, even when models and tasks stay fixed.[1]
A joint model across benchmarks separates benchmark-conditioned capability from the model component expected to persist across sampled benchmarks, tasks, and scaffolds. Proposition 3.1 is dry and useful: holding other facet sample sizes fixed, as the number of tasks (n_i\to\infty), task-indexed error terms vanish, but benchmark- and scaffold-indexed terms remain. Arbitrarily many similarly constructed tasks therefore cannot overcome model–benchmark heterogeneity or model–scaffold coupling. Corollary 3.2 adds the design punchline: at a fixed total task budget (N=n_b n_i), spreading budget across more benchmarks raises reliability whenever benchmark-conditioned differences are non-zero. Breadth reduces contamination by condition-specific advantages; it does not invent extra persistent model signal. The corollary assumes similarly informative benchmark draws, adequate crossing, and no extra benchmark setup cost. It does not claim that arbitrary new benchmarks always beat more tasks, or that benchmarks always outvalue scaffolds.[1]
On the data side, HAL’s nine benchmarks cover web navigation, scientific coding, multi-step assistance, software engineering, customer-service dialogue, and competitive programming (AssistantBench, CORE-Bench Hard, GAIA, Online-Mind2Web, SciCode, ScienceAgentBench, SWE-bench Verified Mini, τ-bench Airline, USACO). The paper reports about 29,923 HAL rollouts across 54 models and 13 scaffolds, with nine models appearing on every benchmark. Harbor Index supplies a complementary pooled analysis and external convergent checks. The observation graph is incomplete at every level—the usual leaderboard reality: evaluations are expensive, so many model–scaffold–benchmark combinations are simply absent. Shared models and scaffolds connect the graph enough for partial separation of effects. Sparseness is not a minor blemish; it is a design fact the method must face. Bayesian regularization and partial pooling stabilize estimates and propagate uncertainty, but they cannot manufacture empirical identification for disconnected components.[1]
Four results: how to read the numbers
1. Reliability follows the measurement goal
Across the nine HAL benchmarks, system ranking reliability (model–scaffold pairs) is Eρ²MA ∈ [0.935, 0.994], while model ranking reliability is Eρ²M ∈ [0.148, 0.841]. These are not conflicting readings of the same score. System reliability counts scaffold differences and compatibility as signal; model reliability requires differences that persist after averaging over scaffolds. A leaderboard can therefore be reliable for ranking deployable combinations and unreliable for ranking the underlying model alone. The introduction occasionally rounds system reliability to ([0.94, 0.99]); the abstract and results use 0.935–0.994.[1]
Recommendation, short form: specify whether the object is the model or a model–scaffold system, and whether interpretation is rank comparison or absolute scores; then estimate reliability for that claim rather than reporting one generic statistic. Shared multi-benchmark totals make this especially easy to blur.[1]
2. Changing the scaffold can change conclusions
Inter-scaffold reliability (with 95% HDIs) ranges from 0.151 [0.004, 0.577] on OnlineMind2Web to 0.852 [0.631, 0.947] on CORE-Bench Hard; posterior medians across benchmarks fall roughly in 0.151–0.852. At the low end, changing only the scaffold can substantially change which models look best, even with models and tasks fixed. The pooled decomposition further shows that the model-versus-scaffold main-effect contrast is not one-sided (directional posterior probability about 50%), while task-specific variation across scaffolds tends to exceed task-specific variation across models (posterior probability about 84% that (\sigma^2_{IA[B]}>\sigma^2_{IM[B]})). Scaffold choice can reshuffle which tasks get solved without making one scaffold uniformly stronger on the whole benchmark—a familiar engineering pattern when a tool layer or prompt orchestration changes and the failure set moves while the headline score barely does.[1]
If the goal is to evaluate the underlying model, test the same models across multiple scaffolds, report how much conclusions move, and avoid attributing scaffold-specific advantages to the model. If the deployed model–scaffold system is the object, scaffold variation can instead count as part of what is being measured—“ranks flipped after a harness change” is then not automatically an evaluation failure.[1]
3. More tasks cannot always resolve uncertainty
Adding tasks helps when the main uncertainty comes from differences across tasks. It does not fix uncertainty driven by other facets, such as limited scaffold coverage. Under the observed scaffold coverage, even infinitely many similarly constructed tasks would improve model-ranking reliability by at most about 0.10; the abstract states the same ceiling as 0.097—prefer 0.097, treating “~0.10” in the body as a rounded restatement. Only CORE-Bench Hard and SciCode can exceed Eρ² = 0.75 through task scaling alone. OnlineMind2Web is the ceiling example: reliability moves only from about 0.148 to 0.153.[1]
A benchmark can run out of useful information before it runs out of tasks. More of the same construction often makes the existing evaluation apparatus more precise without making a broader model claim much more reliable. Low reliability does not mean the benchmark measures an unimportant capability; it means that, for the models being compared, scores do not reliably distinguish the quantity of interest. Re-estimate reliability as the competitor population changes. Before expanding a benchmark, estimate how much reliability task additions can still buy—so you do not confuse busy workload with claim-level gains.[1]
4. Pooling diverse benchmarks: same budget, higher projected reliability
When the claim is ranking models across diverse agentic tasks, pooling benchmarks that test different kinds of work supplies more information about which differences persist across settings. At a fixed task budget, benchmark breadth averages condition-specific model advantages that within-benchmark replication leaves untouched. Spreading evaluations across the nine-benchmark battery raises projected model-ranking reliability from about 0.44 for a single benchmark to about 0.75. A complementary Harbor Index pooled analysis supports the same qualitative pattern. On cost: the full HAL battery exceeds about $47,000; a balanced allocation reaches comparable reliability for about $19,000; for an illustrative target of SNR ≈ 2.5 (Eρ² ≈ 0.71), about 14 tasks per benchmark cost about $8,144—an estimated 83% reduction. Appendix task-subsampling checks corroborate diminishing returns and whether reduced designs retain the ranking information D-studies predict.[1]
Reliability-adjusted latent model effects also agree more closely with four held-out agent benchmarks (BFCL v4, Terminal-Bench 2.0, SWE-bench Verified, τ²-bench Core) than a simple mean of benchmark scores—mean Kendall’s τ about 0.789 versus 0.463. That is convergent predictive evidence, not proof of a universal one-dimensional agent capability. Breadth helps most when the design is connected and benchmarks are informative; CORE-Bench Hard (especially with its own CORE Agent scaffold) carries the highest SNR in these data, and removing it raises pooled uncertainty more than other leave-one-out removals. Strong discriminating benchmarks can anchor a common scale when connectivity is sufficient. In imbalanced designs, finding informative bridges often beats hanging a few more sparsely evaluated models on the edge of the graph.[1]
When reading these four results, one confusion is common: high system reliability does not automatically mean the model claim is nearly ready. It often only says that, under the current task and scaffold sampling, fixed combinations rank stably. If the external sentence you need is “the model still leads after a reasonable harness swap,” look at model reliability, inter-scaffold reliability, and how much headroom the task-only ceiling still has. Conversely, if procurement or launch decisions target one fixed combination, a 0.93+ system (E\rho^2) may already be enough—and further spend on “pure” model ranking can answer the wrong question. Clear claims change how money should be spent.[1]
Another misread is treating “add 100 similar tasks” as a universal fix. The proposition plus the OnlineMind2Web example make a concrete point: when uncertainty is already dominated by scaffold (or benchmark) facets, more similar tasks mostly make the same conditioning louder; they do not restore missing facet coverage. On an engineering roadmap, benchmark expansion and orchestration coverage / benchmark breadth are different ledgers. Merging them produces the familiar report where task counts look healthy and claim reliability barely moves.[1]
A practitioner checklist for designing evals and reading HAL-style boards
Fold the paper’s recommendations into checklist form:
- Write the claim first. Is the object the underlying model, a fixed model–scaffold system, or model ranking across diverse agentic tasks? Do you need relative ranks or absolute scores / capability thresholds? Without a claim, a reliability number has no home. Prefer object language over “total score” language in system cards and methods sections.[1]
- Diagnose the limiting variance. Use (or demand) a variance decomposition: is uncertainty mainly task sampling, scaffold coverage, benchmark conditioning, or model×condition interactions? Decision studies (D-studies) project which additional samples buy the most reliability before collecting more data, and they should carry uncertainty in the variance estimates rather than plugging in a single point.[1]
- Do not default to more tasks. Add tasks when task sampling is the bottleneck. If the ceiling is already scaffold- or benchmark-limited (OnlineMind2Web’s 0.148→0.153), spend on broader scaffold coverage or benchmark breadth. Compute the task-only asymptotic reliability before expanding, so effort is not mistaken for claim-level gain.[1]
- On multi-scaffold boards, report splits. Separate system ranking from model ranking; when needed, report inter-scaffold reliability and sensitivity of conclusions to scaffold choice. If the deployment object is the system, count the scaffold inside the object. If you are comparing models, do not write a single-scaffold advantage as a model win.[1]
- For cross-task claims, prefer pooling over exhausting each benchmark. When connectivity and information content are adequate, spreading the same task budget across diverse benchmarks usually fits “differences that persist across settings” better than stacking tasks inside one benchmark. Use reliability analysis to decide how much evaluation each benchmark still needs while preserving enough task diversity. When a full battery is expensive, reliability-aware allocations (e.g., the illustrative ~14 tasks per benchmark) can approach a target SNR at far lower projected cost.[1]
- Treat absolute scores separately. Reliable rankings are not reliable scores. Decisions that need numeric thresholds need absolute reliability analysis; absolute reliability is no greater than the corresponding ranking reliability and is substantially lower on several benchmarks.[1]
Caveats: which claims stay fragile on sparse boards
The method faces the obvious: leaderboards are sparse, imbalanced, and only partially crossed. Bayesian regularization and partial pooling stabilize variance estimates and propagate uncertainty, but they cannot invent empirical identification for disconnected components—connectivity and estimability diagnostics matter. Per-benchmark scaffold counts are often too small to pin all scaffold-related components in isolation; those claims lean on shared scaffolds and pooling across benchmarks, so per-benchmark scaffold-variance precision remains weaker. Robustness checks localize the cost of sparseness; they do not prove sparseness is harmless.[1]
Generalization universes have edges. Generalizing to new tasks constructed like the benchmark does not establish that the benchmark represents all real deployments; reliable ranking inside the panel’s model population is not a deployment warranty for arbitrary new models. Improved Kendall’s τ on held-out boards is convergent predictive validity, not construct validity or proof of a universal capability. Posterior rank intervals can remain wide; neighboring models can have ordering probabilities near 1/2; some models move across rank quartiles between reported and adjusted rankings. Pooling improves design projections and external association—it does not magically erase ties. When the competitor set changes, re-estimate reliability: yesterday’s task budget may be insufficient once frontier models cluster more tightly.[1]
Relative to other evaluation threads on this site: CIR asks about the causal effect of recovery decisions from the same state; SAGE asks when self-evolving changes are statistically acceptable; Bad Genius / MoMHa ask how harnesses evolve and how multi-objective search should trade accuracy, safety, and tokens; this paper asks whether comparative inferences still stand after swapping tasks, scaffolds, or benchmarks. The lines do not substitute for each other. A system can be repeatable inside one harness while its rank depends heavily on that harness—repeatability and comparative reliability are different questions. If board reading only tracks “does variance shrink after a few more runs,” it will miss rank drift driven by scaffold and benchmark facets.[1]
Closing: more evaluation ≠ better evaluation
Leaderboard numbers look hard because they are written as a single truth. Hardy et al. split that truth back into claim, object, and design. Fixed system rankings can be highly reliable; underlying-model rankings and absolute scores are often much more fragile; stacking similarly constructed tasks has a clear ceiling (abstract upper bound about 0.097, with OnlineMind2Web barely moving); and for cross-task claims, spending budget on diverse benchmarks and the facets that actually limit uncertainty usually beats blindly adding tasks. When designing evaluations or reading HAL-style boards, the order should be: state what a score or ranking should mean → diagnose which variance sources limit that claim → spend evaluation budget on the uncertainties that matter. More evaluation is better only when it addresses those sources. Code and data: github.com/hardy-education/scaffold_eval.[1]
References
[1] Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo. Agent Evaluation Reliability: More Tasks Won’t (Always) Fix An Agent Leaderboard. arXiv:2610.00651, 2026. https://arxiv.org/abs/2610.00651 · https://arxiv.org/html/2610.00651
[2] Site crosslinks: CIR, SAGE, Bad Genius, MoMHa.