Back to articles
The Harness Produces the Capability You Think You Are Buying

The Harness Produces the Capability You Think You Are Buying

A scaffolded small model scored 0.912 where the bigger model scored 0.619. Agent capability is engineered around the model, not purchased with it.

August 24, 2026 · 9 min read
Add as a preferred source on Google

Every model release sets off the same budget conversation. The agent underperforms, a stronger model ships, someone proposes the upgrade, and the roadmap absorbs the cost as though capability were a line item you procure.

Salesforce AI Research ran the experiment that should complicate that meeting. They took GPT-5.4-mini, which scores a macro-average accuracy of 0.488 across four Theory-of-Mind benchmarks, and left its weights alone. Instead, stronger builder models wrote inference-time scaffolds around it: routing logic, prompt templates, deterministic solvers, format enforcement. The best scaffolded run reached 0.912. The unscaffolded GPT-5.4, the bigger model a team in that meeting would have bought, scores 0.619.

Same weak model. No retraining. The gap was opened by engineering the conditions the model works in.

Key takeaways

  • Scaffolding a weak model beat upgrading to a stronger one in Salesforce AI Research’s test-time transfer study: GPT-5.4-mini rose from 0.488 to 0.912 with no weight updates, while unscaffolded GPT-5.4 scored 0.619.
  • Harness surface area belongs to the product team, not the model provider. Instructions, tools, sandboxes, orchestration, hooks and observability are all owned locally, which makes agent reliability an engineering budget rather than a procurement decision.
  • Reading three 2026 harness studies together, I’d argue the knobs that scale spend are the wrong ones. In the Salesforce data, scaffold code size correlated with accuracy at r≈0.22 while deterministic offloading correlated at r=0.72.
  • Extra scaffolding regresses a model that is already close to ceiling. Against the stronger Gemini-3.5-flash target, every builder made at least one benchmark worse, in 9 of 20 matched cases.

Bar chart: GPT-5.4-mini scores 0.488 alone, the larger GPT-5.4 scores 0.619, and GPT-5.4-mini with a harness scores 0.912.

An agent is a model plus everything your team owns

The clearest statement of the anatomy comes from Osmani, Saboo and Kartakis in their May 2026 paper on the new SDLC. Their equation is blunt: agent equals model plus harness. A raw model is not an agent at all; it becomes one once something gives it state, tool execution, feedback loops and enforceable constraints.

Read the inventory as a budget line. Instructions and rule files. Tools and the prose that tells the model when to call them. Sandboxes and execution environments. Orchestration logic for sub-agent spawning and model routing. Deterministic hooks that fire before a tool call or after a file edit. Observability, without which there is no way to tell whether an agent is performing or quietly drifting.

Then the line that reframes the whole cost question: “If that sounds like a lot of surface area, it is. And it is the team’s surface area, not the model provider’s.”

I’d argue that sentence is the actual finding. A model upgrade is a supplier decision. A harness is a product you build, version and own, which is why its reusable pieces, from specs that survive the model changing underneath them to skills treated as durable assets rather than prompt snippets, compound in a way that a subscription tier never does.

Decision flowchart: sort 30 agent failures into compilable or judgment-bound, then choose harness work or a model upgrade.

The scaffolded small model beat the bigger model

The study’s design closes the obvious escape hatch. Each builder model saw only a 195-item validation sample, 5% of the data, and never the 3,900-item hidden test set. Memorised answers would not have transferred.

Across 57 scaffolded runs the mean macro-average was 0.763, an uplift of +0.275, and every single run beat the baseline. The best individual run, GPT-5.5 building on GPT Codex, hit 0.912, an absolute gain of +0.423, or 87% relative. A paired McNemar test over the 3,900 items shows the shape of that gain: the scaffold fixed 1,717 baseline errors while breaking 105 previously correct ones.

Hou and colleagues at the National University of Defense Technology reach the same conclusion from a different direction. Their Harness-G work redesigns the retrieval interface for search agents, replacing free-form query generation with selection from a bounded action menu while the environment builds the query and tracks state. Holding the graph, the reward, the training budget and the GRPO configuration fixed, that interface change alone improves F1 by more than 17 points, and by more than 35 points on MuSiQue under outcome-only training. Free-form querying produced illusory exploration, where query wording stayed diverse while retrieval outcomes collapsed. The share of rollout groups reaching genuinely different evidence fell from about 86% to below 10% by training step 30.

Miyai, Aizawa and Yamasaki put a ceiling on the effect, citing evidence that changing the harness around a fixed model yields up to a 6× performance difference on the same benchmark.

The authors put the comparison in one chart, with the unscaffolded models drawn in as reference bars, so you can check the claim rather than take my word for it.

Table of eleven builder models beside a bar chart where scaffolded GPT-5.4-mini beats both unscaffolded models on all four benchmarks.

Exhibit 1. The eleven builder models and their macro-average scores, beside the per-benchmark comparison of the best scaffold against two unscaffolded models and a human-designed harness. In the table, each row is one builder model that wrote a scaffold for GPT-5.4-mini: R is how many runs were pooled, Avg. is the macro average across the four benchmarks with sd its standard deviation, and the last column is the uplift over the 0.488 no-scaffold baseline in the top row. In the chart, from Salesforce AI Research, where AI is artificial intelligence, four bars per benchmark are named in the legend: GPT-5.4 with no scaffold, GPT-OSS-120B with no scaffold, the best scaffold running on the much weaker GPT-5.4-mini, and UserHarness, the harness a human designed for these tasks; MMToM-QA is the question-answering, or QA, benchmark of the four. The scaffolded small model clears both unscaffolded bars on all four benchmarks, reading 1.00, 0.80, 0.84 and 0.88 against 0.49, 0.68, 0.64 and 0.70 for unscaffolded GPT-5.4, while the human-designed harness still wins three of the four at 0.95, 0.87, 0.98 and 0.96. Source: Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307, Figure 3, p. 7.

The expensive knob is the wrong knob

Here is the pattern none of the three papers states on its own. In every study, the knob that scales spending failed to move the outcome, and the knob that scales structure did.

Salesforce shows it twice. The number of validation evaluations a builder ran was essentially uncorrelated with final accuracy (Pearson r=0.17), while the builder’s own reasoning effort improved scaffold quality monotonically. And scaffold code size related only weakly to accuracy (r≈0.22), while the determinism fraction, the share of items answered entirely by code or rules, correlated at r=0.72. Moving the right work into deterministic structure was what moved the outcome.

Harness-G shows it in the reward signal. Methods that densify credit for free-form querying leave a large residual gap that the interface redesign closes, which means the fix was in how actions are exposed rather than in how they are scored.

Task-CoEvolve shows it in the evaluation loop. Concentrating evaluation on tasks where candidate harnesses actually disagree matched full-set search while cutting evaluations by 80%, and at a 20% budget it beat full-set search outright on text classification. The tasks every candidate already passes, and the ones every candidate still fails, were consuming most of the budget and producing almost no signal.

Bar chart of determinism fraction by benchmark: BigToM 0.94, Hi-ToM 0.51, MMToM-QA 0.44, MuMA-ToM 0.36.

The compilable-failure test

So before signing the upgrade, run what I’d call the compilable-failure test.

Pull your last thirty agent failures from logs. Sort each one into compilable or judgment-bound: compilable means a deterministic rule, a schema check, a format constraint or a routing decision would have caught it, and judgment-bound means the model genuinely had to reason its way through. Then read the split as a budget. The compilable share is harness headroom, available now at engineering cost. The judgment-bound share is the only part a model upgrade actually addresses.

My decision rule: if more than half your sampled failures are compilable, a model upgrade is the more expensive way to buy accuracy you could engineer. Spend on the harness first.

The Salesforce data supports the split rather than stating it. Compilability varied enormously by task. BigToM was almost fully offloadable at a mean determinism fraction around 0.94, Hi-ToM sat near 0.51, MMToM-QA near 0.44, and MuMA-ToM, built on free-form dialogue reasoning, near 0.36. The residual errors concentrated exactly where compilation failed.

The prevalence table in the Salesforce study puts format enforcement in 100% of scaffolds, temperature control in 98% and benchmark routing in 95%. Those are a reliability floor rather than a differentiator; they make higher performance possible without explaining it. The techniques that separated strong scaffolds from weak ones were the ones requiring real task analysis: polarity and negation logic (+0.09), structured state extraction (+0.06), hybrid fallback (+0.04).

Sequence diagram: the harness sandboxes the model, runs tests, routes failures back, and blocks commits via hooks.

Where more harness makes it worse

The honest limit on all of this is that scaffolding is a competence-recovery mechanism, not a capability generator. Realized uplift tracked the target’s available headroom at r=0.75, what the authors call a headroom law. Harness-G’s own scale split reproduces that pattern in a different architecture, raising average F1 by 10.74 points at 1.5B against 3.98 points at 3B.

Which means it can backfire. Against the weak GPT-5.4-mini target, no builder regressed below baseline on any benchmark, 0 of 20 matched cases. Against the already-strong Gemini-3.5-flash, every builder regressed somewhere, 9 of 20 cases, with Hi-ToM down 0.04 and near-saturated MuMA-ToM down 0.02 on average. Extra routing and rules disrupted correct behaviour more often than they repaired errors.

So gate harness work by remaining headroom on the specific path, and freeze the harness on paths already near ceiling. Build variance is real: the mean standard deviation across repeats was 0.036 against a +0.275 mean uplift, but the widest setting spanned 0.201, with the widest spreads coming from deterministic-solver strategies where one logic error shifts accuracy by tens of points. The authors’ own recipe is to build two or three scaffolds and ship the best validation performer, which is cheap insurance for a team already running an agentic SDLC.

The practical move this quarter is smaller than a migration. Sample thirty failures, sort them, and see what your own split says before anyone reprices the model line.

References

  • Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307.
  • Hou, Y., Chen, H., Zhou, S., Chen, X., Liu, X., Yuan, D., Meng, L., Liu, Q., & Huang, J. (2026). Harness-G: A Graph-Structured Harness for Search Agents. National University of Defense Technology, arXiv:2607.27652.
  • Miyai, A., Aizawa, K., & Yamasaki, T. (2026). Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection. The University of Tokyo, arXiv:2608.20169.
  • Osmani, A., Saboo, S., & Kartakis, S. (2026). The New SDLC With Vibe Coding: From ad-hoc prompting to Agentic Engineering. May 2026.

Frequently asked questions

What is an agent harness?

The harness is everything wrapped around the model that lets an agent finish work: instructions and rule files, tools, sandboxes, orchestration and model routing, deterministic hooks, and observability. Osmani, Saboo and Kartakis state it as an equation, agent equals model plus harness, and note that a raw model only becomes an agent once a harness gives it state, tool execution, feedback loops and enforceable constraints.

Does improving the harness really beat upgrading the model?

It did in Salesforce AI Research's 2026 test-time transfer study. Scaffolding lifted GPT-5.4-mini from 0.488 to 0.912 macro-average accuracy with no weight updates, while the unscaffolded and stronger GPT-5.4 scored 0.619. Task-CoEvolve cites evidence that harness changes alone can produce up to a 6× performance difference on the same benchmark.

When is a model upgrade the right call instead of harness work?

When your failures need judgment rather than structure. Sort recent agent failures into ones a deterministic rule, schema check or routing decision would have caught, and ones where the model genuinely had to reason. Only the second group is what a stronger model buys, and in the Salesforce data the residual errors concentrated exactly where reasoning resisted compilation into rules.

Can adding more harness make an agent worse?

Yes, on tasks the model already handles well. Against the stronger Gemini-3.5-flash target, every builder in the Salesforce study regressed on at least one benchmark, 9 of 20 matched cases, versus 0 of 20 against the weaker target. Realized uplift tracked available headroom at r=0.75, so scaffolding paths already near ceiling tends to disrupt correct behaviour.

Evidence

Then the line that reframes the whole cost question: "If that sounds like a lot of surface area, it is. And it is the team's surface area, not the model provider's."

If that sounds like a lot of surface area, it is. And it is the team's surface area, not the model provider's.

Osmani, A., Saboo, S., & Kartakis, S. (2026). The New SDLC With Vibe Coding: From ad-hoc prompting to Agentic Engineering. May 2026. (p. 28)

The number of validation evaluations a builder ran was essentially uncorrelated with final accuracy (Pearson r=0.17), while the builder's own reasoning effort improved scaffold quality monotonically.

the amount of refinement itself is not predictive: the number of validation iterations is essentially uncorrelated with final full-set accuracy (Pearson 𝑟 = 0.17, right panel). In short, builder quality matters much more than how often the builder probes the validation set.

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 8)

Concentrating evaluation on tasks where candidate harnesses actually disagree matched full-set search while cutting evaluations by 80%, and at a 20% budget it beat full-set search outright on text classification.

Task-CoEvolve consistently outperforms fixed-subset baselines and matches the final performance of full-set search while reducing the number of evaluations during optimization by 80%.

Miyai, A., Aizawa, K., & Yamasaki, T. (2026). Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection. The University of Tokyo, arXiv:2608.20169. (p. 1)

Instructions and rule files. Tools and the prose that tells the model when to call them. Sandboxes and execution environments. Orchestration logic for sub-agent spawning and model routing. Deterministic hooks that fire before a tool call or after a file edit.

Instructions and Rule Files: The text that defines who the agent is... Tools: The functions, MCP servers, and APIs the agent can call, plus the prose around them that tells the model when and how to call them. Sandboxes and execution environments... Orchestration logic: Sub-agent spawning, model routing, hand-offs between specialists... Guardrails or Hooks: Deterministic code that runs at specific lifecycle points: before a tool call, after a file edit, before a commit.

Osmani, A., Saboo, S., & Kartakis, S. (2026). The New SDLC With Vibe Coding: From ad-hoc prompting to Agentic Engineering. May 2026. (p. 28)

Their equation is blunt: agent equals model plus harness.

Figure 7: Harness Anatomy | Agent = Model + Harness

Osmani, A., Saboo, S., & Kartakis, S. (2026). The New SDLC With Vibe Coding: From ad-hoc prompting to Agentic Engineering. May 2026. (p. 27)

The tasks every candidate already passes, and the ones every candidate still fails, were consuming most of the budget and producing almost no signal.

Tasks that every candidate can already solve, or that no candidate can yet solve, continue to consume the evaluation budget while providing little signal for optimization.

Miyai, A., Aizawa, K., & Yamasaki, T. (2026). Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection. The University of Tokyo, arXiv:2608.20169. (p. 2)

The techniques that separated strong scaffolds from weak ones were the ones requiring real task analysis: polarity and negation logic (+0.09), structured state extraction (+0.06), hybrid fallback (+0.04).

Among the better-powered contrasts, the largest positive associations are polarity/negation logic (+0.09), structured extraction (+0.06), hybrid fallback (+0.04), and deterministic solving. These are not generic prompting tricks; they directly target the main failure modes of the ToM benchmarks

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 13)

Against the weak GPT-5.4-mini target, no builder regressed below baseline on any benchmark, 0 of 20 matched cases.

For GPT-5.4-mini, no builder regresses below baseline on any benchmark (0/20 matched cases).

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 11)

The share of rollout groups reaching genuinely different evidence fell from about 86% to below 10% by training step 30.

the fraction of rollout groups spanning multiple retrieval-

Hou, Y., Chen, H., Zhou, S., Chen, X., Liu, X., Yuan, D., Meng, L., Liu, Q., & Huang, J. (2026). Harness-G: A Graph-Structured Harness for Search Agents. National University of Defense Technology, arXiv:2607.27652. (p. 1)

Harness-G's own scale split reproduces that pattern in a different architecture, raising average F1 by 10.74 points at 1.5B against 3.98 points at 3B.

Harness-G raises average F1 from 40.09 to 50.83, a 10.74- ... point improvement over Graph-R1, and outperforms it on ... cate that restricting retrieval to executable, structured actions is particularly helpful at limited model capacity

Hou, Y., Chen, H., Zhou, S., Chen, X., Liu, X., Yuan, D., Meng, L., Liu, Q., & Huang, J. (2026). Harness-G: A Graph-Structured Harness for Search Agents. National University of Defense Technology, arXiv:2607.27652. (p. 6)

Harness-G shows it in the reward signal.

IGPO (Wang et al. 2025a) densifies free-query credit but leaves a large residual gap to the menu; Menu+SNC is best on all three multi-hop datasets.

Hou, Y., Chen, H., Zhou, S., Chen, X., Liu, X., Yuan, D., Meng, L., Liu, Q., & Huang, J. (2026). Harness-G: A Graph-Structured Harness for Search Agents. National University of Defense Technology, arXiv:2607.27652. (p. 6)

Each builder model saw only a 195-item validation sample, 5% of the data, and never the 3,900-item hidden test set.

Each builder additionally receives a 195-item (5%) validation sample drawn by a fixed random seed. The primary metric is the unweighted macro average of the four per-benchmark full set accuracies

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 6)

Against the already-strong Gemini-3.5-flash, every builder regressed somewhere, 9 of 20 cases, with Hi-ToM down 0.04 and near-saturated MuMA-ToM down 0.02 on average.

For Gemini-3.5-flash, by contrast, every builder regresses on at least one benchmark (9/20 cases), especially on tasks where the baseline is already high: Hi-ToM (−0.04 on average) and near-saturated MuMA-ToM (−0.02 on average). This illustrates the risk of over-scaffolding

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 11)

BigToM was almost fully offloadable at a mean determinism fraction around 0.94, Hi-ToM sat near 0.51, MMToM-QA near 0.44, and MuMA-ToM, built on free-form dialogue reasoning, near 0.36.

is almost fully offloadable, with mean determinism around ... partially offloadable (≈ 0.51), typically through symbolic

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 14)

The prevalence table in the Salesforce study puts format enforcement in 100% of scaffolds, temperature control in 98% and benchmark routing in 95%.

two techniques are nearly universal: robust format enforcement, which ... format enforcement, greedy decoding, routing, and forced

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 9)

The best scaffolded run reached 0.912.

The best individual run, produced by GPT-5.5 on GPT Codex, reaches 0.912, an uplift of +0.423 (87% relative).

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 7)

A paired McNemar test over the 3,900 items shows the shape of that gain: the scaffold fixed 1,717 baseline errors while breaking 105 previously correct ones.

the gain over the GPT-5.4-mini no-scaffold baseline is overwhelmingly significant under a paired McNemar test over the 3900 evaluation items, with 𝜒2 ≫ 104 and 𝑝 < 10−4 . The direction of the item-level changes is also highly asymmetric: the scaffold fixes 1717 baseline errors while breaking only 105 previously correct items.

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 13)

A raw model is not an agent at all; it becomes one once something gives it state, tool execution, feedback loops and enforceable constraints.

A raw model is not an agent. It becomes one once a harness gives it state, tool execution, feedback loops, and enforceable constraints.

Osmani, A., Saboo, S., & Kartakis, S. (2026). The New SDLC With Vibe Coding: From ad-hoc prompting to Agentic Engineering. May 2026. (p. 27)

And scaffold code size related only weakly to accuracy (r≈0.22), while the determinism fraction, the share of items answered entirely by code or rules, correlated at r=0.72.

is only weakly related to accuracy ... what matters is not simply writing more code, but writing code

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 14)

The authors' own recipe is to build two or three scaffolds and ship the best validation performer, which is cheap insurance for a team already running an agentic SDLC.

cal recipe: because failures are usually visible on vali- ... scaffolds and selecting the best validation performer

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 8)

Holding the graph, the reward, the training budget and the GRPO configuration fixed, that interface change alone improves F1 by more than 17 points, and by more than 35 points on MuSiQue under outcome-only training.

The action menu improves F1 by more than 17 points over free-query under both credit regimes ... and by more than 35 points on MuSiQue under outcome-only training

Hou, Y., Chen, H., Zhou, S., Chen, X., Liu, X., Yuan, D., Meng, L., Liu, Q., & Huang, J. (2026). Harness-G: A Graph-Structured Harness for Search Agents. National University of Defense Technology, arXiv:2607.27652. (p. 6)

Miyai, Aizawa and Yamasaki put a ceiling on the effect, citing evidence that changing the harness around a fixed model yields up to a 6× performance difference on the same benchmark.

Indeed, simply changing the harness around a fixed LLM has been shown to yield up to a 6× difference in performance on the same benchmark (Tian et al., 2026), suggesting that the harness can matter as much as the underlying model itself.

Miyai, A., Aizawa, K., & Yamasaki, T. (2026). Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection. The University of Tokyo, arXiv:2608.20169. (p. 2)

Realized uplift tracked the target's available headroom at r=0.75, what the authors call a headroom law.

Uplift follows a headroom law. As illustrated in Figure 8(b), across all builder×benchmark×target settings, realized uplift is strongly predicted by the target's available headroom on that benchmark, 1 − baseline (Pearson 𝑟 = 0.75).

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 11)

The unscaffolded GPT-5.4, the bigger model a team in that meeting would have bought, scores 0.619.

many scaffolded GPT-5.4-mini configurations surpass the no- ... scaffold can sometimes yield gains larger than upgrading to a

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 6)

Across 57 scaffolded runs the mean macro-average was 0.763, an uplift of +0.275, and every single run beat the baseline.

Across all 57 scaffolded GPT-5.4-mini runs, the mean macro-average accuracy is 0.763, corresponding to an uplift of +0.275 over the baseline, and 100% of runs exceed the baseline.

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 7)

Build variance is real: the mean standard deviation across repeats was 0.036 against a +0.275 mean uplift, but the widest setting spanned 0.201, with the widest spreads coming from deterministic-solver strategies where one logic error shifts accuracy by tens of points.

The mean standard deviation of all the macro-average is 0.036, roughly an order of magnitude smaller than the +0.275 mean uplift. At the same time, the widest setting has a repeat range of 0.201, indicating that the build process is still not fully deterministic.

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 7)

The honest limit on all of this is that scaffolding is a competence-recovery mechanism, not a capability generator.

Scaffolding acts primarily as a competence-recovery mechanism: its payoff is governed by how much latent ability the target fails to deploy... This suggests that scaffolding primarily recovers latent competence that the target model already possesses but does not reliably deploy

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 11)

They took GPT-5.4-mini, which scores a macro-average accuracy of 0.488 across four Theory-of-Mind benchmarks, and left its weights alone.

Vanilla: Each target model is called directly with the same naive prompt, without any task-specific scaffolding. This setting yields a macro-average accuracy of 0.488 for GPT-5.4-mini and 0.761 for Gemini-3.5-flash.

Qian, C., Zhao, W., Yang, L., Wang, H., Qiu, J., Ji, H., Savarese, S., Wang, H., & Heinecke, S. (2026). AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses. Salesforce AI Research preprint, arXiv:2608.12307. (p. 6)