second brain
source
← 首页

external-source

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Note: body below is original English text extracted from arXiv abs / HTML. Do not treat this file as a translation.

arXiv:2609.20474 · published 2026-09-17 · submitted 17 Sep 2026

Abstract

Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61% of Retail oracle-invalid episodes while withholding 17% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

Authors

Yukun Zhang, Kemu Xu, Yishen Chen

Key claims (verbatim-leaning English extract)

Structure (section headings from HTML)

Remainder

Full original English text: see html_url / source_url / pdf_url in frontmatter.