Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

Summary (Overview)

  • Problem definition: This paper introduces intent-execution correspondence (IEC) — the property that the executed action matches the action the emitted tool call denotes under the tool contract. Tool calls pass through multiple hops (serialization, wrapper, shell parse, process interface, target parser), and any hop can silently alter the call.

  • Key finding: In 47,828 production shell calls, Claude Code's Bash tool changes 12.0% of calls carrying code, escape sequences, or long text. For 80.7% of calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change calls.

  • Two contributions: The IEC protocol observes what each hop received without executing the call and names the first hop that changed it (using the receiver's own parser). IntAct repairs the path by delivering the call in a form the named hop cannot alter, or refuses the call.

  • Evaluation: On the new IEC-Bench benchmark, the path raises token cost per passed task 2.4× (up to 12.3×). Trajectory-based judgment attributes 95.1% of production failures to the LLM, though the path caused more than half. IntAct recovers 79.2% of failures with a changed call, deployed in a commercial product.


Introduction and Theoretical Foundation

Background and Motivation

LLM-based agents build and run software through tool calls. A tool call reaches its program over an execution path: the harness serializes arguments, a host wrapper passes them to an interpreter, the OS forwards them, and the program parses them again. Each hop can change the call without notice, causing the program to execute an action the agent did not intend.

Problems of command execution account for 25.0% of 3,864 bug reports across 3 coding agents. These failures cost users time and money through failed runs and retries. Critically, benchmarks score the selected call or final state and assume the call executes as written, so path-caused failures are incorrectly attributed to the LLM.

Three Research Gaps

  1. Trajectory-based analyses do not observe the path — trajectories record the call and result but not what each hop received.
  2. Studies of the path examine only one hop — QuoteBench records only the wrapper; an agent's path has 5 hops, and a hop that changes a call hides changes after it.
  3. Real-world use is unmeasured — existing studies run benchmark tasks or analyze bug reports, never showing how often and at which hop calls change in production.

Formal Definition of IEC

The intended action aa of a tool call is the action the emitted call denotes under the tool contract. The call reaches the program over the execution path Π=⟨b1,…,bm⟩\Pi = \langle b_1, \ldots, b_m \rangle. At each hop, one side serializes and the receiving side parses. Let rir_i be the action the receiver of hop bib_i parses.

The path corresponds for aa when every hop receives the intended action:

ri=a for all ir_i = a \text{ for all } i

The first divergence is:

FD(a,Π)=min⁡{i∣ri≠a}\mathsf{FD}(a, \Pi) = \min\{i \mid r_i \neq a\}

A divergence is silent when the tool result contains no error flag or parser/quoting error message, and loud otherwise.

Four Change Mechanisms

  1. Quote consumption — text is parsed a second time (e.g., console.log("test") arrives as console.log(test))
  2. Re-expansion — a receiver reinterprets a delivered literal (%PATH% in batch, !x in interactive Bash)
  3. Truncation — Git Bash cuts arguments at 8,186 UTF-16 code units; cmd.exe ends at first line break
  4. Arity change — an operand is dropped or split

Methodology

The IEC Protocol

The protocol observes each hop without executing the call, using witnesses:

  • Bash: a file named by BASH_ENV reports the received text after passing through the wrapper
  • PowerShell: a parser computes the script powershell.exe would execute
  • Process interface: an argv probe records the argument vector a native program receives
  • Structured tools: witnesses record bound values the tool function receives

The protocol compares by the receiver's own parser (not string equality), comparing operators, words, heredoc bodies, parse trees, and argument vectors. This yields verdicts for 99.98% of 32,259 production Bash calls.

Failure Attribution

The IEC protocol attributes loud failures by:

  1. Parsing the agent's text with execution disabled — text that doesn't parse = generation error
  2. Delivering text through the witness — differences = path-caused
  3. Errors at lines beyond the agent's text = path-caused (wrapper-appended text)
  4. Text that parses, arrives intact, but invokes a further interpreter = nested hop

IntAct: Repair at the Named Hop

IntAct delivers the call through a channel the named hop does not parse. Key renderings include:

LaunchNamed hopRendering
Claude Code's Bashhost wrappersource a file
Qwen Code's cmd.exeshell parserefuse multi-line text
One-argument launchprocess interfacenative rendering of each literal
Appended launchwrapper + process interface-EncodedCommand + native rendering
Interactive Bashshell parsedisable history expansion/completion

The native rendering function for PowerShell 5.1:

native(s)={" "s=ε," esc(s) β(s)"s contains white space,esc(s)otherwise,\text{native}(s) = \begin{cases} \text{" "} & s = \varepsilon, \\ \text{" esc}(s) \, \beta(s) \text{"} & s \text{ contains white space}, \\ \text{esc}(s) & \text{otherwise}, \end{cases}

where esc writes \" for double quotes and doubles backslashes before them, and β(s)\beta(s) repeats the backslash run ending ss so the closing quote stays a delimiter.

IEC-Bench

Built from changes observed in real-world use:

  • 67 pattern pairs (hazard vs. control variants differing in one literal)
  • 135 ToolHop chains (3–7 steps)
  • 46 Terminal-Bench 2 tasks (paired across Linux/Windows containers)

Runs under 6 execution paths with 4 LLMs (qwen3-coder-plus, Qwen2.5-Coder-32B, Qwen2.5-72B, Hermes-3-Llama-3.1-70B).


Empirical Validation / Results

RQ-1: Real-World Change Rates

FeatureChanged-call rateCount
Backslash pair67.0%846/1,263
Heredoc9.3%546/5,901
1,000+ chars10.6%470/4,450
8,000+ chars100.0%64/64
None of these0.0%0/24,768

Key results:

  • 10.0% of 11,982 exposed calls change in production
  • Changed calls occur in 32.7% of sessions with exposed calls
  • Of 426 calls with parse errors, the IEC protocol attributes 306 to the path, 84 to generation, 36 to nested hops
  • All 10 harnesses and 10 of 22 MCP servers change calls

RQ-2: Impact on Scores and Costs

  • Qwen2.5-72B leads Hermes-3 by 29 chains with direct JSON calls, but falls 2–3 chains behind on PowerShell launches; IntAct restores the lead to 5–4 chains
  • Token cost per passed task: 2.4× overall, 12.3× for chains on the appended launch
  • In the commit task, only 1 of 39 runs commits the intended message; 19 runs end in false confirmation
  • Agents report 665 of 949 failed runs as completed (70.1% false-confirmation rate)

RQ-3: Attribution Accuracy

ApproachCorrectly attributed to path (of 56)Incorrectly (of 46)κ
Who&When all-at-once000.00
Who&When step-by-step29200.04
Who&When binary search25200.01
No witness (error message alone)56460.00
IEC protocol (ours)5410.77

The IEC protocol agrees with human labels on all 115 action-change pairs (κ = 1.00). QuoteBench's reference path detects 37.2% of path-caused losses — all at the wrapper, none at the process interface.

RQ-4: Repair Effectiveness

With recorded actions held fixed: IntAct recovers 137 of 173 lost pairs with a changed call (79.2%), and 0 of 187 without one (control, p < .001).

On the appended launch: process-interface repair alone recovers 0/69 pairs, wrapper alone 31, both together 69.

With agent in the loop (Table 7):

Launch, taskNoneIntActTemp. fileBoundaryFeedbackReflexion (10 trials)IntAct (10 trials)
One-arg, pairs (201)10412910385101148172
Appended, chains (405)974395836150
  • IntAct gains are significant by exact McNemar test (p < .001 for Qwen2.5 LLMs)
  • Evidence about the change does not help the agent (evidence ablation shows no significant difference)
  • Reflexion gains only 11.4–55.0% of IntAct's gains at 463–1,180 retries per cell

Theoretical and Practical Implications

Theoretical Contributions

  1. IEC as a formal property: The paper establishes intent-execution correspondence as a measurable, testable property of tool-call execution paths, distinct from generation correctness.

  2. Hop-hiding phenomenon: A hop that changes a call hides changes after it — 55.1% of failures on the appended launch appear only after the first hop is repaired. This invalidates single-hop analyses.

  3. Attribution correction: Trajectory-based judgment attributes 95.1% of production failures to the LLM, though the path caused more than half. Benchmarks comparing LLMs must hold the launch fixed.

Practical Contributions

  1. IntAct deployment: Released as a Claude Code mod, shell shim, MCP server, and deployed in a commercial desktop agent's execution kernel. Adds only 2.1–396 ms per call.

  2. Design guidance: Harnesses should be designed and tested hop-by-hop so correct calls execute as intended or are refused. A benchmark should report its launch and hold it fixed when comparing LLMs.

  3. Repair principles: The hop determines the repair — requoting every argument restores only 3 of 30 injected faults. Refusing a call is the only outcome that preserves correspondence when no rendering exists.


Conclusion

This paper presents the IEC protocol for measuring intent-execution correspondence and IntAct for repairing it. Key findings:

  • 10.0% of exposed production calls change; 80.7% of backslash changes are silent
  • The path, not the LLM, causes most tool-call failures; trajectory-based analysis misattributes them
  • IntAct recovers 79.2% of failures with changed calls and reduces token costs 2.4×

Future directions: Extend measurement to terminal-typing harnesses and new launches; grow IEC-Bench with each observed change; apply the protocol to structured tools and MCP servers. The central insight: a changed call comes from the launch, not the LLM — harnesses must establish intent-execution correspondence independently of the agent.

Related papers