Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Summary (Overview)
-
Problem definition: This paper introduces intent-execution correspondence (IEC) — the property that the executed action matches the action the emitted tool call denotes under the tool contract. Tool calls pass through multiple hops (serialization, wrapper, shell parse, process interface, target parser), and any hop can silently alter the call.
-
Key finding: In 47,828 production shell calls, Claude Code's Bash tool changes 12.0% of calls carrying code, escape sequences, or long text. For 80.7% of calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change calls.
-
Two contributions: The IEC protocol observes what each hop received without executing the call and names the first hop that changed it (using the receiver's own parser). IntAct repairs the path by delivering the call in a form the named hop cannot alter, or refuses the call.
-
Evaluation: On the new IEC-Bench benchmark, the path raises token cost per passed task 2.4× (up to 12.3×). Trajectory-based judgment attributes 95.1% of production failures to the LLM, though the path caused more than half. IntAct recovers 79.2% of failures with a changed call, deployed in a commercial product.
Introduction and Theoretical Foundation
Background and Motivation
LLM-based agents build and run software through tool calls. A tool call reaches its program over an execution path: the harness serializes arguments, a host wrapper passes them to an interpreter, the OS forwards them, and the program parses them again. Each hop can change the call without notice, causing the program to execute an action the agent did not intend.
Problems of command execution account for 25.0% of 3,864 bug reports across 3 coding agents. These failures cost users time and money through failed runs and retries. Critically, benchmarks score the selected call or final state and assume the call executes as written, so path-caused failures are incorrectly attributed to the LLM.
Three Research Gaps
- Trajectory-based analyses do not observe the path — trajectories record the call and result but not what each hop received.
- Studies of the path examine only one hop — QuoteBench records only the wrapper; an agent's path has 5 hops, and a hop that changes a call hides changes after it.
- Real-world use is unmeasured — existing studies run benchmark tasks or analyze bug reports, never showing how often and at which hop calls change in production.
Formal Definition of IEC
The intended action of a tool call is the action the emitted call denotes under the tool contract. The call reaches the program over the execution path . At each hop, one side serializes and the receiving side parses. Let be the action the receiver of hop parses.
The path corresponds for when every hop receives the intended action:
The first divergence is:
A divergence is silent when the tool result contains no error flag or parser/quoting error message, and loud otherwise.
Four Change Mechanisms
- Quote consumption — text is parsed a second time (e.g.,
console.log("test")arrives asconsole.log(test)) - Re-expansion — a receiver reinterprets a delivered literal (
%PATH%in batch,!xin interactive Bash) - Truncation — Git Bash cuts arguments at 8,186 UTF-16 code units; cmd.exe ends at first line break
- Arity change — an operand is dropped or split
Methodology
The IEC Protocol
The protocol observes each hop without executing the call, using witnesses:
- Bash: a file named by
BASH_ENVreports the received text after passing through the wrapper - PowerShell: a parser computes the script powershell.exe would execute
- Process interface: an argv probe records the argument vector a native program receives
- Structured tools: witnesses record bound values the tool function receives
The protocol compares by the receiver's own parser (not string equality), comparing operators, words, heredoc bodies, parse trees, and argument vectors. This yields verdicts for 99.98% of 32,259 production Bash calls.
Failure Attribution
The IEC protocol attributes loud failures by:
- Parsing the agent's text with execution disabled — text that doesn't parse = generation error
- Delivering text through the witness — differences = path-caused
- Errors at lines beyond the agent's text = path-caused (wrapper-appended text)
- Text that parses, arrives intact, but invokes a further interpreter = nested hop
IntAct: Repair at the Named Hop
IntAct delivers the call through a channel the named hop does not parse. Key renderings include:
| Launch | Named hop | Rendering |
|---|---|---|
| Claude Code's Bash | host wrapper | source a file |
| Qwen Code's cmd.exe | shell parse | refuse multi-line text |
| One-argument launch | process interface | native rendering of each literal |
| Appended launch | wrapper + process interface | -EncodedCommand + native rendering |
| Interactive Bash | shell parse | disable history expansion/completion |
The native rendering function for PowerShell 5.1:
where esc writes \" for double quotes and doubles backslashes before them, and repeats the backslash run ending so the closing quote stays a delimiter.
IEC-Bench
Built from changes observed in real-world use:
- 67 pattern pairs (hazard vs. control variants differing in one literal)
- 135 ToolHop chains (3–7 steps)
- 46 Terminal-Bench 2 tasks (paired across Linux/Windows containers)
Runs under 6 execution paths with 4 LLMs (qwen3-coder-plus, Qwen2.5-Coder-32B, Qwen2.5-72B, Hermes-3-Llama-3.1-70B).
Empirical Validation / Results
RQ-1: Real-World Change Rates
| Feature | Changed-call rate | Count |
|---|---|---|
| Backslash pair | 67.0% | 846/1,263 |
| Heredoc | 9.3% | 546/5,901 |
| 1,000+ chars | 10.6% | 470/4,450 |
| 8,000+ chars | 100.0% | 64/64 |
| None of these | 0.0% | 0/24,768 |
Key results:
- 10.0% of 11,982 exposed calls change in production
- Changed calls occur in 32.7% of sessions with exposed calls
- Of 426 calls with parse errors, the IEC protocol attributes 306 to the path, 84 to generation, 36 to nested hops
- All 10 harnesses and 10 of 22 MCP servers change calls
RQ-2: Impact on Scores and Costs
- Qwen2.5-72B leads Hermes-3 by 29 chains with direct JSON calls, but falls 2–3 chains behind on PowerShell launches; IntAct restores the lead to 5–4 chains
- Token cost per passed task: 2.4× overall, 12.3× for chains on the appended launch
- In the commit task, only 1 of 39 runs commits the intended message; 19 runs end in false confirmation
- Agents report 665 of 949 failed runs as completed (70.1% false-confirmation rate)
RQ-3: Attribution Accuracy
| Approach | Correctly attributed to path (of 56) | Incorrectly (of 46) | κ |
|---|---|---|---|
| Who&When all-at-once | 0 | 0 | 0.00 |
| Who&When step-by-step | 29 | 20 | 0.04 |
| Who&When binary search | 25 | 20 | 0.01 |
| No witness (error message alone) | 56 | 46 | 0.00 |
| IEC protocol (ours) | 54 | 1 | 0.77 |
The IEC protocol agrees with human labels on all 115 action-change pairs (κ = 1.00). QuoteBench's reference path detects 37.2% of path-caused losses — all at the wrapper, none at the process interface.
RQ-4: Repair Effectiveness
With recorded actions held fixed: IntAct recovers 137 of 173 lost pairs with a changed call (79.2%), and 0 of 187 without one (control, p < .001).
On the appended launch: process-interface repair alone recovers 0/69 pairs, wrapper alone 31, both together 69.
With agent in the loop (Table 7):
| Launch, task | None | IntAct | Temp. file | Boundary | Feedback | Reflexion (10 trials) | IntAct (10 trials) |
|---|---|---|---|---|---|---|---|
| One-arg, pairs (201) | 104 | 129 | 103 | 85 | 101 | 148 | 172 |
| Appended, chains (405) | 9 | 74 | 39 | 5 | 8 | 36 | 150 |
- IntAct gains are significant by exact McNemar test (p < .001 for Qwen2.5 LLMs)
- Evidence about the change does not help the agent (evidence ablation shows no significant difference)
- Reflexion gains only 11.4–55.0% of IntAct's gains at 463–1,180 retries per cell
Theoretical and Practical Implications
Theoretical Contributions
-
IEC as a formal property: The paper establishes intent-execution correspondence as a measurable, testable property of tool-call execution paths, distinct from generation correctness.
-
Hop-hiding phenomenon: A hop that changes a call hides changes after it — 55.1% of failures on the appended launch appear only after the first hop is repaired. This invalidates single-hop analyses.
-
Attribution correction: Trajectory-based judgment attributes 95.1% of production failures to the LLM, though the path caused more than half. Benchmarks comparing LLMs must hold the launch fixed.
Practical Contributions
-
IntAct deployment: Released as a Claude Code mod, shell shim, MCP server, and deployed in a commercial desktop agent's execution kernel. Adds only 2.1–396 ms per call.
-
Design guidance: Harnesses should be designed and tested hop-by-hop so correct calls execute as intended or are refused. A benchmark should report its launch and hold it fixed when comparing LLMs.
-
Repair principles: The hop determines the repair — requoting every argument restores only 3 of 30 injected faults. Refusing a call is the only outcome that preserves correspondence when no rendering exists.
Conclusion
This paper presents the IEC protocol for measuring intent-execution correspondence and IntAct for repairing it. Key findings:
- 10.0% of exposed production calls change; 80.7% of backslash changes are silent
- The path, not the LLM, causes most tool-call failures; trajectory-based analysis misattributes them
- IntAct recovers 79.2% of failures with changed calls and reduces token costs 2.4×
Future directions: Extend measurement to terminal-typing harnesses and new launches; grow IEC-Bench with each observed change; apply the protocol to structured tools and MCP servers. The central insight: a changed call comes from the launch, not the LLM — harnesses must establish intent-execution correspondence independently of the agent.
Related papers
- The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Self-improving LLM loops on small reused evaluation sets inflate reported gains by 13–20 points, with most proposals after the first rewrite being harmful.
- Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.
- Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.