# Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

> Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.

- **Source:** [arXiv](https://arxiv.org/abs/2610.04375)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/ppRVBb
- **Whiteboard:** https://picx.dev/p/ppRVBb/image

## Summary

# Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

## Summary (Overview)

- **Problem definition**: This paper introduces **intent-execution correspondence (IEC)** — the property that the executed action matches the action the emitted tool call denotes under the tool contract. Tool calls pass through multiple hops (serialization, wrapper, shell parse, process interface, target parser), and any hop can silently alter the call.

- **Key finding**: In 47,828 production shell calls, Claude Code's Bash tool changes **12.0%** of calls carrying code, escape sequences, or long text. For **80.7%** of calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change calls.

- **Two contributions**: The **IEC protocol** observes what each hop received without executing the call and names the first hop that changed it (using the receiver's own parser). **IntAct** repairs the path by delivering the call in a form the named hop cannot alter, or refuses the call.

- **Evaluation**: On the new IEC-Bench benchmark, the path raises token cost per passed task **2.4×** (up to 12.3×). Trajectory-based judgment attributes **95.1%** of production failures to the LLM, though the path caused more than half. IntAct recovers **79.2%** of failures with a changed call, deployed in a commercial product.

---

## Introduction and Theoretical Foundation

### Background and Motivation

LLM-based agents build and run software through tool calls. A tool call reaches its program over an **execution path**: the harness serializes arguments, a host wrapper passes them to an interpreter, the OS forwards them, and the program parses them again. Each hop can change the call without notice, causing the program to execute an action the agent did not intend.

Problems of command execution account for **25.0%** of 3,864 bug reports across 3 coding agents. These failures cost users time and money through failed runs and retries. Critically, **benchmarks score the selected call or final state and assume the call executes as written**, so path-caused failures are incorrectly attributed to the LLM.

### Three Research Gaps

1. **Trajectory-based analyses do not observe the path** — trajectories record the call and result but not what each hop received.
2. **Studies of the path examine only one hop** — QuoteBench records only the wrapper; an agent's path has 5 hops, and a hop that changes a call hides changes after it.
3. **Real-world use is unmeasured** — existing studies run benchmark tasks or analyze bug reports, never showing how often and at which hop calls change in production.

### Formal Definition of IEC

The intended action $a$ of a tool call is the action the emitted call denotes under the tool contract. The call reaches the program over the execution path $\Pi = \langle b_1, \ldots, b_m \rangle$. At each hop, one side serializes and the receiving side parses. Let $r_i$ be the action the receiver of hop $b_i$ parses.

The path **corresponds** for $a$ when every hop receives the intended action:
$$r_i = a \text{ for all } i$$

The **first divergence** is:
$$\mathsf{FD}(a, \Pi) = \min\{i \mid r_i \neq a\}$$

A divergence is **silent** when the tool result contains no error flag or parser/quoting error message, and **loud** otherwise.

### Four Change Mechanisms

1. **Quote consumption** — text is parsed a second time (e.g., `console.log("test")` arrives as `console.log(test)`)
2. **Re-expansion** — a receiver reinterprets a delivered literal (`%PATH%` in batch, `!x` in interactive Bash)
3. **Truncation** — Git Bash cuts arguments at 8,186 UTF-16 code units; cmd.exe ends at first line break
4. **Arity change** — an operand is dropped or split

---

## Methodology

### The IEC Protocol

The protocol observes each hop without executing the call, using **witnesses**:

- **Bash**: a file named by `BASH_ENV` reports the received text after passing through the wrapper
- **PowerShell**: a parser computes the script powershell.exe would execute
- **Process interface**: an argv probe records the argument vector a native program receives
- **Structured tools**: witnesses record bound values the tool function receives

The protocol compares by the **receiver's own parser** (not string equality), comparing operators, words, heredoc bodies, parse trees, and argument vectors. This yields verdicts for **99.98%** of 32,259 production Bash calls.

### Failure Attribution

The IEC protocol attributes loud failures by:
1. Parsing the agent's text with execution disabled — text that doesn't parse = **generation error**
2. Delivering text through the witness — differences = **path-caused**
3. Errors at lines beyond the agent's text = **path-caused** (wrapper-appended text)
4. Text that parses, arrives intact, but invokes a further interpreter = **nested hop**

### IntAct: Repair at the Named Hop

IntAct delivers the call through a channel the named hop does not parse. Key renderings include:

| Launch | Named hop | Rendering |
|---|---|---|
| Claude Code's Bash | host wrapper | source a file |
| Qwen Code's cmd.exe | shell parse | refuse multi-line text |
| One-argument launch | process interface | native rendering of each literal |
| Appended launch | wrapper + process interface | `-EncodedCommand` + native rendering |
| Interactive Bash | shell parse | disable history expansion/completion |

The native rendering function for PowerShell 5.1:

$$\text{native}(s) = \begin{cases} \text{" "} & s = \varepsilon, \\ \text{" esc}(s) \, \beta(s) \text{"} & s \text{ contains white space}, \\ \text{esc}(s) & \text{otherwise}, \end{cases}$$

where `esc` writes `\"` for double quotes and doubles backslashes before them, and $\beta(s)$ repeats the backslash run ending $s$ so the closing quote stays a delimiter.

### IEC-Bench

Built from changes observed in real-world use:
- **67 pattern pairs** (hazard vs. control variants differing in one literal)
- **135 ToolHop chains** (3–7 steps)
- **46 Terminal-Bench 2 tasks** (paired across Linux/Windows containers)

Runs under 6 execution paths with 4 LLMs (qwen3-coder-plus, Qwen2.5-Coder-32B, Qwen2.5-72B, Hermes-3-Llama-3.1-70B).

---

## Empirical Validation / Results

### RQ-1: Real-World Change Rates

| Feature | Changed-call rate | Count |
|---|---|---|
| Backslash pair | 67.0% | 846/1,263 |
| Heredoc | 9.3% | 546/5,901 |
| 1,000+ chars | 10.6% | 470/4,450 |
| 8,000+ chars | 100.0% | 64/64 |
| None of these | 0.0% | 0/24,768 |

**Key results**:
- **10.0%** of 11,982 exposed calls change in production
- Changed calls occur in **32.7%** of sessions with exposed calls
- Of 426 calls with parse errors, the IEC protocol attributes **306 to the path**, 84 to generation, 36 to nested hops
- All 10 harnesses and 10 of 22 MCP servers change calls

### RQ-2: Impact on Scores and Costs

- Qwen2.5-72B leads Hermes-3 by 29 chains with direct JSON calls, but falls **2–3 chains behind** on PowerShell launches; IntAct restores the lead to 5–4 chains
- Token cost per passed task: **2.4×** overall, **12.3×** for chains on the appended launch
- In the commit task, only **1 of 39 runs** commits the intended message; **19 runs end in false confirmation**
- Agents report **665 of 949** failed runs as completed (70.1% false-confirmation rate)

### RQ-3: Attribution Accuracy

| Approach | Correctly attributed to path (of 56) | Incorrectly (of 46) | κ |
|---|---|---|---|
| Who&When all-at-once | 0 | 0 | 0.00 |
| Who&When step-by-step | 29 | 20 | 0.04 |
| Who&When binary search | 25 | 20 | 0.01 |
| No witness (error message alone) | 56 | 46 | 0.00 |
| **IEC protocol (ours)** | **54** | **1** | **0.77** |

The IEC protocol agrees with human labels on all 115 action-change pairs (κ = 1.00). QuoteBench's reference path detects **37.2%** of path-caused losses — all at the wrapper, none at the process interface.

### RQ-4: Repair Effectiveness

**With recorded actions held fixed**: IntAct recovers **137 of 173** lost pairs with a changed call (79.2%), and **0 of 187** without one (control, p < .001).

**On the appended launch**: process-interface repair alone recovers 0/69 pairs, wrapper alone 31, both together 69.

**With agent in the loop** (Table 7):

| Launch, task | None | IntAct | Temp. file | Boundary | Feedback | Reflexion (10 trials) | IntAct (10 trials) |
|---|---|---|---|---|---|---|---|
| One-arg, pairs (201) | 104 | 129 | 103 | 85 | 101 | 148 | 172 |
| Appended, chains (405) | 9 | 74 | 39 | 5 | 8 | 36 | 150 |

- IntAct gains are significant by exact McNemar test (p < .001 for Qwen2.5 LLMs)
- Evidence about the change does **not** help the agent (evidence ablation shows no significant difference)
- Reflexion gains only 11.4–55.0% of IntAct's gains at 463–1,180 retries per cell

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **IEC as a formal property**: The paper establishes intent-execution correspondence as a measurable, testable property of tool-call execution paths, distinct from generation correctness.

2. **Hop-hiding phenomenon**: A hop that changes a call hides changes after it — **55.1%** of failures on the appended launch appear only after the first hop is repaired. This invalidates single-hop analyses.

3. **Attribution correction**: Trajectory-based judgment attributes **95.1%** of production failures to the LLM, though the path caused more than half. Benchmarks comparing LLMs must hold the launch fixed.

### Practical Contributions

1. **IntAct deployment**: Released as a Claude Code mod, shell shim, MCP server, and deployed in a commercial desktop agent's execution kernel. Adds only **2.1–396 ms** per call.

2. **Design guidance**: Harnesses should be designed and tested **hop-by-hop** so correct calls execute as intended or are refused. A benchmark should report its launch and hold it fixed when comparing LLMs.

3. **Repair principles**: The hop determines the repair — requoting every argument restores only 3 of 30 injected faults. Refusing a call is the only outcome that preserves correspondence when no rendering exists.

---

## Conclusion

This paper presents the IEC protocol for measuring intent-execution correspondence and IntAct for repairing it. Key findings:

- **10.0%** of exposed production calls change; **80.7%** of backslash changes are silent
- The path, not the LLM, causes most tool-call failures; trajectory-based analysis misattributes them
- IntAct recovers **79.2%** of failures with changed calls and reduces token costs **2.4×**

**Future directions**: Extend measurement to terminal-typing harnesses and new launches; grow IEC-Bench with each observed change; apply the protocol to structured tools and MCP servers. The central insight: *a changed call comes from the launch, not the LLM* — harnesses must establish intent-execution correspondence independently of the agent.

---

_Markdown view of https://picx.dev/p/ppRVBb, served by PicX — AI-generated visual whiteboard summaries of research papers._
