Back to News

BenchShield Framework Detects Reward Hacking in LLM-Agent Evaluations

#reward hacking#llm agents#benchmark evaluation#formal verification

Researchers from Dartmouth, UC Berkeley, and BenchFlow AI introduced BenchShield, a framework for detecting reward hacking in LLM-agent evaluation infrastructure. It extends checks to the entire scoring process, using static analysis to find exploit paths and runtime evidence to determine if agents attempted or used them. The paper highlights a case where a Lean proof verifier gave full marks to an incorrect proof due to disabled kernel type checking.

Coverage timeline

  1. 机器之心机器之心

    BenchJack 之后,reward hacking 检测还缺什么? Agent 在 benchmark 上拿到高分,未必真正完成了任务。它也可能利用评测或奖励机制中的漏洞,绕开题目要求来提高得分。这类行为通常被称为 reward hacking。 此前,BenchJack 对 10 个 Agent benchmark 开展系统审计,发现了 219 处漏洞,并构造出能在多数 benchmark 上获得接近满分、却不解决实际任务的利用方式。研究团队将常见问题归纳为八类漏洞模式,整理成 Agent-Eval Checklist,供 benchmark 设计者检查和修补评测系统。 发现漏洞之后,还需要判断具体的一次运行:Agent 有没有使用这个漏洞?它的操作是否影响了最终分数?有漏洞的任务中也可能出现合规的执行过程,而一个通过测试的答案,也可能来自任务不允许的获取方式。 论文:《BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure》 论文链接:https://arxiv.org/abs/2609.11028v1 来自达特茅斯学院、加州大学伯克利分校、BenchFlow AI 等机构的研究者提出了 BenchShield。论文题为《BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure》。 BenchShield 将检查扩展到产生分数的完整过程:先明确各阶段应满足的要求,再用静态分析查找潜在利用路径,并根据运行时证据判断 Agent 是否尝试或实际使用了这些路径。 被隔离的 verifier,为什么仍然会给出满分? 绕过类型检查,让错误的证明通过验证 论文中的一个 Lean 定理证明任务,已经将负责检查提交结果的 verifier 与 Agent 的运行环境隔离。Agent 提交了一份错误的证明,却仍获得了满分(reward = 1.00)。 原因在于,提交的 Lean 源代码启用了 debug.skipKernelTC,关闭了 Lean 内核的类型