Full text not available for this paper

Summary of "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"

Summary (Overview)

  • Core finding: Under matched conditions (same time budget, hardware, and LLM backbone), a minimal-harness coding agent baseline ("Malena") matches or outperforms state-of-the-art open-source MLE harnesses on MLE-bench and NatureBench.
  • Key ablation result: The coding agent environment (filesystem + shell access) is the single largest performance driver; all additional machinery (search trees, multi-agent orchestration, memory hierarchies) yields no statistically significant gains.
  • Backbone-centric conclusion: Performance is driven primarily by the LLM backbone's agentic capability, not harness scaffolding—echoing the "bitter lesson" (Sutton, 2019).
  • Trace analysis evidence: Capable coding agents autonomously perform adaptive search, balancing rare technique exploration against performance-oriented refinement, without harness-imposed structure.
  • Practical implication: Effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.

Introduction and Theoretical Foundation

The paper addresses the growing complexity of autonomous Machine Learning Engineering (MLE) agents. Since Jiang et al. (2025) framed MLE as tree search over code, the dominant design has been the harness—an outer program that queries an LLM for code, executes it, and orchestrates exploration. Harnesses have grown increasingly elaborate, adding:

  • Search trees over candidate solutions
  • Populations of competing programs
  • Memory hierarchies and context management
  • Teams of specialized agents

The authors identify two premises underlying this machinery that deserve reconsideration:

  1. Single chat completions cannot interact with execution environments—but modern LLMs are post-trained for tool use and can operate inside coding agent environments (e.g., OpenCode), where they read/write files, execute code, and debug within a self-contained session.

  2. MLE is treated as a long-horizon problem requiring dedicated memory/context management—but with million-token context windows now affordable, this premise is questionable.

The central question: Which parts of the harness still earn their keep?

The paper also notes a methodological gap: harness performance has risen alongside backbone capability, making gains hard to attribute to harness design. Published results use different backbones/hardware, seed counts rarely exceed three, and headline margins can be smaller than run-to-run spread.


Methodology

Experimental Setup

  • Framework: OpenCode v1.15.6 coding agent framework
  • Primary benchmark: MLE-bench (75 tasks, with a fixed subset used for evaluation)
  • Secondary benchmark: NatureBench (40 open scientific-research tasks)
  • Hardware: 1× A100 80GB, 12 CPU cores, 144GB RAM (fixed across all experiments)
  • Backbones: GLM 5.2 (primary), Kimi K3, DeepSeek V4 Flash/Pro, Gemma 4 31B, GPT OSS 120B
  • Evaluation metrics: Medal rate (any-medal, %), mean percentile (vs. human leaderboard entries)
  • Selection protocols: Self-select (best validation score) vs. Oracle (best hidden test score)

Ablation Ladder (Section 4)

The authors re-implement intervention families in a single codebase, isolating four classes:

  1. Coding agent environment (Section 4.1): Comparing Chat (single prompt, text-only response) vs. Oneshot (single coding agent session with filesystem/execution access)

  2. Search primitives (Section 4.2): Four tree-search strategies at matched 24h budget:

    • Chain: always picks most recent node as parent (pure iterative improvement)
    • Greedy: picks best validation score node (pure exploitation)
    • UCB1: upper-confidence-bound rule over visits and scores (explore-exploit trade-off)
    • Best-of-N: independent, diverse exploration
  3. Autonomy (Section 4.3): Malena—a single long-running agent session with minimal tools:

    • Submission tool: registers submission files (no test scores communicated)
    • System tool: reports hardware resources and remaining time
    • Jobs tools: background bash execution, stdout inspection, job cancellation
  4. Multi-agent orchestration (Section 4.4): Three interventions on top of Malena:

    • +D (Delegation): background subagents via native task tool (hierarchical planner–executor)
    • +P3 (Parallelism): three independent Malena iterations on same machine
    • +B (Broadcast tool): message broadcasting between parallel agents

Production Harness Comparison (Section 5)

Malena compared against four open-source state-of-the-art harnesses:

  • MLEvolve (Du et al., 2026): older paradigm, minimal unit = single completion
  • AiScientist (Chen et al., 2026a): orchestrates multiple agentic sessions with tools
  • Arbor (Jin et al., 2026): tree-structured search over agentic sessions
  • ScienceFlow (Zhao et al., 2026): stagnation-triggered idea switching

Statistical Protocol

  • Macro-task averages with 95% bootstrap CIs
  • Paired-by-task comparisons for method differences
  • Fixed task sets (default29, fixed30, fixed14 splits)

Empirical Validation / Results

4.1 Coding Agent Environment Impact

  • Frontier LLMs (GLM 5.2, Kimi K3) dramatically improve with direct execution environment access
  • Even weaker models (Gemma 4) benefit—tool-use fluency matters more than raw model size
  • DeepSeek V4 models show larger improvement vs. their Preview versions (post-training for tool use)

4.2 Search Primitives

  • Search strategy has very small effect on performance (largely overlapping CIs)
  • Largest pair gap: Best-of-N vs. UCB1, with UCB1 ahead by 2.71 pp, 95% CI (−1.44, +6.60)—not statistically significant
  • Best-of-N's self-select falls 6.832pp short of oracle (largest selection gap among strategies)

4.3 Autonomy (Malena)

  • Malena achieves the highest average medal rate among all iterations
  • Significant medal rate advantage over all manual search algorithms (95% CI lower bounds: 0.4%–5.6%)
  • Lowest validation gap of all iterations: 1.876pp
  • Percentile differences vs. UCB1/Greedy are statistically indistinguishable

4.4 Multi-Agent Orchestration

ConditionPercentile (Self-select)Percentile (Oracle)Medal rate % (Self-select)Medal rate % (Oracle)
base66.51 [62.46, 69.97]69.29 [65.82, 72.13]55.7 [48.2, 62.9]60.4 [53.2, 66.8]
+D63.62 [59.27, 68.26]67.40 [63.22, 71.98]45.2 [38.1, 54.8]47.0 [38.1, 56.0]
+P367.77 [64.58, 70.62]71.79 [69.34, 73.83]52.4 [42.8, 61.9]60.7 [53.6, 67.3]
+B+P358.91 [53.50, 64.05]68.57 [65.77, 71.14]33.3 [23.8, 40.5]53.0 [45.8, 60.1]

Key finding: Base Malena performs no significantly worse than any intervention. Parallelism raises oracle ceiling but increases selection gap, negating self-select benefits.

5.1 Production Harness Comparison

MLE-bench (fixed30): Across 17 harness–backbone pairs, Malena matches or beats every harness on every tested pair at frontier backbones. With GLM 5.2: Malena earns medals on 62.5% of competitions vs. 47.1% for the best external harness.

NatureBench (Surpassed-SOTA rate, %)

BackboneMalenaAiScientistMLEvolve
GLM-5.221.7 [15.0, 27.5]17.5 [12.5, 22.5]10.8 [7.5, 15.0]
Kimi-K326.7 [20.0, 32.5]27.1 [20.0, 35.0]11.7 [7.5, 15.0]

Only exception: At smaller Gemma 4 31B backbone, MLEvolve leads by mean estimate (CIs still overlap).

Cost analysis: Malena's modeled USD cost (12.12)≈6.2×AiScientist′s(12.12) ≈ 6.2× AiScientist's (1.95) due to cache-read tokens from its single ever-growing session. Hardware cost dominates overall (24–24–85 for 24h A100 80GB).

5.2 Trace Analysis

  • Malena's workflow: Produces a single general pipeline spanning task stages, then performs hyperparameter search over code patches rather than parameter values
  • Technique rarity: Malena and MLEvolve are consistently top-2 in technique rarity (using techniques appearing in fewer other groups' runs)
  • Late-run behavior: Malena becomes more performance-oriented (ensembling, model selection, post-processing) while AiScientist shows a milder shift; MLEvolve stabilizes after ~10% of run
  • Notable example: Malena independently rediscovers pseudo-labeling under two different backbones

Theoretical and Practical Implications

  1. Bitter lesson confirmation: As agentic capability grows (via coding-agent post-training), hand-crafted workflow priors become redundant. Weaker backbones still benefit from harness structure, but frontier models do not.

  2. Primitive substitution: Coding-agent post-training substitutes for the scaffolding that harnesses hand-design. The "unit of work" has changed from single completions to autonomous sessions.

  3. Search is emergent: A capable coding agent performs adaptive search autonomously—balancing exploration of rare techniques against exploitation of performance gains—without harness-imposed structure.

  4. Cost trade-off: Minimal harnesses may have higher inference costs (cache-read tokens) but eliminate engineering overhead; hardware remains the dominant cost factor.

  5. Research prioritization: Effort should shift from harness engineering toward:

    • Improving backbone agentic capabilities
    • Runtime/tooling design (shell, filesystem access)
    • Reducing self-selection gaps (6.832pp for Best-of-N)

Conclusion

Main takeaways:

  • Almost all MLE performance gain comes from the runtime and LLM backbone; almost none from harness scaffolding
  • Giving the model a shell and filesystem (vs. chat interface) is the single largest measured effect
  • A single long-running agent session with minimal machinery is a strong baseline; no standard harness component improves on it
  • At frontier backbones, no published harness significantly outperforms the plain agent session

Future directions:

  • Deeper exploration of inter-agent coordination protocols (left as open question)
  • Mitigation strategies for self-selection gaps
  • Testing on benchmarks beyond MLE-bench/NatureBench (less saturated, less contaminated)

Limitations acknowledged:

  • MLE-bench has task-preparation issues and contamination risk (mitigated via task fixes, contamination checks, NatureBench evaluation)
  • Statistical power limits strong effect-size claims (though no-worse-than claims are well supported)
  • Scope limited to single-worker interventions
  • Possible asymmetry in tuning effort between Malena and external harnesses (documented in Appendix D.3)

Related papers