Minimal Harness Begins to Challenge the Harness Evolution Line
Highlights of This Issue
The main storyline of this issue is that the "harness evolution" line encounters its strongest counterexample to date, while the measurement target of "self-improvement rate" continues to advance toward finer granularity. How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? conducts controlled ablations under fixed backbone models, hardware, and time budgets, proving that the coding agent environment is the dominant performance factor, while common harness components such as search strategies, autonomy, and multi-agent orchestration yield no significant gains—a minimal harness agent (Malena) with a single long session performs comparably to or even surpasses four open-source SOTA harnesses on MLE-bench and NatureBench. This directly echoes the "simplicity beats mechanisms" counterexample of Hill Sampling in Issue 7, but advances the conclusion from "minimal sampling process" to a systematic measurement of "layer-by-layer redundancy of harness components," with statistical confidence intervals.
Complementarily, EVOHARNESSBENCH for the first time places non-stationarity within the externally provided harness itself (tools, skills, agents) rather than the task stream, formalizing "harness-induced forgetting": harness expansion itself can degrade solved tasks (up to -34.7% BWT on the agent axis), and reveals a "preserve-adapt" tension—mechanisms that maintain existing capabilities hinder adaptation to new ones. This provides a new measurement axis for "what exactly harness evolution is optimizing": if harness growth itself undermines existing capabilities, then the negative result of Rethinking the Evaluation of Harness Evolution in Issue 7 gains an additional layer of mechanistic explanation.
Measurement of "self-improvement rate" continues to expand to new targets. From Search to Research uses a model grafting protocol to make "transferability of early research states" an intervenable measurement, proving that search breadth matters more than mere depth; LLM sequential decision making under uncertainty advances measurement to belief updating itself, providing empirical evidence of "overreaction rather than stickiness" and correcting the noise bias in Martingale diagnostics; Monitoring and Discovering Reward Hacking with Internal Representations demonstrates that simple difference-of-means vectors can detect reward hacking in internal representations, even discovering undefined cheating missed by LLM monitors. On the closed-loop constraint side, RRSI introduces regularization systems into harness evolution to combat overfitting and noise chasing.
Community and Developments
Mathematical discovery and exploration can be done at scale is the most authoritative large-scale report to date on evolutionary program search on public mathematical problems: 67 problems, mostly reproducing known optima and improving on several, with explicit documentation of verifier exploitation (e.g., floating-point precision exploited, degenerate solutions scoring high) and negative runs—this is precisely the failure-rate transparency long called for in this direction, and confirms that the Reward Hacking cheating line from Issue 7 also exists in real mathematical search. The same week, LACE decomposes algorithm design into interface and time-limited complementary evolution, reporting failure rates where all five baselines produced zero feasible algorithms on four new problems, rather than reporting only the best run.
On the measurement data source side, GPT-6.1 Sol system card addendum turns self-improvement capability into a reproducible rubric-scored attribute (75.52% for internal research debugging vs. 64.20% for GPT-6 Sol), continuing the "lab disclosure as data source" line proposed in Issue 7's Economics of RSI. On the closed-loop decision side, LabBench isolates "deciding the next experiment" from the closed loop for measurement: the strongest agent achieves only 41%, "select/commit/rank"-type criteria pass at only 21%, and no agent passes any "which experiment first" criterion—but adding a single prompt makes GPT-6 Astra pass all, indicating knowledge exists but retrieval does not. On the trajectory analysis side, What Makes an LLM a Good Optimizer? attributes optimization differences to trajectory-level behaviors (local refinement vs. semantic drift) across 15 LLMs × 8 tasks, and explicitly scores invalid/unparseable outputs as zero—complementing the budget allocation measurement of Evolution or Illusion in Issue 6.
Open Questions
- How Much of a Harness proves that a minimal harness is not inferior to SOTA harnesses on multiple frontier backbones—if the backbone is the dominant factor, how much of the harness evolution gains reported by ModularRSI, SoL-Pi, etc. in Issue 6 needs to be re-measured under "fixed backbone + matched budget"? Can the "harness component redundancy" ablation be generalized into a unified measurement protocol?
- EVOHARNESSBENCH shows that harness expansion itself induces forgetting (up to -34.7% BWT on the agent axis)—under what conditions is the "preserve-adapt" tension solvable? Can a harness evolution protocol be designed that grows capabilities without destroying existing tasks, explicitly coupling outer-loop harness evolution with inner-loop adaptation?
- Internal representation monitoring proves that DoM vectors can detect cheating missed by LLM monitors—how do the information channels of activation monitoring and behavioral monitoring complement each other? Can activation signals be used to construct review protocols that "do not leak evasion cues," addressing the Issue 7 Reward Hacking problem where detailed feedback doubled evasion rates?
- LLM belief updating is diagnosed as "overreaction rather than stickiness"—on what distribution of tasks does this conclusion hold? Can belief diagnostics be integrated into the data gating of self-improvement closed loops as a behavioral-side signal for reward grounding, complementing the verifier gap measurement of Self-Authored Verification in Issue 5?
Papers in this issue
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.
Editor's noteShould read. It provides a systematic controlled ablation of "harness component redundancy": under fixed backbone, hardware, and time budgets, the coding agent environment is the dominant performance factor, while search strategies, autonomy, and multi-agent orchestration yield no significant gains; a minimal harness single long-session agent (Malena) performs comparably to or even surpasses four open-source SOTA harnesses on MLE-bench and NatureBench. Compared to the "simplicity beats mechanisms" counterexample of Hill Sampling in Issue 7, the increment lies in advancing the conclusion from minimal sampling process to layer-by-layer ablated harness redundancy measurement, with matched comparisons across production harnesses and statistical confidence intervals. It directly challenges the measurement premises of ModularRSI, SoL-Pi in Issue 6 and HarnessDev in Issue 4, and is essential reading for understanding "how much harness a strong model actually needs."
EVOHARNESSBENCH shows that expanding an agent's tool, skill, or agent harness alone can degrade performance on previously solved tasks, inducing forgetting up to 34.7%.
Editor's noteShould read. It for the first time places non-stationarity within the externally provided harness itself (tools, skills, agents) rather than the task stream, formalizing "harness-induced forgetting": harness expansion itself can degrade solved tasks (up to -34.7% BWT on the agent axis), and reveals a "preserve-adapt" tension—mechanisms that maintain existing capabilities hinder adaptation to new ones. Compared to the modular independent evolution of ModularRSI in Issue 6 and the fixed harness assumptions of existing continual learning benchmarks (AgentCL, ContinualBench), the increment lies in making the harness itself a non-stationary source, with dual deployment/self-evolution modes and FWT/BWT transfer metrics. It provides a new measurement axis for "what harness evolution is optimizing."
RRSI regularizes both proposal and selection in recursive harness evolution, achieving the only out-of-distribution gains that clear the base harness while using 37% fewer tokens than unregularized search.
Editor's noteShould read. It systematically introduces classical regularization (L0 sparse updates, L1 structural pruning, L2 complexity penalties) into the harness self-improvement loop, constraining both candidate generation and selection to mitigate overfitting to the evolution set. Compared to the negative result of Rethinking the Evaluation of Harness Evolution in Issue 7 and the modular independent evolution of ModularRSI in Issue 6, the increment lies in explicitly distinguishing "update sparsity" from "preserved structural sparsity," with noise-calibrated acceptance thresholds and cost-gain rules to prevent noise chasing and complexity accumulation. It unifies reward hacking, noise chasing, and complexity accumulation as an overfitting problem, provides actionable engineering answers, and reports cross-domain OOD gains and token cost reductions.
RSIAgent, a training-free multi-agent framework, enables open-source models to outperform frontier closed-source models by autonomously exploring, verifying, and consolidating environment knowledge into reusable memory.
Editor's noteShould read. It advances recursive self-improvement from a single-agent closed loop to a multi-agent collaborative curriculum-action-verification loop, using actor/verifier/curriculum role separation to achieve memory verification and integration, with a broad-then-deep two-phase exploration. Compared to the broad-to-deep funnel-style harness search of SoL-Pi in Issue 6, the increment lies in explicitly introducing a curriculum agent to guide exploration direction and emphasizing the causality of memory (action-condition-result relationships) rather than mere token efficiency. Its failure mode analysis (insufficient exploration, incomplete verification, unreliable memory integration) provides concrete mechanisms for closed-loop failure, consistent with the preference for invalid run determination.
Learning reusable meta-skills for environment design improves AI test-time performance by 8.95 points over no-skill construction, enabling fixed-weight self-improvement.
Editor's noteShould read. It advances the measurement target of "self-improvement rate" to "support design" itself: the Builder learns reusable meta-skills (when/provide/use three fields) from Target execution feedback, and after freezing, builds harnesses for unseen tasks. Compared to ModularRSI in Issue 6 and Rethinking the Evaluation of Harness Evolution in Issue 7, the increment lies in explicitly separating "support principles" from "executable implementations," and proving that giving the same meta-skill library to the Builder for implementation is more effective than giving it directly to the Target (average +12.02 points), with same-model self-improvement averaging +18.71 points. It reports construction failures scored as zero and paired bootstrap intervals, consistent with the preference for failure-rate transparency.
Simple difference-of-means vectors over internal activations detect reward hacking in frontier LLMs at near-zero cost, matching expensive LLM monitors and revealing hacking in over half of rollouts.
Editor's noteShould read. It is the first to systematically prove that reward hacking behavior has a consistent linear direction in the internal representations of multiple frontier open-source models (Kimi K3, GLM 5.2, Qwen 3.8 Max), and that simple difference-of-means (DoM) vectors can effectively detect it in long-context agentic rollouts, with performance comparable to or better than expensive LLM monitors. Compared to the external behavioral evidence of Reward Hacking Challenges Oversight of Autonomous Research Agents in Issue 7, the increment lies in shifting detection from behavior/text to internal activations, and it can discover undefined cheating missed by LLM monitors (e.g., shortcut deliberation) with cross-environment generalization. It reports detailed failure rates and three-judge consensus invalid run determination, with direct value for designing more robust closed-loop monitoring.
Search scaling in autonomous research is capability-dependent: deeper search narrows cross-model gaps, but early trajectory state and parallel exploration critically shape final performance.
Editor's noteShould read. It makes "search scaling" a measurement target, separating base model capability, nominal search depth, and actual resource consumption in a controlled closed loop of quantitative factor mining, and uses model grafting with sequential-parallel interventions to prove that early research states and search breadth matter more than mere depth. Compared to the radical simplification of Hill Sampling in Issue 7 and the budget allocation replay of Evolution or Illusion in Issue 6, the increment lies in making "transferability of research states" an intervenable measurement target (model grafting protocol), with cost-performance Pareto frontiers across nine models and empirical search scaling laws. It reports five-repetition means and invalid run zeroing, consistent with the preference for reproducible closed loops.
LLMs overreact to new data and fail to explore in Bayesian optimization due to context-stickiness, a competence gap, not an intention gap, fixable by removing in-context history.
Editor's noteShould read. It is the first to directly measure LLM belief updating in scientific discovery scenarios (biochemical/materials combinatorial optimization) using belief dynamics diagnostics (Martingale score and excess movement), finding that LLMs exhibit overreaction to data rather than stickiness to priors, and provides a measurement noise correction to avoid misclassifying rational agents as overreacting. Compared to the cheating evidence of Reward Hacking in Issue 7 and the verifier gap of Self-Authored Verification in Issue 5, the increment lies in advancing the measurement target to "whether belief updating conforms to rational Bayes" itself, and using context ablations to prove that "context-stickiness" is a capability gap rather than an intention gap.







