Full text not available for this paper
Summary (Overview)
- Novel Benchmark & Protocol: Introduces a training-free, hypothesis-constrained protocol for measuring "search scaling" in autonomous research, using 50 quantitative factor-mining tasks grounded in financial research reports from the Chinese A-share market.
- Key Finding on Capability vs. Search: Initial performance is more strongly associated with base model capability, while deeper search can narrow cross-model gaps—weaker models can partially catch up with more iterations.
- Model Grafting Insight: Early research-state quality materially shapes final performance; transferring intermediate research states between models reveals that the initial trajectory direction is a critical determinant of outcome.
- Parallel vs. Sequential: Parallel search outperforms sequential search under the same iteration budget, consistent with the benefit of broader coverage of the search space.
- Trajectory Analysis: Higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates, connecting qualitative behavior to quantitative outcomes.
Introduction and Theoretical Foundation
Background and Motivation
The paper extends the principle of inference scaling—where increased test-time computation improves LLM performance—to autonomous LLM agents through increased search budgets, termed search scaling. While prior work has characterized the mechanisms and limits of inference scaling in reasoning tasks, much less is known about these questions in the context of autonomous research, where agents must:
- Observe experimental results
- Diagnose failures
- Choose the next experiment
- Iteratively refine artifacts
The authors argue that search scaling is fundamentally different from simply generating more tokens or independent candidate answers: "added computation pays off only when the agent converts feedback into useful experiments, so the return to search may depend on the base model."
Theoretical Foundation
The paper distinguishes three key quantities:
- Base-model capability : inherent quality of the model
- Nominal search budget : externally controlled number of research iterations
- Realized resource consumption : tokens, tool calls, elapsed time, and API cost incurred at depth
The central question: How does base-model capability shape the marginal value of additional experimentation and the performance level that iterative search can reach?
Why Quantitative Factor Mining?
The authors choose factor construction as a testbed because:
- It requires end-to-end research loops (hypothesis → implementation → evaluation → refinement)
- It is training-free—additional iterations change research behavior without increasing model-training compute
- Economic hypotheses constrain the search space, discouraging revisions that abandon the intended hypothesis
- Noisy financial feedback and local optima preserve a substantive generalization problem, making held-out evaluation essential
Methodology
Problem Formulation
Let denote an autonomous factor-research objective, a fixed base model, and a shared general-purpose agent harness. The system follows a continuous trajectory on task . At search depth , it selects an incumbent factor using development-period evidence. An independent evaluator measures held-out quality:
The checkpoint is an artifact produced by research, whereas is observed only by the evaluator.
Controlled Iterative Search Protocol
- An external harness loop advances each model–task pair along a single research trajectory
- Each loop invocation permits one research cycle: act on current state, run experiments, inspect development-period feedback, update candidate
- The next cycle inherits workspace artifacts, memory, and interaction history
- The agent is not told the eventual trajectory length
- Checkpoints are frozen after every iteration; evaluation is separate from search and never feeds back into agent observations
Measuring Search Scaling
Aggregate held-out outcomes over the same tasks at each depth:
The paper proposes a shifted, saturating power law to describe the budget-performance curve:
Where:
- : plateau parameter
- : scale of remaining gains
- : decay rate with depth
- : shared offset fixed before fitting
Experimental Setup
- 50 tasks: grounded in quantitative research reports, spanning fundamental and technical ideas (momentum, reversal, valuation, earnings, growth, volatility, liquidity)
- Environment: Chinese A-share market
- Development period: Jan 1, 2013 – Jun 30, 2024
- Evaluation period: Jul 1, 2024 – Jul 17, 2026
- Models: 9 models evaluated along continuous 10-iteration trajectories
- Interventions: model grafting (transferring intermediate research states between models) and sequential–parallel comparison
Empirical Validation / Results
1. Capability vs. Search Depth
The paper finds that:
- Initial performance is more strongly associated with model capability
- Deeper search can narrow cross-model gaps—weaker models benefit more from additional iterations in relative terms
- The marginal gains from search are not simply proportional to initial quality; the relationship between capability and search returns is nuanced
2. Model Grafting
Model grafting—transferring the research state (checkpoint) from one model to another mid-trajectory—reveals that:
- The early research state materially shapes final performance
- A strong model starting from a weak model's state does not fully recover the performance it would achieve from its own trajectory
- Conversely, a weak model starting from a strong model's state inherits significant benefits
This suggests that path dependence matters: the direction set early in the research process constrains later outcomes.
3. Parallel vs. Sequential Search
Under the same total iteration budget:
- Parallel search (running multiple independent trajectories and selecting the best) outperforms sequential search (one continuous trajectory)
- This is consistent with the benefit of broader coverage of the search space—parallel exploration reduces the risk of getting stuck in local optima
4. Trajectory Analysis
Higher-performing models exhibit qualitatively different behaviors:
- More effective failure diagnosis: they correctly attribute feedback to its causes
- Better search redirection: they revise search directions when evidence warrants
- Hypothesis preservation: they retain candidates consistent with the assigned economic hypothesis, avoiding "reward hacking" where revisions abandon the intended hypothesis to chase backtest scores
Empirical Search Scaling Law
The fitted power law:
shows diminishing returns with search depth across all models, but the rate of saturation () and the plateau () differ by model capability. Higher-capability models reach higher plateaus and exhibit more favorable cost-performance trade-offs.
Theoretical and Practical Implications
Theoretical Implications
-
Search scaling is capability-dependent: The marginal value of additional experimentation is not uniform across models. This extends the test-time scaling literature from independent sampling to dependent empirical search.
-
Path dependence in research: The model grafting results demonstrate that research trajectories are path-dependent—early decisions constrain later outcomes. This has implications for how we think about the "state" of an autonomous research agent.
-
Coverage vs. depth trade-off: The parallel-vs-sequential result provides evidence that search breadth can be as important as depth, aligning with theoretical results on exploration–exploitation trade-offs.
Practical Implications
-
Resource allocation: Organizations deploying autonomous research agents should consider that weaker models may benefit more from additional search budget, while stronger models may warrant more parallel exploration.
-
Adaptive policies: The findings suggest that future progress requires "stronger models together with adaptive policies for deploying test-time computation throughout the research process."
-
Evaluation design: The protocol establishes a template for measuring search scaling that separates development-period search from held-out evaluation, avoiding feedback contamination.
-
Risk of overfitting: The hypothesis-constrained protocol highlights the risk of "adaptive selection that favors noise"—agents must be constrained to preserve intended economic hypotheses to avoid backtest overfitting.
Conclusion
This paper establishes a rigorous framework for studying search scaling in autonomous research, using quantitative factor mining as a controlled yet realistic testbed. Key takeaways:
-
Capability and search are complementary: Initial quality tracks model capability, but deeper search can partially close cross-model gaps.
-
Early state matters: Model grafting shows that the research trajectory's early direction significantly shapes final outcomes, implying that initial planning quality is crucial.
-
Breadth helps: Parallel search outperforms sequential search under the same budget, favoring broader exploration.
-
Behavioral differences drive outcomes: Higher-performing models exhibit better failure diagnosis, search redirection, and hypothesis preservation.
Future Directions
- Adaptive budget allocation: Developing policies that dynamically allocate search compute based on model capability and task difficulty
- Better search organizations: Exploring hybrid sequential–parallel strategies that combine the benefits of both
- Stronger base models: The ceiling on search scaling is partly set by base-model capability; improving base models remains essential
- Generalization to other domains: Extending the protocol to other autonomous research settings beyond finance
The paper concludes that "future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process." The search scaling framework provides a foundation for understanding how these two levers interact.
Related papers
- On-Demand Attention: Language Models Know When to Recall
On-demand attention uses a lightweight recall head to predict when global attention helps, recovering most quality with up to 2.65x decoding throughput.
- EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH shows that expanding an agent's tool, skill, or agent harness alone can degrade performance on previously solved tasks, inducing forgetting up to 34.7%.
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.