Full text not available for this paper
LLM Sequential Decision Making Under Uncertainty in Biochemical Domains
Summary (Overview)
- This paper benchmarks five frontier LLMs against published statistical baselines in Bayesian Optimization (BO) settings across seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly, and catalysis.
- Through a prompt ablation that progressively strips chemical context (default → alias → blind), the authors separate memorization from chemical reasoning from bare categorical optimization, finding that prior chemical knowledge helps on average but with high variance and occasionally harms performance.
- Using belief-movement and Martingale diagnostics (corrected for measurement-noise bias), the authors show LLMs overreact to incoming data rather than entrenching on their priors in scientific discovery contexts.
- A key finding is that LLM actions are exploitative despite models sincerely intending to explore—this failure is identified as a competence gap arising from "context-stickiness" (clustering selections around candidates already in the context window), not a preference or intention gap.
- Removing in-context history restores exploration behavior, demonstrating that priors and data must be decoupled for effective LLM-driven discovery.
Introduction and Theoretical Foundation
Background and Motivation
Scientific research involves searching enormous experimental spaces under cost and time constraints. Discovery unfolds as sequential decisions under partial knowledge where new information often contradicts prior belief. Large language models (LLMs) are increasingly deployed in this loop to propose, test, and revise hypotheses in response to data.
The authors formalize this discovery loop as Bayesian Optimization (BO), where a prior over the search space is updated as data accumulate, and each new query is drawn from the resulting posterior. BO has become standard for biochemical and materials discovery due to sample efficiency in costly evaluation settings where acquisition functions make the exploration-exploitation tradeoff explicit.
Conflicting Literature
The literature reports mixed signals on LLM performance as BO optimizers:
- Positive results: Reinhart et al. (2024) showed Claude 3.5 outperforms active-learning and genetic-algorithm pipelines on macromolecule design.
- Negative results: Gupta et al. (2025) showed statistical models outperform open-source LLMs, with LLMs appearing insensitive to experimental feedback (no worse with randomly permuted labels).
- Intermediate results: The authors' earlier work found LLMs overfixate on irrelevant context, where withholding information actually improved performance.
Theoretical Framework: Belief Diagnostics
The paper applies two diagnostic frameworks from cognition and economics:
-
Belief-movement (Augenblick–Rabin): Formalizes how far a rational agent's beliefs should move as new evidence arrives. The excess-movement statistic compares realized belief movement to realized uncertainty reduction.
-
Martingale score (He et al.): The slope in the regression of belief updates on prior beliefs. For a rational Bayesian agent, ; positive movement and negative slope indicate over-reaction, while the reverse indicates entrenchment.
Key theoretical insight: Two opposing pathologies exist in different settings:
- In semantically rich domains (forecasting, paper review): LLMs entrench (beliefs drift toward prior)
- In abstract, semantically inert domains (card-drawing tasks): LLMs overreact to weak signals
Scientific discovery is both at once—semantically rich with only a few noisy points in a vast search space—making it unclear which regime governs LLM behavior.
Methodology
Optimization Domains and Datasets
Seven datasets spanning five biochemical domains:
| Domain | Dataset | Search Space Size | Dimensions | Data Points | Coverage |
|---|---|---|---|---|---|
| Protein | TrpB | 159,129 | 4 | 149,361 | 99.5% |
| Protein | GB1 | 149,361 | 4 | 159,129 | 93.4% |
| Reaction | direct_arylation | 1,728 | 5 | 1,728 | 100% |
| Reaction | aryl_amination_2b | 792 | 4 | 792 | 100% |
| Peptide | tripeptides | 8,000 | 3 | ~8,000 | — |
| NFA | NFA | 301 | 3 | 301 | 78.4% |
| Catalysis | OER | 2,121 | 10-binned 6-simplex | 2,121 | — |
Prompt Ablation Design
Three prompting modes progressively strip prior knowledge without changing the optimization function:
- Default: Full biochemical names, descriptors (charge, size, pKa), and reaction context
- Alias: Reaction class preserved but component names replaced with generic labels (ligand_1, base_1); descriptors retained
- Blind: Domain-agnostic one-hot search space (e.g., "car, animal, tree" labels); no chemical context
LLM Configurations
Three configurations tested:
- No reasoning: Reasoning turned off, one-shot candidate proposals
- Reasoning: Standard reasoning model with thinking enabled
- Agent: Equipped with a tool wrapping the statistical model, providing independent uncertainty estimates
Models tested: GPT-5.4, Claude Sonnet 4.6, Kimi 2.5, Qwen3.5, GLM-4.7
Belief Estimation
Stated beliefs: An LLM judge reads reasoning traces and assigns marginal probabilities that each label belongs to the best candidate.
Action beliefs: Derived from actions via smoothed marginal over proposed batches:
Measurement-Noise Correction
The authors identify and correct a critical bias: both diagnostics are computed from noisy belief estimates, introducing errors-in-variables biases that mislabel rational agents as overreacting.
For the Martingale slope, observed belief with noise energy :
For a rational agent with :
Correction: Use lagged belief as an instrument:
For belief movement, the per-step bias is estimated as where , and corrected statistic .
Decision Decomposition
Actions are decomposed into value (V) and uncertainty (U) features via conditional logit model:
where are z-standardized features from a shared Gaussian-process surrogate. Six statistical acquisition strategies serve as behavioral baselines: UCB, EI, Thompson sampling, -greedy, greedy, and directed evolution (DE).
Empirical Validation / Results
Prior Knowledge: Helpful but Unreliable
- Default and alias modes yield similar performance, indicating LLMs use chemical descriptors rather than memory of domain-specific terms
- Weak positive correlation between prior information and performance (Spearman , one-sided )
- Prior value relative to blind: ~1.0 batch (CI 95% [−0.21, 2.58])—high variance
- LLMs beat best statistical model on NFA () and tripeptide self-assembly () but fall short elsewhere
- No configuration decisively beats the mean statistical baseline across domains
Belief Dynamics: Overreaction, Not Entrenchment
- Stated and action beliefs agree well (pooled Pearson , CI 95% [+0.754, +0.764])
- All models, on all datasets, overreact to incoming data in all prompt modes (positive , negative at every step)
- Beliefs swing ~3× more on the first update as models move from prior-driven to data-driven beliefs
- Entrenchment is ruled out—models draw more information from new data than a rational agent would
Context-Stickiness: The Core Failure Mode
- Reasoning LLMs select more low-uncertainty candidates than any statistical model tested
- Non-reasoning LLMs match or are more uncertainty-avoiding than directed evolution
- Reasoning LLMs achieve less search-space coverage than DE
- Mean Hamming distance from proposals to history confirms reasoning LLMs select candidates close to those in context
- Reasoning enforces the bias against exploration rather than relieving it
The Intention-Ability Gap
- Reasoning LLMs invoke exploration in 92% of all traces and act on that intent
- However, correlation between stated exploration intent and realized exploration is zero for rarity and novelty, only weakly positive for coverage gain
- Providing uncertainty-informed tools (agent mode) makes all three correlations positive
- The failure is an intention-ability gap, not an intention-action gap
Causal Test: Removing History Restores Exploration
- Removing history from the agent's prompt moves behavior from uncertainty-avoiding to the U-range of a Thompson sampler
- No-history agents show strengthened correlation between exploration intent and realized information gain
- Performance effects are mixed and dataset-dependent, with correlation between improvement and UCB-DE difference (, CI 95% [−0.25, +0.96])
Rational-Agent Validation of Noise Correction
| Diagnostic | Estimator | Estimate | 95% CI | Prediction |
|---|---|---|---|---|
| Martingale | naive M | −0.0352 | [−0.0406, −0.0301] | −0.0397 |
| Martingale | +0.0021 | [−0.0029, +0.0068] | 0 | |
| Martingale | , edge-heavy | +0.0003 | [−0.0011, +0.0016] | 0 |
| Belief movement | naive Z | +0.0241 | [+0.0211, +0.0271] | +0.0250 (2) |
| Belief movement | −0.0013 | [−0.0041, +0.0015] | 0 | |
| Belief movement | , edge-heavy | −0.0005 | [−0.0029, +0.0020] | 0 |
The corrections successfully restore the null (zero) for a rational agent, confirming the naive estimators would mislabel rational agents as overreacting.
Theoretical and Practical Implications
Theoretical Contributions
-
Methodological: Belief- and uncertainty-dynamics analysis extended to combinatorial search spaces with marginalized belief updates provides finer resolution than prior LLM-BO work.
-
Noise correction: The measurement-noise correction removes a bias in Martingale-based belief diagnostics that otherwise flags rational agents as overreacting—reusable by anyone diagnosing belief dynamics.
-
Behavioral characterization: The paper resolves the entrenchment-vs-overreaction debate for scientific discovery domains: LLMs overreact to weak signals rather than entrenching on priors in rich-prior, weak-signal regimes.
Practical Implications
-
Competence gap vs. preference: The distinction matters for intervention design. Strategy/intention gaps respond to prompt instructions and reward shaping (RL fine-tuning), but competence gaps do not—explicit exploration instructions leave exploration metrics unchanged.
-
Context-stickiness is structural: Present regardless of whether the prior is correct or how strong it is, and orthogonal to domain prior.
-
Design guidance for LLM-BO pipelines: Successful LLM-BO systems should decouple LLM priors from LLM decision-making under uncertainty—either by converting LLM priors to statistical scores for use with statistical models, or by converting numerical data and uncertainty proxies into semantic text.
-
Nuance on feedback sensitivity: The claim that "LLM-based agents show no sensitivity to experimental feedback" (from label-shuffling experiments) may be misleading—LLMs are sensitive but context-sticky. Shuffling labels leaves the context set intact, so search strategy may not change.
-
Scaling caveat: More information in larger context windows does not simply scale capability; it may change model behavior, and not always positively—relevant to the single-large-context-agent vs. multi-agent debate.
Conclusion
This paper provides a comprehensive diagnostic of LLM behavior in sequential decision-making in prior- and data-heavy domains. The key findings:
-
Prior knowledge helps on average but is unreliable—high variance, occasionally harmful, and no configuration beats the mean statistical baseline across domains.
-
LLMs overreact to incoming data rather than entrenching on priors in scientific discovery contexts, with beliefs moving most dramatically in the first step.
-
Under-exploration is a competence gap, not a preference—LLMs sincerely intend to explore but cannot identify high-uncertainty regions due to context-stickiness.
-
Removing in-context history restores exploration, demonstrating that priors and data must be decoupled for effective LLM-driven discovery.
Future work should focus on designing methods that allow LLMs to reason rationally across data and semantic priors simultaneously. The study is limited to combinatorial search spaces and small budgets, but context-stickiness would likely occur in other domains given similar exploration failures in bandit and forecasting settings.
Related papers
- Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Frozen weights do not guarantee safe saturation in agentic coding; auditors must monitor scaffold expansion, not just checkpoint freezing.
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.
- Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.