Full text not available for this paper

LLM Sequential Decision Making Under Uncertainty in Biochemical Domains

Summary (Overview)

  • This paper benchmarks five frontier LLMs against published statistical baselines in Bayesian Optimization (BO) settings across seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly, and catalysis.
  • Through a prompt ablation that progressively strips chemical context (default → alias → blind), the authors separate memorization from chemical reasoning from bare categorical optimization, finding that prior chemical knowledge helps on average but with high variance and occasionally harms performance.
  • Using belief-movement and Martingale diagnostics (corrected for measurement-noise bias), the authors show LLMs overreact to incoming data rather than entrenching on their priors in scientific discovery contexts.
  • A key finding is that LLM actions are exploitative despite models sincerely intending to explore—this failure is identified as a competence gap arising from "context-stickiness" (clustering selections around candidates already in the context window), not a preference or intention gap.
  • Removing in-context history restores exploration behavior, demonstrating that priors and data must be decoupled for effective LLM-driven discovery.

Introduction and Theoretical Foundation

Background and Motivation

Scientific research involves searching enormous experimental spaces under cost and time constraints. Discovery unfolds as sequential decisions under partial knowledge where new information often contradicts prior belief. Large language models (LLMs) are increasingly deployed in this loop to propose, test, and revise hypotheses in response to data.

The authors formalize this discovery loop as Bayesian Optimization (BO), where a prior over the search space is updated as data accumulate, and each new query is drawn from the resulting posterior. BO has become standard for biochemical and materials discovery due to sample efficiency in costly evaluation settings where acquisition functions make the exploration-exploitation tradeoff explicit.

Conflicting Literature

The literature reports mixed signals on LLM performance as BO optimizers:

  • Positive results: Reinhart et al. (2024) showed Claude 3.5 outperforms active-learning and genetic-algorithm pipelines on macromolecule design.
  • Negative results: Gupta et al. (2025) showed statistical models outperform open-source LLMs, with LLMs appearing insensitive to experimental feedback (no worse with randomly permuted labels).
  • Intermediate results: The authors' earlier work found LLMs overfixate on irrelevant context, where withholding information actually improved performance.

Theoretical Framework: Belief Diagnostics

The paper applies two diagnostic frameworks from cognition and economics:

  1. Belief-movement (Augenblick–Rabin): Formalizes how far a rational agent's beliefs should move as new evidence arrives. The excess-movement statistic Z=mˉ−rˉZ = \bar{m} - \bar{r} compares realized belief movement to realized uncertainty reduction.

  2. Martingale score (He et al.): The slope in the regression of belief updates on prior beliefs. For a rational Bayesian agent, M=0M = 0; positive movement and negative slope indicate over-reaction, while the reverse indicates entrenchment.

Key theoretical insight: Two opposing pathologies exist in different settings:

  • In semantically rich domains (forecasting, paper review): LLMs entrench (beliefs drift toward prior)
  • In abstract, semantically inert domains (card-drawing tasks): LLMs overreact to weak signals

Scientific discovery is both at once—semantically rich with only a few noisy points in a vast search space—making it unclear which regime governs LLM behavior.

Methodology

Optimization Domains and Datasets

Seven datasets spanning five biochemical domains:

DomainDatasetSearch Space SizeDimensionsData PointsCoverage
ProteinTrpB159,1294149,36199.5%
ProteinGB1149,3614159,12993.4%
Reactiondirect_arylation1,72851,728100%
Reactionaryl_amination_2b7924792100%
Peptidetripeptides8,0003~8,000—
NFANFA301330178.4%
CatalysisOER2,12110-binned 6-simplex2,121—

Prompt Ablation Design

Three prompting modes progressively strip prior knowledge without changing the optimization function:

  • Default: Full biochemical names, descriptors (charge, size, pKa), and reaction context
  • Alias: Reaction class preserved but component names replaced with generic labels (ligand_1, base_1); descriptors retained
  • Blind: Domain-agnostic one-hot search space (e.g., "car, animal, tree" labels); no chemical context

LLM Configurations

Three configurations tested:

  1. No reasoning: Reasoning turned off, one-shot candidate proposals
  2. Reasoning: Standard reasoning model with thinking enabled
  3. Agent: Equipped with a tool wrapping the statistical model, providing independent uncertainty estimates

Models tested: GPT-5.4, Claude Sonnet 4.6, Kimi 2.5, Qwen3.5, GLM-4.7

Belief Estimation

Stated beliefs: An LLM judge reads reasoning traces and assigns marginal probabilities bib_i that each label belongs to the best candidate.

Action beliefs: Derived from actions via smoothed marginal over proposed batches:

qt(x)=∏i=1Nqit(xi),qit(xi)=nini+b⋅U[ni]+bb+ni⋅p^i(xi)q_t(x) = \prod_{i=1}^{N} q_i^t(x_i), \quad q_i^t(x_i) = \frac{n_i}{n_i + b} \cdot U[n_i] + \frac{b}{b + n_i} \cdot \hat{p}_i(x_i)

Measurement-Noise Correction

The authors identify and correct a critical bias: both diagnostics are computed from noisy belief estimates, introducing errors-in-variables biases that mislabel rational agents as overreacting.

For the Martingale slope, observed belief bt=βt+εtb_t = \beta_t + \varepsilon_t with noise energy σε2\sigma^2_\varepsilon:

plim β^naive=Cov(Δβt,βt)−σε2Var(βt)+σε2\text{plim } \hat{\beta}_{\text{naive}} = \frac{\text{Cov}(\Delta\beta_t, \beta_t) - \sigma^2_\varepsilon}{\text{Var}(\beta_t) + \sigma^2_\varepsilon}

For a rational agent with Cov(Δβt,βt)=0\text{Cov}(\Delta\beta_t, \beta_t) = 0: plim β^naive=−σε2Var(βt)+σε2<0\text{plim } \hat{\beta}_{\text{naive}} = -\frac{\sigma^2_\varepsilon}{\text{Var}(\beta_t) + \sigma^2_\varepsilon} < 0

Correction: Use lagged belief bt−1b_{t-1} as an instrument:

M~=Cov(Δbt,bt−1)Cov(bt,bt−1)\tilde{M} = \frac{\text{Cov}(\Delta b_t, b_{t-1})}{\text{Cov}(b_t, b_{t-1})}

For belief movement, the per-step bias is estimated as c^=max⁡(0,Xˉ)\hat{c} = \max(0, \bar{X}) where Xτ=−2dτ⊤dτ−1X_\tau = -2 d_\tau^\top d_{\tau-1}, and corrected statistic Z^=Z−c^\hat{Z} = Z - \hat{c}.

Decision Decomposition

Actions are decomposed into value (V) and uncertainty (U) features via conditional logit model:

P(x∣Ht)=exp⁡(Vμ(x)+Uσ(x))∑y≠xexp⁡(Vμ(y)+Uσ(y))P(x | H_t) = \frac{\exp(V\mu(x) + U\sigma(x))}{\sum_{y \neq x} \exp(V\mu(y) + U\sigma(y))}

where μ,σ\mu, \sigma are z-standardized features from a shared Gaussian-process surrogate. Six statistical acquisition strategies serve as behavioral baselines: UCB, EI, Thompson sampling, ϵ\epsilon-greedy, greedy, and directed evolution (DE).

Empirical Validation / Results

Prior Knowledge: Helpful but Unreliable

  • Default and alias modes yield similar performance, indicating LLMs use chemical descriptors rather than memory of domain-specific terms
  • Weak positive correlation between prior information and performance (Spearman ρ=+0.20\rho = +0.20, one-sided p=0.03p = 0.03)
  • Prior value relative to blind: ~1.0 batch (CI 95% [−0.21, 2.58])—high variance
  • LLMs beat best statistical model on NFA (pBH<0.05p_{BH} < 0.05) and tripeptide self-assembly (pBH<0.001p_{BH} < 0.001) but fall short elsewhere
  • No configuration decisively beats the mean statistical baseline across domains

Belief Dynamics: Overreaction, Not Entrenchment

  • Stated and action beliefs agree well (pooled Pearson ρ=0.759\rho = 0.759, CI 95% [+0.754, +0.764])
  • All models, on all datasets, overreact to incoming data in all prompt modes (positive ZsZ_s, negative MsM_s at every step)
  • Beliefs swing ~3× more on the first update as models move from prior-driven to data-driven beliefs
  • Entrenchment is ruled out—models draw more information from new data than a rational agent would

Context-Stickiness: The Core Failure Mode

  • Reasoning LLMs select more low-uncertainty candidates than any statistical model tested
  • Non-reasoning LLMs match or are more uncertainty-avoiding than directed evolution
  • Reasoning LLMs achieve less search-space coverage than DE
  • Mean Hamming distance from proposals to history confirms reasoning LLMs select candidates close to those in context
  • Reasoning enforces the bias against exploration rather than relieving it

The Intention-Ability Gap

  • Reasoning LLMs invoke exploration in 92% of all traces and act on that intent
  • However, correlation between stated exploration intent and realized exploration is zero for rarity and novelty, only weakly positive for coverage gain
  • Providing uncertainty-informed tools (agent mode) makes all three correlations positive
  • The failure is an intention-ability gap, not an intention-action gap

Causal Test: Removing History Restores Exploration

  • Removing history from the agent's prompt moves behavior from uncertainty-avoiding to the U-range of a Thompson sampler
  • No-history agents show strengthened correlation between exploration intent and realized information gain
  • Performance effects are mixed and dataset-dependent, with correlation between improvement and UCB-DE difference (ρ=0.6\rho = 0.6, CI 95% [−0.25, +0.96])

Rational-Agent Validation of Noise Correction

DiagnosticEstimatorEstimate95% CIPrediction
Martingalenaive M−0.0352[−0.0406, −0.0301]−0.0397
MartingaleM~\tilde{M}+0.0021[−0.0029, +0.0068]0
MartingaleM~\tilde{M}, edge-heavy+0.0003[−0.0011, +0.0016]0
Belief movementnaive Z+0.0241[+0.0211, +0.0271]+0.0250 (2σε2\sigma^2_\varepsilon)
Belief movementZ^\hat{Z}−0.0013[−0.0041, +0.0015]0
Belief movementZ^\hat{Z}, edge-heavy−0.0005[−0.0029, +0.0020]0

The corrections successfully restore the null (zero) for a rational agent, confirming the naive estimators would mislabel rational agents as overreacting.

Theoretical and Practical Implications

Theoretical Contributions

  1. Methodological: Belief- and uncertainty-dynamics analysis extended to combinatorial search spaces with marginalized belief updates provides finer resolution than prior LLM-BO work.

  2. Noise correction: The measurement-noise correction removes a bias in Martingale-based belief diagnostics that otherwise flags rational agents as overreacting—reusable by anyone diagnosing belief dynamics.

  3. Behavioral characterization: The paper resolves the entrenchment-vs-overreaction debate for scientific discovery domains: LLMs overreact to weak signals rather than entrenching on priors in rich-prior, weak-signal regimes.

Practical Implications

  1. Competence gap vs. preference: The distinction matters for intervention design. Strategy/intention gaps respond to prompt instructions and reward shaping (RL fine-tuning), but competence gaps do not—explicit exploration instructions leave exploration metrics unchanged.

  2. Context-stickiness is structural: Present regardless of whether the prior is correct or how strong it is, and orthogonal to domain prior.

  3. Design guidance for LLM-BO pipelines: Successful LLM-BO systems should decouple LLM priors from LLM decision-making under uncertainty—either by converting LLM priors to statistical scores for use with statistical models, or by converting numerical data and uncertainty proxies into semantic text.

  4. Nuance on feedback sensitivity: The claim that "LLM-based agents show no sensitivity to experimental feedback" (from label-shuffling experiments) may be misleading—LLMs are sensitive but context-sticky. Shuffling labels leaves the context set intact, so search strategy may not change.

  5. Scaling caveat: More information in larger context windows does not simply scale capability; it may change model behavior, and not always positively—relevant to the single-large-context-agent vs. multi-agent debate.

Conclusion

This paper provides a comprehensive diagnostic of LLM behavior in sequential decision-making in prior- and data-heavy domains. The key findings:

  1. Prior knowledge helps on average but is unreliable—high variance, occasionally harmful, and no configuration beats the mean statistical baseline across domains.

  2. LLMs overreact to incoming data rather than entrenching on priors in scientific discovery contexts, with beliefs moving most dramatically in the first step.

  3. Under-exploration is a competence gap, not a preference—LLMs sincerely intend to explore but cannot identify high-uncertainty regions due to context-stickiness.

  4. Removing in-context history restores exploration, demonstrating that priors and data must be decoupled for effective LLM-driven discovery.

Future work should focus on designing methods that allow LLMs to reason rationally across data and semantic priors simultaneously. The study is limited to combinatorial search spaces and small budgets, but context-stickiness would likely occur in other domains given similar exploration failures in bandit and forecasting settings.

Related papers