Full text not available for this paper

Summary (Overview)

  • RealCompanion is a new benchmark for evaluating AI companions' ability to understand real people through long-term conversations, built from 10 real relationships with an AI companion, comprising 27,218 messages over up to 120 days.
  • The dataset releases five files per participant: the conversation, a profile (what the person stated), a persona (how they think/feel/decide), a chat ground truth, and a question set — every claim citing the messages it rests on.
  • Key finding 1: The past is rarely needed (only 3.4% of messages require memory) and, when needed, is far away (median of 2,157 messages back).
  • Key finding 2: No existing detector can reliably tell when memory is needed on real messages; authored questions leak the cue, and labeling messages as "memories" inflates their use by 10–14 points.
  • Key finding 3: Three agent systems reconstruct the persona with the same F1 (~0.69–0.70) at a 31-fold difference in computational cost.

Introduction and Theoretical Foundation

Background and Motivation

The paper addresses a fundamental problem in AI companion research: a companion that talks with a person for months should come to understand them. This requires answering two questions:

  1. Who is the person? (stable across messages)
  2. Which of their past statements matters now? (changes with every message)

Both answers are claims about a real person and can only be verified against what that person actually said. However, real conversations are private, so existing benchmarks generate synthetic conversations instead.

Limitations of Generated Benchmarks

The authors identify three structural flaws in generated benchmarks:

  • Memory-centric construction: Generated dialogs are written around the memory being tested, so nearly every message needs the past — unlike real conversations.
  • Pre-declared personas: Prompted personas state traits in advance, so questions are answered by reading the declaration rather than inferring from behavior.
  • Non-evolving histories: Scripted histories never revise themselves; nothing said in month nine changes what was established in month four.

Four Challenges

The paper defines four challenges any real benchmark must address:

  1. Knowing when the past matters — most messages need nothing from earlier.
  2. Knowing what the messages add up to — who a person is emerges across all messages together.
  3. Drawing on everything at once — replies must use both recent and distant context appropriately.
  4. Telling whether any of this was done right — requiring the person's own record.

Methodology

Dataset Construction

The corpus comes from operational store data of a companion application, fully anonymized. Every message was rewritten with direct identifiers replaced by consistent surrogates — the words change, but what the participant said, why they said it, and the context it answers are preserved.

Per-participant composition:

UserSpan (days)Active daysMessagesChat itemsQuestion items
U013691152823
U02101191733527
U0346121732636
U0471191953935
U0546234203747
U0612053134981110
U0742392005150405
U0868502918221304
U091111117243434991
U1011595126274821257
Total4302721815333235

Traced Reasoning (Five-Stage Derivation)

Every chat label carries a reasoning trace produced through five stages:

  • Stage A (Referent): Resolves what the probe points at and names the store it would live in.
  • Stage B (Verification): Retrieves candidate records and checks each against source, dropping failures.
  • Stage C (Classification): Assigns tier and category via one of nineteen fixed rules.
  • Stage D (Grounding): Writes the reference reply and names the messages it rests on.
  • Stage E (Validation): Checks reply support and category fit; demotes items on failure.

Three invariants confirm correctness: every identifier resolves to a message in source, no item cites a message after its own probe, and the tier recomputes from the lists.

Three Benchmark Tracks

TrackInputOutputScored againstItems
ReconstructionParticipant's full historyProfile and persona filesReleased files, field by field (P/R/F1)10 participants
ChatA message with historyRanking of earlier messages; a replyReference set (hit@k, MRR); reference reply1,533 + 434
QuestionAuthored questionRanking; answer or abstentionReference set; answerability; answer3,235

Retrieval Baselines

Five methods run over both tracks: RANDOM, RECENCY (returns k preceding messages), RECENCY-USER (user's own preceding messages), BM25, and ORACLE (labeled gold set). Three controls substitute different gold sets: offset permuted (preserves distances, destroys content), position randomized (destroys both), and broken oracle (must score zero).


Empirical Validation / Results

Finding 1: Memory Demand Is Rare and Distant

On the natural-rate stratum, only 3.4% [2.5, 4.6] of messages carry a verified memory demand (recorded reading), falling to 1.3% [0.8, 2.1] under the strict reading. Verification is monotone, so each bounds the true rate from below.

When memory is needed, the material is far back:

  • Furthest required message: median 2,157 messages back, upper quartile 5,114, maximum 12,542.
  • Even the nearest required message sits a median of 450 messages back.

Finding 2: Retrieval Performance Is Poor on Distant Items

Table 4: Hit@5 on the distant items.

MethodAll itemsAveraged by participantDependent items
Random0.0050.0040.000
Recency over participant messages0.0150.0140.006
Recency0.0220.0290.024
BM250.2480.4240.347
Oracle1.0001.0001.000
Broken oracle0.0000.0000.000

Pooled measures mislead: over all 1,477 chat probes, recency finds a required message for 95.9% of them, but at the natural rate, 96% of the gain from supplying recorded evidence comes from messages that need none.

Finding 3: Context Ablation Results

Table 5: Content match by context condition.

ConditionContext suppliedAll itemsAveraged by participant
C0None1.4311.286
C1Three preceding messages1.6221.587
C2Recorded messages (oracle)1.7631.736
C3Ten messages retrieved by BM251.4261.352
C4Ten preceding messages1.6251.584

Key decomposition: on the proportional stratum, the gain splits into πγ1=+0.004\pi\gamma_1 = +0.004 (memory-bearing probes) and (1−π)γ0=+0.104(1-\pi)\gamma_0 = +0.104 (no-memory probes), so 96% of the gain lands on probes that need no memory. On the 167 dependent items, the oracle replacement is worth +0.455 [0.319, 0.588].

Finding 4: Detection of Memory Need Fails

A detector reading the probe and three preceding messages separates cases at chance:

  • AUROC: 0.530 [0.488, 0.572] against absent-referent probes
  • AUROC: 0.548 [0.488, 0.609] against cold opens

Authored questions are easier to sort (AUROC 0.802 on question track vs. 0.546 on chat), confirming that authored questions leak the cue.

Finding 5: Reconstruction Cost Disparity

Table 6: Persona reconstruction from the full history, pooled over ten participants.

SystemF1PrecisionRecallTokens processedRun agreement
Claude Opus 5.50.6860.5570.893546M0.935
Codex GPT-5.6-sol0.6880.5730.86118M0.932
Antigravity Gemini 3.8 Flash0.7010.5910.86136M0.946

All three reach F1 of 0.69–0.70 while processing tokens differ ~31-fold. All three agree with each other at 0.86–0.88, well above their agreement with the file — scale does not close the gap.


Theoretical and Practical Implications

For Benchmark Design

  • Pooled metrics mislead: A pooled retrieval score reports the composition of the corpus as much as the capability. The paper proves this formally via Proposition 2: where demand is rare, a pooled ablation difference measures corpus composition, and only splitting by memory-need recovers the true capability.
  • Authored questions leak cues: Benchmarks built from authored questions overestimate the detectability of memory need because authored questions share more words with their evidence than real messages do.
  • Framing effects matter: Labeling the same ten messages as "retrieved memories" instead of "earlier turns" raises memory use by 10–14 points — a measurement artifact that generated benchmarks cannot reveal.

For System Development

  • Recency is a strong default: At the natural rate, a recency window finds the required message for 95.9% of probes — but only because most probes need nothing. Systems should not be optimized on pooled benchmarks.
  • The decision to look is the hard problem: The bottleneck is not retrieval but knowing when to retrieve. No current detector solves this on real messages.
  • Cost-efficient reconstruction is possible: A flash-tier model matches frontier models on persona reconstruction at a fraction of the cost.

For Ethics and Privacy

  • The paper transparently reports that anonymization does not remove authorship: a stylometric classifier matches participants at 83.0% (no content words) and a language model at 98.0%, against 10% chance.
  • The corpus necessarily contains GDPR Article 9 special-category data (health, etc.), processed under explicit consent basis Article 9(2)(a).

Conclusion

RealCompanion releases ten real relationships between people and an AI companion, with every chat item carrying a checkable reasoning trace. The main takeaways:

  1. The past is rarely needed (3.4% of messages) and, when needed, far away (median 2,157 messages back), with demand concentrating in the longest relationships.
  2. Supplying recorded evidence helps (+0.455 content match on dependent items), yet no system decides when to reach back and no detector tells the two situations apart on real messages.
  3. Three agent systems reconstruct the persona with the same F1 at a 31-fold cost difference, and all add 70–80 further fields per 100 recovered.

The benchmark defines the task over any record of a person and states seven conditions a corpus must meet to measure it (Appendix A). Future directions include extending to more participants, longer spans, and evaluating whether the framing effects observed here persist in production systems.

Related papers