Full text not available for this paper
Summary (Overview)
- RealCompanion is a new benchmark for evaluating AI companions' ability to understand real people through long-term conversations, built from 10 real relationships with an AI companion, comprising 27,218 messages over up to 120 days.
- The dataset releases five files per participant: the conversation, a profile (what the person stated), a persona (how they think/feel/decide), a chat ground truth, and a question set — every claim citing the messages it rests on.
- Key finding 1: The past is rarely needed (only 3.4% of messages require memory) and, when needed, is far away (median of 2,157 messages back).
- Key finding 2: No existing detector can reliably tell when memory is needed on real messages; authored questions leak the cue, and labeling messages as "memories" inflates their use by 10–14 points.
- Key finding 3: Three agent systems reconstruct the persona with the same F1 (~0.69–0.70) at a 31-fold difference in computational cost.
Introduction and Theoretical Foundation
Background and Motivation
The paper addresses a fundamental problem in AI companion research: a companion that talks with a person for months should come to understand them. This requires answering two questions:
- Who is the person? (stable across messages)
- Which of their past statements matters now? (changes with every message)
Both answers are claims about a real person and can only be verified against what that person actually said. However, real conversations are private, so existing benchmarks generate synthetic conversations instead.
Limitations of Generated Benchmarks
The authors identify three structural flaws in generated benchmarks:
- Memory-centric construction: Generated dialogs are written around the memory being tested, so nearly every message needs the past — unlike real conversations.
- Pre-declared personas: Prompted personas state traits in advance, so questions are answered by reading the declaration rather than inferring from behavior.
- Non-evolving histories: Scripted histories never revise themselves; nothing said in month nine changes what was established in month four.
Four Challenges
The paper defines four challenges any real benchmark must address:
- Knowing when the past matters — most messages need nothing from earlier.
- Knowing what the messages add up to — who a person is emerges across all messages together.
- Drawing on everything at once — replies must use both recent and distant context appropriately.
- Telling whether any of this was done right — requiring the person's own record.
Methodology
Dataset Construction
The corpus comes from operational store data of a companion application, fully anonymized. Every message was rewritten with direct identifiers replaced by consistent surrogates — the words change, but what the participant said, why they said it, and the context it answers are preserved.
Per-participant composition:
| User | Span (days) | Active days | Messages | Chat items | Question items |
|---|---|---|---|---|---|
| U01 | 36 | 9 | 115 | 28 | 23 |
| U02 | 101 | 19 | 173 | 35 | 27 |
| U03 | 46 | 12 | 173 | 26 | 36 |
| U04 | 71 | 19 | 195 | 39 | 35 |
| U05 | 46 | 23 | 420 | 37 | 47 |
| U06 | 120 | 53 | 1349 | 81 | 110 |
| U07 | 42 | 39 | 2005 | 150 | 405 |
| U08 | 68 | 50 | 2918 | 221 | 304 |
| U09 | 111 | 111 | 7243 | 434 | 991 |
| U10 | 115 | 95 | 12627 | 482 | 1257 |
| Total | 430 | 27218 | 1533 | 3235 |
Traced Reasoning (Five-Stage Derivation)
Every chat label carries a reasoning trace produced through five stages:
- Stage A (Referent): Resolves what the probe points at and names the store it would live in.
- Stage B (Verification): Retrieves candidate records and checks each against source, dropping failures.
- Stage C (Classification): Assigns tier and category via one of nineteen fixed rules.
- Stage D (Grounding): Writes the reference reply and names the messages it rests on.
- Stage E (Validation): Checks reply support and category fit; demotes items on failure.
Three invariants confirm correctness: every identifier resolves to a message in source, no item cites a message after its own probe, and the tier recomputes from the lists.
Three Benchmark Tracks
| Track | Input | Output | Scored against | Items |
|---|---|---|---|---|
| Reconstruction | Participant's full history | Profile and persona files | Released files, field by field (P/R/F1) | 10 participants |
| Chat | A message with history | Ranking of earlier messages; a reply | Reference set (hit@k, MRR); reference reply | 1,533 + 434 |
| Question | Authored question | Ranking; answer or abstention | Reference set; answerability; answer | 3,235 |
Retrieval Baselines
Five methods run over both tracks: RANDOM, RECENCY (returns k preceding messages), RECENCY-USER (user's own preceding messages), BM25, and ORACLE (labeled gold set). Three controls substitute different gold sets: offset permuted (preserves distances, destroys content), position randomized (destroys both), and broken oracle (must score zero).
Empirical Validation / Results
Finding 1: Memory Demand Is Rare and Distant
On the natural-rate stratum, only 3.4% [2.5, 4.6] of messages carry a verified memory demand (recorded reading), falling to 1.3% [0.8, 2.1] under the strict reading. Verification is monotone, so each bounds the true rate from below.
When memory is needed, the material is far back:
- Furthest required message: median 2,157 messages back, upper quartile 5,114, maximum 12,542.
- Even the nearest required message sits a median of 450 messages back.
Finding 2: Retrieval Performance Is Poor on Distant Items
Table 4: Hit@5 on the distant items.
| Method | All items | Averaged by participant | Dependent items |
|---|---|---|---|
| Random | 0.005 | 0.004 | 0.000 |
| Recency over participant messages | 0.015 | 0.014 | 0.006 |
| Recency | 0.022 | 0.029 | 0.024 |
| BM25 | 0.248 | 0.424 | 0.347 |
| Oracle | 1.000 | 1.000 | 1.000 |
| Broken oracle | 0.000 | 0.000 | 0.000 |
Pooled measures mislead: over all 1,477 chat probes, recency finds a required message for 95.9% of them, but at the natural rate, 96% of the gain from supplying recorded evidence comes from messages that need none.
Finding 3: Context Ablation Results
Table 5: Content match by context condition.
| Condition | Context supplied | All items | Averaged by participant |
|---|---|---|---|
| C0 | None | 1.431 | 1.286 |
| C1 | Three preceding messages | 1.622 | 1.587 |
| C2 | Recorded messages (oracle) | 1.763 | 1.736 |
| C3 | Ten messages retrieved by BM25 | 1.426 | 1.352 |
| C4 | Ten preceding messages | 1.625 | 1.584 |
Key decomposition: on the proportional stratum, the gain splits into (memory-bearing probes) and (no-memory probes), so 96% of the gain lands on probes that need no memory. On the 167 dependent items, the oracle replacement is worth +0.455 [0.319, 0.588].
Finding 4: Detection of Memory Need Fails
A detector reading the probe and three preceding messages separates cases at chance:
- AUROC: 0.530 [0.488, 0.572] against absent-referent probes
- AUROC: 0.548 [0.488, 0.609] against cold opens
Authored questions are easier to sort (AUROC 0.802 on question track vs. 0.546 on chat), confirming that authored questions leak the cue.
Finding 5: Reconstruction Cost Disparity
Table 6: Persona reconstruction from the full history, pooled over ten participants.
| System | F1 | Precision | Recall | Tokens processed | Run agreement |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 0.686 | 0.557 | 0.893 | 546M | 0.935 |
| Codex GPT-5.6-sol | 0.688 | 0.573 | 0.861 | 18M | 0.932 |
| Antigravity Gemini 3.8 Flash | 0.701 | 0.591 | 0.861 | 36M | 0.946 |
All three reach F1 of 0.69–0.70 while processing tokens differ ~31-fold. All three agree with each other at 0.86–0.88, well above their agreement with the file — scale does not close the gap.
Theoretical and Practical Implications
For Benchmark Design
- Pooled metrics mislead: A pooled retrieval score reports the composition of the corpus as much as the capability. The paper proves this formally via Proposition 2: where demand is rare, a pooled ablation difference measures corpus composition, and only splitting by memory-need recovers the true capability.
- Authored questions leak cues: Benchmarks built from authored questions overestimate the detectability of memory need because authored questions share more words with their evidence than real messages do.
- Framing effects matter: Labeling the same ten messages as "retrieved memories" instead of "earlier turns" raises memory use by 10–14 points — a measurement artifact that generated benchmarks cannot reveal.
For System Development
- Recency is a strong default: At the natural rate, a recency window finds the required message for 95.9% of probes — but only because most probes need nothing. Systems should not be optimized on pooled benchmarks.
- The decision to look is the hard problem: The bottleneck is not retrieval but knowing when to retrieve. No current detector solves this on real messages.
- Cost-efficient reconstruction is possible: A flash-tier model matches frontier models on persona reconstruction at a fraction of the cost.
For Ethics and Privacy
- The paper transparently reports that anonymization does not remove authorship: a stylometric classifier matches participants at 83.0% (no content words) and a language model at 98.0%, against 10% chance.
- The corpus necessarily contains GDPR Article 9 special-category data (health, etc.), processed under explicit consent basis Article 9(2)(a).
Conclusion
RealCompanion releases ten real relationships between people and an AI companion, with every chat item carrying a checkable reasoning trace. The main takeaways:
- The past is rarely needed (3.4% of messages) and, when needed, far away (median 2,157 messages back), with demand concentrating in the longest relationships.
- Supplying recorded evidence helps (+0.455 content match on dependent items), yet no system decides when to reach back and no detector tells the two situations apart on real messages.
- Three agent systems reconstruct the persona with the same F1 at a 31-fold cost difference, and all add 70–80 further fields per 100 recovered.
The benchmark defines the task over any record of a person and states seven conditions a corpus must meet to measure it (Appendix A). Future directions include extending to more participants, longer spans, and evaluating whether the framing effects observed here persist in production systems.
Related papers
- When do data mixtures improve scaling laws? Insights from high-dimensional regression
Data mixtures provably accelerate scaling laws only when auxiliary data has heavier-tailed spectra and intermediate relative sample growth, with ridge regression achieving the optimal rate.
- EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH shows that expanding an agent's tool, skill, or agent harness alone can degrade performance on previously solved tasks, inducing forgetting up to 34.7%.
- ScAn-Bench: Evaluating Scaling Analysis Methodology
ScAn-Bench introduces the first surrogate benchmarks for scaling analysis, revealing no universally optimal data acquisition and extrapolation strategy exists across LLMs and VLMs.