# RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

> RealCompanion, a benchmark from 10 real AI-companion relationships, shows memory is rarely needed (3.4% of messages) and no detector can reliably identify when it is.

- **Source:** [arXiv](https://arxiv.org/abs/2610.01780)
- **Published:** 2026-10-06
- **Permalink:** https://picx.dev/p/deh2Mg
- **Whiteboard:** https://picx.dev/p/deh2Mg/image

## Summary

## Summary (Overview)

- **RealCompanion** is a new benchmark for evaluating AI companions' ability to understand real people through long-term conversations, built from **10 real relationships** with an AI companion, comprising **27,218 messages** over up to **120 days**.
- The dataset releases five files per participant: the conversation, a profile (what the person stated), a persona (how they think/feel/decide), a chat ground truth, and a question set — every claim citing the messages it rests on.
- **Key finding 1**: The past is rarely needed (only 3.4% of messages require memory) and, when needed, is far away (median of 2,157 messages back).
- **Key finding 2**: No existing detector can reliably tell when memory is needed on real messages; authored questions leak the cue, and labeling messages as "memories" inflates their use by 10–14 points.
- **Key finding 3**: Three agent systems reconstruct the persona with the same F1 (~0.69–0.70) at a **31-fold difference in computational cost**.

---

## Introduction and Theoretical Foundation

### Background and Motivation

The paper addresses a fundamental problem in AI companion research: **a companion that talks with a person for months should come to understand them**. This requires answering two questions:

1. **Who is the person?** (stable across messages)
2. **Which of their past statements matters now?** (changes with every message)

Both answers are claims about a real person and can only be verified against what that person actually said. However, real conversations are private, so existing benchmarks generate synthetic conversations instead.

### Limitations of Generated Benchmarks

The authors identify three structural flaws in generated benchmarks:

- **Memory-centric construction**: Generated dialogs are written *around* the memory being tested, so nearly every message needs the past — unlike real conversations.
- **Pre-declared personas**: Prompted personas state traits in advance, so questions are answered by reading the declaration rather than inferring from behavior.
- **Non-evolving histories**: Scripted histories never revise themselves; nothing said in month nine changes what was established in month four.

### Four Challenges

The paper defines four challenges any real benchmark must address:

1. **Knowing when the past matters** — most messages need nothing from earlier.
2. **Knowing what the messages add up to** — who a person is emerges across all messages together.
3. **Drawing on everything at once** — replies must use both recent and distant context appropriately.
4. **Telling whether any of this was done right** — requiring the person's own record.

---

## Methodology

### Dataset Construction

The corpus comes from **operational store data** of a companion application, fully anonymized. Every message was rewritten with direct identifiers replaced by consistent surrogates — the words change, but what the participant said, why they said it, and the context it answers are preserved.

**Per-participant composition:**

| User | Span (days) | Active days | Messages | Chat items | Question items |
|------|-------------|-------------|----------|------------|----------------|
| U01  | 36          | 9           | 115      | 28         | 23             |
| U02  | 101         | 19          | 173      | 35         | 27             |
| U03  | 46          | 12          | 173      | 26         | 36             |
| U04  | 71          | 19          | 195      | 39         | 35             |
| U05  | 46          | 23          | 420      | 37         | 47             |
| U06  | 120         | 53          | 1349     | 81         | 110            |
| U07  | 42          | 39          | 2005     | 150        | 405            |
| U08  | 68          | 50          | 2918     | 221        | 304            |
| U09  | 111         | 111         | 7243     | 434        | 991            |
| U10  | 115         | 95          | 12627    | 482        | 1257           |
| **Total** | **430** | | **27218** | **1533** | **3235** |

### Traced Reasoning (Five-Stage Derivation)

Every chat label carries a **reasoning trace** produced through five stages:

- **Stage A (Referent)**: Resolves what the probe points at and names the store it would live in.
- **Stage B (Verification)**: Retrieves candidate records and checks each against source, dropping failures.
- **Stage C (Classification)**: Assigns tier and category via one of nineteen fixed rules.
- **Stage D (Grounding)**: Writes the reference reply and names the messages it rests on.
- **Stage E (Validation)**: Checks reply support and category fit; demotes items on failure.

Three invariants confirm correctness: every identifier resolves to a message in source, no item cites a message after its own probe, and the tier recomputes from the lists.

### Three Benchmark Tracks

| Track | Input | Output | Scored against | Items |
|-------|-------|--------|----------------|-------|
| **Reconstruction** | Participant's full history | Profile and persona files | Released files, field by field (P/R/F1) | 10 participants |
| **Chat** | A message with history | Ranking of earlier messages; a reply | Reference set (hit@k, MRR); reference reply | 1,533 + 434 |
| **Question** | Authored question | Ranking; answer or abstention | Reference set; answerability; answer | 3,235 |

### Retrieval Baselines

Five methods run over both tracks: **RANDOM**, **RECENCY** (returns k preceding messages), **RECENCY-USER** (user's own preceding messages), **BM25**, and **ORACLE** (labeled gold set). Three controls substitute different gold sets: offset permuted (preserves distances, destroys content), position randomized (destroys both), and broken oracle (must score zero).

---

## Empirical Validation / Results

### Finding 1: Memory Demand Is Rare and Distant

On the natural-rate stratum, only **3.4% [2.5, 4.6]** of messages carry a verified memory demand (recorded reading), falling to **1.3% [0.8, 2.1]** under the strict reading. Verification is monotone, so each bounds the true rate from below.

When memory *is* needed, the material is far back:
- Furthest required message: **median 2,157 messages back**, upper quartile 5,114, maximum 12,542.
- Even the **nearest** required message sits a median of 450 messages back.

### Finding 2: Retrieval Performance Is Poor on Distant Items

**Table 4: Hit@5 on the distant items.**

| Method | All items | Averaged by participant | Dependent items |
|--------|-----------|------------------------|-----------------|
| Random | 0.005 | 0.004 | 0.000 |
| Recency over participant messages | 0.015 | 0.014 | 0.006 |
| Recency | 0.022 | 0.029 | 0.024 |
| BM25 | 0.248 | 0.424 | 0.347 |
| Oracle | 1.000 | 1.000 | 1.000 |
| Broken oracle | 0.000 | 0.000 | 0.000 |

Pooled measures mislead: over all 1,477 chat probes, recency finds a required message for **95.9%** of them, but at the natural rate, **96% of the gain** from supplying recorded evidence comes from messages that need none.

### Finding 3: Context Ablation Results

**Table 5: Content match by context condition.**

| Condition | Context supplied | All items | Averaged by participant |
|-----------|------------------|-----------|------------------------|
| C0 | None | 1.431 | 1.286 |
| C1 | Three preceding messages | 1.622 | 1.587 |
| C2 | Recorded messages (oracle) | 1.763 | 1.736 |
| C3 | Ten messages retrieved by BM25 | 1.426 | 1.352 |
| C4 | Ten preceding messages | 1.625 | 1.584 |

Key decomposition: on the proportional stratum, the gain splits into $\pi\gamma_1 = +0.004$ (memory-bearing probes) and $(1-\pi)\gamma_0 = +0.104$ (no-memory probes), so **96% of the gain lands on probes that need no memory**. On the 167 dependent items, the oracle replacement is worth **+0.455 [0.319, 0.588]**.

### Finding 4: Detection of Memory Need Fails

A detector reading the probe and three preceding messages separates cases at chance:
- AUROC: **0.530 [0.488, 0.572]** against absent-referent probes
- AUROC: **0.548 [0.488, 0.609]** against cold opens

Authored questions are easier to sort (AUROC 0.802 on question track vs. 0.546 on chat), confirming that **authored questions leak the cue**.

### Finding 5: Reconstruction Cost Disparity

**Table 6: Persona reconstruction from the full history, pooled over ten participants.**

| System | F1 | Precision | Recall | Tokens processed | Run agreement |
|--------|-----|-----------|--------|------------------|---------------|
| Claude Opus 5.5 | 0.686 | 0.557 | 0.893 | 546M | 0.935 |
| Codex GPT-5.6-sol | 0.688 | 0.573 | 0.861 | 18M | 0.932 |
| Antigravity Gemini 3.8 Flash | 0.701 | 0.591 | 0.861 | 36M | 0.946 |

All three reach F1 of 0.69–0.70 while processing tokens differ **~31-fold**. All three agree with each other at 0.86–0.88, well above their agreement with the file — **scale does not close the gap**.

---

## Theoretical and Practical Implications

### For Benchmark Design

- **Pooled metrics mislead**: A pooled retrieval score reports the composition of the corpus as much as the capability. The paper proves this formally via **Proposition 2**: where demand is rare, a pooled ablation difference measures corpus composition, and only splitting by memory-need recovers the true capability.
- **Authored questions leak cues**: Benchmarks built from authored questions overestimate the detectability of memory need because authored questions share more words with their evidence than real messages do.
- **Framing effects matter**: Labeling the same ten messages as "retrieved memories" instead of "earlier turns" raises memory use by 10–14 points — a measurement artifact that generated benchmarks cannot reveal.

### For System Development

- **Recency is a strong default**: At the natural rate, a recency window finds the required message for 95.9% of probes — but only because most probes need nothing. Systems should not be optimized on pooled benchmarks.
- **The decision to look is the hard problem**: The bottleneck is not retrieval but *knowing when to retrieve*. No current detector solves this on real messages.
- **Cost-efficient reconstruction is possible**: A flash-tier model matches frontier models on persona reconstruction at a fraction of the cost.

### For Ethics and Privacy

- The paper transparently reports that **anonymization does not remove authorship**: a stylometric classifier matches participants at 83.0% (no content words) and a language model at 98.0%, against 10% chance.
- The corpus necessarily contains GDPR Article 9 special-category data (health, etc.), processed under explicit consent basis Article 9(2)(a).

---

## Conclusion

RealCompanion releases ten real relationships between people and an AI companion, with every chat item carrying a checkable reasoning trace. The main takeaways:

1. **The past is rarely needed** (3.4% of messages) and, when needed, **far away** (median 2,157 messages back), with demand concentrating in the longest relationships.
2. **Supplying recorded evidence helps** (+0.455 content match on dependent items), yet **no system decides when to reach back** and **no detector tells the two situations apart** on real messages.
3. **Three agent systems reconstruct the persona with the same F1** at a 31-fold cost difference, and all add 70–80 further fields per 100 recovered.

The benchmark defines the task over any record of a person and states seven conditions a corpus must meet to measure it (Appendix A). Future directions include extending to more participants, longer spans, and evaluating whether the framing effects observed here persist in production systems.

---

_Markdown view of https://picx.dev/p/deh2Mg, served by PicX — AI-generated visual whiteboard summaries of research papers._
