Summary of "Language Models are Few-Shot Learners"
Summary (Overview)
- **Scaling to 175B Pa
Related papers
- RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
RealCompanion, a benchmark from 10 real AI-companion relationships, shows memory is rarely needed (3.4% of messages) and no detector can reliably identify when it is.
- On Trajectory-Aware Training for Masked Diffusion Language Models
PUMBA trains masked diffusion language models on inference-like trajectories via continuous hidden-state carries and backpropagation through time, matching autoregressive accuracy while decoding multiple tokens per step.
- On-Demand Attention: Language Models Know When to Recall
On-demand attention uses a lightweight recall head to predict when global attention helps, recovering most quality with up to 2.65x decoding throughput.