DiG-bench: 70-game benchmark measures AI discovery in hidden-rule worlds
Import AI 469 highlights DiG-bench, a new benchmark of 70 games designed to assess AI systems' ability to infer hidden rules and objectives through exploration and curiosity, akin to the ARC visual reasoning benchmark. The newsletter also covers science AI, an RSI simulator, and Mark Zuckerberg's technological pessimism.
Coverage timeline
Import AI (Jack Clark)Jack Clark
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now DiG-bench shows that Fable displays some creative intuition: …The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of their environment… How well can AI systems figure out the rules of their environment through exploration and curiosity, versus being fed them? That’s an important question for better understanding the intuitive and creative capabilities of AI systems and it’s one being asked by DiG-bench (Discovery in Games), a new benchmark of 70 games “designed to map the surface of discovery in well-controlled interactive systems”. Similar to the visual ‘ARC’ game, in DiG-bench “each game is a self-contained miniature world with its own laws, but both the rules and the objective are hidden from the player and must be uncovered through interaction”.
