Crowdsourced game on Olmo 3 reveals exploit patterns in prosocial AI tests
#ai-safety#crowdsourcing#olmo-3#evaluation
A crowdsourced game built on Olmo 3 demonstrated how players can exploit unexpected model behaviors to stress-test prosocial AI evaluations. The findings highlight that open access to a model's internals helps researchers understand why these evaluations fail.
Coverage timeline
Ai2 (Allen Institute for AI)
A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers understand why those tests break.
