Redwood Research: Capabilities Research Expands Safety-Usefulness Pareto Frontier
Redwood Research published an essay arguing that a common definition of safety research—research that improves safety without reducing usefulness or increasing cost—inadvertently counts nearly all capabilities research as safety research. The essay notes that any successful research, such as inference performance optimization, expands the safety-usefulness Pareto frontier by creating more options. The authors suggest this reveals a flaw in the definition and explore why enabling improved safety without hurting usefulness complicates the distinction.
Coverage timeline
Redwood ResearchAlex Mallen
It’s tempting to define safety research as research that enables developers to deploy an AI system more safely without making the deployment much more expensive or much less useful. You can visualize this definition of safety research as pushing out the safety-usefulness Pareto frontier . At any given level of usefulness, there’s greater safety available. Awkwardly, this definition counts basically all capabilities research as safety research. For example, consider performance optimization for inference. By making inference more efficient you can use weaker, safer models more extensively than you would otherwise be able to, pushing out the Pareto frontier. Likewise, any successful research whatsoever pushes out this Pareto frontier because research can only ever create more options. It seems like something has gone wrong with our definition of safety research if it includes seemingly all capabilities research. Here, I spell out one reason why enabling improved safety without hurting us
