The backfiring effect of weak AI safety regulation
Abstract
Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists—those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.
Article Details
Journal Info
Proceedings of the National Academy of Sciences
National Academy of Sciences
Authors (3)
Benjamin Laufer
Department of Information Science, Cornell Tech
Jon Kleinberg
Department of Computer Science and Information Science, Cornell University
Hoda Heidari
Machine Learning Department and the Institute for Software, Systems, and Society, Carnegie Mellon University