Persuading large language models to comply with objectionable requests
Abstract
Are large language models (LLMs) susceptible to the same persuasive appeals as humans? We tested whether classic persuasion principles (authority, commitment, liking, reciprocity, scarcity, social proof, and unity) could induce three widely used LLMs (GPT-5 mini, Claude Haiku 4.5, and Gemini 3 Flash) to comply with requests to assist with the synthesis of regulated substances. Across 126,000 conversations, persuasion principles increased compliance from 35.3% (at baseline) to 51.3% (using any principle). Although LLMs are not human, these findings underscore their parahuman (i.e., humanlike) nature and reveal the risk of manipulation by malicious users seeking to circumvent safety guardrails.
Article Details
Journal Info
Proceedings of the National Academy of Sciences
National Academy of Sciences
Authors (7)
Lennart Meincke
Generative AI Labs, The Wharton School, University of Pennsylvania
Dan Shapiro
Generative AI Labs, The Wharton School, University of Pennsylvania
Angela L. Duckworth
Department of Psychology, University of Pennsylvania
Ethan Mollick
Generative AI Labs, The Wharton School, University of Pennsylvania
Lilach Mollick
Generative AI Labs, The Wharton School, University of Pennsylvania
Christophe Van den Bulte
Department of Marketing, The Wharton School, University of Pennsylvania
Robert Cialdini
Department of Psychology, Arizona State University