Project page for Prompt Injection as Role Confusion, accepted to ICML 2026. We show prompt injections are driven by a flaw in how LLMs perceive roles. This lets us create new attacks, explain mech interp results, and predict when attacks succeed. We then discuss what roles are and why they matter, and share research ideas for a science of roles.
Make this part of your paper trail
Save this paper to a shelf, write a review, and keep your own notes.