Modern large language models possess impressive capabilities but remain vulnerable to various attacks capable of manipulating their responses, causing confidential data leaks, or bypassing restrictions. The main focus is on analyzing “prompt injection” attacks, which allow circumventing model limitations, extracting hidden data, or forcing the model to follow malicious instructions.
Bezzateev et al. (Mon,) studied this question.