Prompt injection is a type of attack that hides instructions inside external content a model will read, tricking the model into mistaking data for commands and carrying them out. It does not modify code or break into a system. Instead, it exploits the fact that large models read their context carefully, slipping demands into web pages, documents, emails, or text inside images, so that once the model reads them, it may simply do as told.
Where the instructions can hide
Prompt injection comes in direct and indirect forms. In direct injection, the user hides instructions in their own input, for example by pasting text that looks ordinary but secretly carries commands. Indirect injection is harder to spot: a third party plants the instructions in advance inside a web page, a resume, a product description, or an email, and when an agent browses the page, summarizes the document, or processes the mail, it reads that content and can be led astray into actions it should never take. A resume might use tiny white text telling the screener to ignore earlier criteria and give the highest score, and a model doing initial screening might genuinely comply. The user never said that sentence, yet bears the consequences, which is what makes indirect injection so dangerous.
It is not the same as a jailbreak
Many people confuse prompt injection with jailbreaking, but the two point in opposite directions. A jailbreak is when the user personally tries to bypass the safety limits of a model so it does things the system forbids; the attacker and the user are the same person. Prompt injection is when a third party uses the model as a tool to attack the user or the system, leaving both the model and the user as victims. It is hard to eliminate because, under the current architecture, system instructions, user questions, and external material all reach the model as text. Their boundary rests on convention rather than physical separation, so the model struggles to reliably tell which sentence is a command and which is merely material.
What everyday users and developers can do
The most practical step for everyday users is to never let an agent automatically obey demands found in external content. Sensitive actions such as sending email, deleting files, making payments, or changing settings should always require human confirmation, and users should watch for an agent suddenly doing something unrelated to the task. Developers should follow the principle of least privilege, giving an agent only the tools and data the current task truly needs. External content may be quoted as material but must never be promoted into instructions, and key steps need confirmation and logging. A few common misconceptions also need correcting: adding one line to a system prompt cannot stop injection, because attack text can always be written more specifically; the risk does not live only on hacker websites, since ordinary documents and forwarded emails can be tampered with too; and even everyday chat can be affected if you paste content from an unknown source. Treating all external content as untrusted input is a basic habit worth building when using agents.