Prompt injection is an attack technique that exploits a structural limitation of language models: the inability to distinguish between instructions to execute and data to process.
For an LLM, everything entering the context carries equal weight: the system prompt written by the application developer, the user query, a corporate document retrieved via RAG, and the response from an external tool. Because no input possesses a higher privilege level than the others, any sentence phrased as a command can be treated as a legitimate instruction. From there, an attacker can bypass rules, access confidential information, or force the agent to perform unauthorized actions.
This is not a defect in a specific model, but rather the fundamental nature of how GenAI operates. For this reason, defenses must be designed at the architectural level rather than relying solely on writing safer prompts.
Prompt Injection as an Architectural Property of LLMs
A traditional web application separates code from data: a parameterized query neutralizes SQL injection because the database recognizes what constitutes an instruction versus user input. An LLM lacks this boundary: system prompts, user messages, document content retrieved from a RAG system, and tool outputs all reach the model as a single stream of tokens. While techniques like instruction hierarchy train the model to prioritize system instructions, this remains a statistical preference rather than a guaranteed security boundary.
This explains why prompt injection has occupied the top spot (LLM01) in the OWASP Top 10 for LLM Applications for two consecutive editions (2023 and 2025), and why no definitive patch exists.
It is not an implementation flaw tied to a specific vendor; it is the inherent cost of a natural language interface. Effective defenses do not eliminate the vulnerability, they contain it. They reduce the volume of hostile instructions reaching the model and, crucially, limit the damage those instructions can cause if they slip through. Designing under the assumption that some payload will inevitably bypass controls results in systems that are far more robust than attempting to build a perfect filter.
The practical implication for enterprise architecture design is that every untrusted input reaching an LLM must be treated as potentially executable code, not plain text.
Direct vs. Indirect Injection: Comparing Attack Vectors
To understand the attack surfaces to which LLMs are exposed, it is essential to distinguish between two main entry vectors:
- Direct Injection: The attacker is the user directly entering the payload into the chat interface.
- Indirect Injection: The hostile instruction is hidden within content retrieved automatically by the system (e.g., support tickets, emails, PDFs, web pages). This is the most dangerous form because it scales effortlessly and can leverage elements invisible to the human eye (such as white text on a white background, transparent Unicode characters, image alt text, or audio tracks).
The table below compares the primary attack vectors, highlighting their entry methods, typical goals, and detection complexity:
| Vector | Entry Method | Typical Goal | Detection Complexity |
| Direct Injection / Jailbreaking | User enters the payload in chat (“Ignore previous instructions…”) | Bypass safety policies, extract the system prompt | Medium (inspectable prior to call) |
| Indirect Injection | RAG documents, emails, web pages, tool outputs | Execute unauthorized actions via the agent | High (content appears “trusted”) |
| Data Exfiltration | Commands forcing external transmission of confidential data | Exfiltrate PII, secrets, knowledge base data | High (traffic appears legitimate) |
| Agentic / Tool Abuse | Injection that propagates to agent API and tool calls | Execute commands using the agent’s privileges | Very High (impact extends outside LLM boundary) |
While jailbreaking aims to make the model ignore content policies (generating inappropriate outputs), indirect injection on an agent connected to enterprise systems aims to cause an actual data breach.
Agentic Risk: When Text Becomes Action
Agentic risk emerges when a model goes beyond text generation to invoke tools and execute actions on behalf of the user, such as reading emails, querying databases, or updating CRM records.
Real-world vulnerabilities discovered by security researchers in production products demonstrate the scope of this challenge.
- EchoLeak (CVE-2025-32711): a single malicious email allowed automatic data exfiltration from Microsoft 365 Copilot. The payload bypassed filters and used Markdown links to transmit information externally.
- CurXecute (CVE-2025-54135): an injection in coding agents enabled arbitrary code execution on developer machines through manipulated MCP configuration files.
- ForcedLeak (Salesforce Agentforce): hidden prompts embedded in web-to-lead forms exposed and transmitted customer data to external domains.
The issue escalates with the adoption of the Model Context Protocol (MCP) and tool-calling systems, where every connected server or tool becomes a potential attack surface (tool poisoning).
To mitigate the impact of vulnerabilities in AI systems, adhering to Simon Willison’s concept of the ‘lethal trifecta’ is critical: an agent should never simultaneously combine access to sensitive data, exposure to untrusted content, and the ability to communicate externally or modify systems.
Data Exfiltration and PII Data Protection
In enterprise GenAI architectures, sensitive data leakage is a structural factor rather than an accidental one: every call to a third-party LLM transfers prompt context beyond corporate network boundaries. This affects data residency and compliance, requiring management through both contractual mechanisms (DPAs, zero data retention policies) and technical controls.
However, exfiltration via prompt injection poses a distinct threat: the attacker manipulates the model into embedding confidential information (from Knowledge Bases, chat history, or tool outputs) into a freely accessible egress channel.
Commonly exploited techniques include:
- Malicious Markdown Images: The model is instructed to render an image whose URL contains the exfiltrated data, which the client automatically requests upon rendering.
- Abuse of Network Tools: The system is tricked into invoking a web search tool while appending sensitive data into the search query string.
- Unauthorized Responses: The model directly outputs sensitive information that the user lacks authorization to view.
In all these scenarios, outgoing traffic appears legitimate unless explicitly inspected. Mitigating this risk requires three key countermeasures: sanitization or allowlisting of URLs and images in outputs, restrictive Content Security Policies (CSPs) on the client side, and egress allowlists for corporate tools.
Preventing Prompt Injection: Defense-in-Depth
Securing GenAI workflows in enterprise environments requires a paradigm shift. Security cannot rely on a single filter; it must adapt to a dynamic, multi-layered threat landscape.
To prevent malicious input from compromising enterprise systems, an effective defense combines inbound and outbound traffic inspection with an architecture designed to strictly contain the blast radius of a breach.
Core mitigation pillars include:
- Traffic Filtering and Anonymization: Block known patterns using deterministic rules, employ verifier models (LLM-as-a-Judge) for semantic threats, and pseudonymize PII (e.g., replacing tax IDs and IBANs with placeholders) across both prompts and responses.
- Instruction-Data Separation: Isolate external content from system instructions (e.g., via delimiter tags or Dual-LLM architectures) to prevent untrusted data from being interpreted as executable commands.
- Least Privilege and Action Isolation: Restrict agent blast radius (e.g., requiring human-in-the-loop approval before sending emails or calling critical APIs) and sanitize outputs before passing them to downstream systems.
Managing these pillars is more than a technical challenge; it quickly becomes a core governance priority with direct legal and regulatory ramification.
Governance and Compliance
In risk management committees, the fundamental question centers on accountability. If a compromised corporate agent exfiltrates personal data, it constitutes a data breach under the GDPR. This triggers a requirement to notify the supervisory authority within 72 hours of awareness, unless the risk to data subjects is unlikely.
The AI Act reinforces this framework: for high-risk systems starting in December 2027, the absence of input controls and logging will make it difficult to demonstrate compliance with requirements for robustness and automated event traceability.
This reality shifts focus from application-level controls to platform-level protection. Deploying fragmented guardrails across individual applications results in inconsistent coverage that is complex to audit and difficult to update. Conversely, centralizing defenses through an AI Gateway on the underlying infrastructure ensures uniform policy enforcement and a single point of maintenance.
This is the context in which the Radicalbit AI Gateway operates, designed to deliver a unified layer of control and protection:
- Centralized Controls: intercepts all traffic bound for LLM providers by enforcing bidirectional input and output guardrails—including PII de-identification (leveraging engines such as Microsoft Presidio for NER, regex, and tax/VAT codes), semantic evaluations, and access controls.
- Real-Time Audit and Traceability: exposes aggregated policy metrics, instantly delivering data required for compliance audits without needing to reconstruct logs from scattered microservices.
- Consumption Management: enforces rate limits, budget caps, and token thresholds per project directly at the network transit point.
- On-Premise Security and Data Residency: installs directly within the corporate infrastructure, keeping governance and control planes fully within the private boundary, even when leveraging external models.
No single solution completely eliminates prompt injection, and the Radicalbit AI Gateway is no exception. However, it shifts defense to a centralized location where controls are measurable, auditable, and updateable across the entire enterprise in a single operation.
The solution is also available as open source on GitHub, and the Radicalbit team is available for dedicated evaluations in enterprise settings.
Frequently Asked Questions About Prompt Injection
Jailbreaking is a form of direct injection that aims to bypass the model’s content safety policies. In agentic applications, the focus shifts: prompt injection aims to override the application logic to force the agent into performing unintended actions. The former is a model alignment issue, whereas the latter is an application integrity vulnerability.
Because it requires no access to the user interface and leverages content the system inherently trusts by design: indexed documents, emails, web pages, and tool outputs. The attacker drops the payload and waits for automated processes to ingest it, often with zero human interaction.
No: le istruzioni difensive contenute nel system prompt risiedono nello stesso flusso di dati del payload ostile e possono essere aggirate con riformulazioni, offuscamento o traduzioni.
By combining deterministic filters for known signature patterns with semantic evaluations using verifier models. Monitoring guardrail metrics allows teams to detect anomalies or reconnaissance attempts. In initial stages, running in observation-only mode is recommended to fine-tune thresholds and prevent false positives.
An AI Gateway serves as a centralized proxy for all GenAI traffic. As a mandatory transit point, it consistently applies guardrails, PII anonymization, and logging across all enterprise applications—eliminating the need to duplicate security logic across individual codebases.
Key Takeaways
- Prompt injection is an architectural limitation of LLMs that cannot be resolved with a patch; it requires defenses designed at the system level.
- Indirect injection is the most severe threat because it exploits trusted data sources without requiring direct user action.
- In agentic systems, risk becomes operational, transforming manipulated input into unauthorized execution of actions and API calls.
- Security requires a multi-layered strategy combining deterministic filters, PII anonymization, semantic controls, and least-privilege agent permissions.
- Deploying an AI Gateway solution centralizes enterprise governance, allowing defense policies to be managed and updated from a single control point.
