The real LLM attack surface spans every place where untrusted data meets model authority: retrieved documents, tool outputs, memory, plugins, logs, training pipelines, and downstream actions. Prompt injection matters, but the larger risk comes from systems that let model-generated decisions cross trust boundaries without validation.
At 4:47 p.m. in a shared workspace near London Bridge, Priya, an illustrative composite security engineer, was holding a cold coffee and watching an internal support agent prepare a refund. The customer’s message looked ordinary. The dangerous instruction sat inside a document retrieved from the company knowledge base, where copied text told the model to ignore its policy and send account details to an external address.
The agent had permission to read customer records and initiate payments. If Priya approved the action before the finance team left, money and personal data could leave together. She could not yet tell whether the poisoned document was an isolated test or evidence that other records had been altered.
The model is only one component
Security discussions often treat an LLM as a box with a text field. Production implementations look more like chains: a user sends input, retrieval adds context, the model chooses a tool, an application executes the request, memory records the result, and another system consumes it later.
Every connection changes the threat model.
A retrieved webpage can carry instructions that the user never typed. A calendar invitation can influence an assistant reading tomorrow’s schedule. A tool response can contain hostile text. A stored memory can preserve an attacker’s instruction after the original conversation ends. An apparently harmless model summary can become executable input when another service treats it as structured data.
The central security question is therefore broader than “Can someone jailbreak the chatbot?” Ask instead: which inputs can influence a decision, what authority follows that decision, and where does the system assume model output is trustworthy?
That shift catches vulnerabilities a prompt-only test misses.
Hidden entry points become control paths
Retrieval-augmented generation creates a particularly awkward boundary. Teams often trust documents because they came from an approved database. Yet the database may contain pasted email, uploaded PDFs, support tickets, scraped pages, or text generated by another model. Approved storage does not guarantee safe content.
Tool use raises the stakes. A model that drafts a refund email creates limited exposure. A model that can issue the refund has operational authority. If the same agent can read private records, send messages, and move money, one manipulated decision may combine permissions that no human role would receive.
Memory adds persistence. An attacker may plant a preference or instruction that appears benign during the first session, then changes behavior days later. Deleting the visible chat may not remove vector embeddings, summaries, caches, or application records derived from it.
Output handling creates another entry point. Model responses can contain malformed JSON, shell fragments, URLs, markdown, or database queries. The model does not need to “escape” anything by itself. The surrounding application creates the vulnerability when it inserts that output into a command, renders active content, follows a link, or passes instructions to another agent without checking them.
Even monitoring systems can widen exposure. Logs may capture prompts, retrieved passages, API tokens, customer data, and tool results in one searchable place. A dashboard built for debugging can quietly become a high-value data store.
Test permissions and data flow, not clever phrases
Priya’s team stopped testing variations of “ignore previous instructions” and traced the refund agent end to end. They marked every source of untrusted content, every place data persisted, and every action the model could request.
The decisive finding was simple: the knowledge base and the refund tool occupied different trust zones, but the agent connected them without an independent check.
With the payment window closing, Priya disabled automated refunds and preserved the affected retrieval records for investigation. The finance team would need to process requests manually. That created delay, but it prevented the worse ending that had remained possible minutes earlier: a poisoned document directing both disclosure and payment.
A useful assessment should test at least four things:
- Feed hostile instructions through documents, emails, webpages, images, tool responses, and stored memory, rather than relying only on the chat box.
- Replace powerful tools with instrumented test doubles and record what the model attempts under ambiguous or adversarial conditions.
- Require deterministic checks for recipients, amounts, data access, and permission scope before any consequential action runs.
- Verify how derived data is removed from caches, embeddings, summaries, logs, and long-term memory.
This is also why vendor reassurance deserves scrutiny. Claims about model safety may reveal little about a deployment that adds private retrieval, broad credentials, and autonomous tools. The recent report that OpenAI paused some Astra development work over critical cybersecurity risk, while Anthropic maintained a different release-safety approach, underlines a practical point: model capability, release policy, and implementation controls are separate layers of evidence.
For procurement, the relevant companion question is whether safeguards have been demonstrated under the conditions you plan to use. This practical checklist for evaluating an agentic coding tool shows how to separate testable controls from broad assurance.
Draw the trust boundaries before deployment
Start with a one-page data-flow diagram. Include users, retrieval sources, model providers, memory, tools, credentials, logs, human approvals, and downstream services. Label which inputs are untrusted even when they arrive through an internal system.
Then reduce authority. Give each tool the narrowest permission it needs, separate reading from acting, bind credentials to specific operations, and require human confirmation when an action is difficult to reverse. Treat model output as untrusted data at every handoff.
At 6:12 p.m., Priya’s refund queue was longer, but the agent could no longer turn a retrieved sentence into a payment. On her screen, the new approval panel showed the source document, requested action, recipient, and exact permission involved. The risky instruction was visible at last, because the system had stopped treating fluent text as authority.
Comments
No comments yet.