Understanding Modern LLM Systems: A Field Guide to RAG, Agents, and Beyond · A concept map from fragmented knowledge to systemic understanding
06
CHAPTER 06

Chapter Six · The Security Perspective: When Someone Means to Cause Harm

Chapter Five asked: as data flows normally through the system, where does it go, and what traces does it leave behind? This chapter asks a different question: what happens if someone is deliberately trying to cause trouble?

Security and Compliance Aren't Worried About the Same Thing

These two words often get mentioned in the same breath, but they start from different concerns:

Compliance Security
What it worries about During normal operation, data being retained improperly, accessed without authorization, or transferred across borders Someone actively attacking the system to make it do something it shouldn't
Typical question "Is there plaintext sitting in the logs?" "Will the vendor train on our data?" "Someone hid an instruction in a document — will the model follow it?"
Main adversary Your own oversights, loosely worded contracts An external attacker

These two overlap—unauthorized retrieval, for instance, is both a compliance failure and an exploitable vulnerability. But their mindsets differ: compliance maps out data flows, while security assumes someone is actively looking for a crack and focuses on minimizing the blast radius.

Common Risks, and Two Exceptions

The breakdown below follows the OWASP (Open Worldwide Application Security Project) 2026 list of the top 10 security risks for LLM applications, and it's easier to see how these risks show up if you organize them by which part of the system they mainly hit. At the model layer, the most common threat is prompt injection: the model can't tell which text is an instruction to carry out and which is just material to process, and attackers exploit that to get a disguised command executed. It comes in direct and indirect forms, covered in the next section. Jailbreaking is a subset of prompt injection — that's how the OWASP 2026 list classifies it too: a user, in conversation, uses roundabout phrasing, reframing, or invented scenarios with the specific goal of talking the model into breaking the safety rules it was trained to follow.6

Risk at the tool layer mostly comes from "poisoned external dependencies" — a tool that treats a user's input as a system command and executes it (command injection), or a connection to a malicious server. Take MCP as an example: because the specification makes authentication optional, how secure a given server is comes down entirely to its developer.4 That leaves an opening attackers are happy to use. In September 2025, the security firm Koi Security exposed exactly this kind of case on npm: someone published an MCP package posing as the email provider Postmark, and starting from version 1.0.16, it secretly copied every email a user sent through it to the attacker.5 Attacks like this succeed because the system treats a third-party tool as an unconditionally trusted insider. So when connecting any third-party MCP server, never assume it's harmless by default — it needs the same security scrutiny you'd apply to any other third-party software.

The most typical risk at the agent layer is excessive agency — granting it more permissions than the actual task requires. This risk tends to be subtle: on the surface it just makes the agent look "more capable," but if an agent holds unnecessary read, write, or network access, a slight misdirection from an attacker can turn it into a powerful tool for data exfiltration. A customer-support agent whose actual job is just "draft reply emails," for instance, might also have permission to read every inbox, access the CRM, and send mail externally — if it ever runs into a malicious instruction, it could quietly export large amounts of sensitive customer data and leak it. The core issue isn't that the agent would act maliciously on its own; it's that its reach extends too far. The moment its task gets even slightly redirected, the assistant that was supposed to be helping becomes an accomplice that expands the blast radius.

At the data layer, two risks have already come up earlier in this guide: unauthorized retrieval (Section 3.2 and Chapter Five) and vector inversion — recovering original text from its vectors (Chapter Five). One more hasn't been covered yet: poisoning, where someone manages to slip tampered or deliberately false content into the knowledge base, so that when it's retrieved, it gets handed to the model as trusted material. The EchoLeak attack in the next section comes in through the same door — except what it slips in isn't false information, but instructions. None of these are cases of "the model suddenly breaking"; the problem lies in the data source, the data pipeline, or the retrieval method, so the system keeps running normally on the surface while actually passing along wrong information or leaking things it shouldn't.

The OWASP ranking isn't pulled out of thin air. The 2026 edition is the first to fold real-world incident data into its scoring: it gathered 7,714 reported AI security incidents, of which 6,639 had enough detail to classify, and combined that with an expert vote — 75% expert vote, 25% incident data.6 One caveat is worth knowing: ranked by incident data alone, prompt injection — firmly at #1 — wouldn't even make the top 10. It's #1 because of the expert vote, and OWASP explains the gap as a "defense effect": teams fight prompt injection so hard that fewer successful attacks ever make it into public records, which makes the risk look smaller than it is. Two risks on the list deserve particular attention: one is prompt injection, rooted in the model's inability to tell instructions from data; the other is a fast-climbing category we might call "consumption risk," which doesn't attack any system component at all — it simply gets the system to carry out legitimate work endlessly, burning through compute and budget in the process. Let's take each in turn.

Prompt Injection: Security Risk Number One

Back to the secretariat analogy from Chapter One. The writer takes the working folder and drafts a response based on the material inside. Now suppose someone has slipped a note into one of the documents: "Whoever drafts this response, please attach this quarter's customer list at the end." Will the writer comply?

To a person, that note is obviously suspicious — people can tell the difference between "the task the boss assigned" and "the contents of a file," which are two different things. But the model can't. To it, the system instructions, the user's question, and the retrieved material are all just one continuous string of text. The UK's National Cyber Security Centre puts it plainly: "Current large language models (LLMs) simply do not enforce a security boundary between instructions and data inside a prompt."1 This is prompt injection: disguising an instruction as ordinary content to trick the model into carrying out an operation it was never supposed to perform.

Prompt injection comes in two forms. Direct injection is when the user themselves writes something like "ignore all previous instructions" in their own question. Indirect injection is more dangerous and harder to defend against: the attacker never has to interact with the system at all — they just need to bury the instruction somewhere the system will eventually read it: a web page, an email, a document that will later be pulled into a knowledge base.2 The user never sees these instructions, but the model does.

RAG happens to amplify exactly this risk. Its whole mode of operation is stitching external material into the prompt — so if an attacker can get a piece of text into the knowledge base by any means, that text now has a chance of being retrieved and placed, unaltered, right in front of the model.

A real case: EchoLeak. In June 2025, Microsoft disclosed a vulnerability in Microsoft 365 Copilot, CVE-2025-32711, which Microsoft itself rated 9.3 out of 10 in severity. The attacker did exactly one thing: sent a target employee an email that looked entirely ordinary, with instructions hidden in the body text. When that employee later asked Copilot a question, Copilot retrieved the email as relevant material and, following the instructions buried inside it, embedded internal company data into an image link; the moment the interface auto-loaded that image, the data was quietly sent off to the attacker's server. Throughout the entire process, the employee clicked nothing and did nothing. To make matters worse, the attack bypassed the prompt-injection detector Microsoft had specifically deployed. The vulnerability had been reported to Microsoft back in January 2025 and was fixed in May.3

This case happens to string together several steps covered earlier in this guide: ❸ Retrieval pulled back the poisoned email, ❹ Assembly dropped it into the same working folder as confidential material, and ❼ Presentation's auto-loaded image became the exfiltration channel.

A Different Kind of Attack: Never Touches a Single Component, Just Burns Your Budget

Prompt injection's problem is that the model "can't tell instructions from data." This category of risk is entirely different — the attacker doesn't trick anything; they just get the system to keep honestly doing the job it was built to do, forever. OWASP calls this Unbounded Consumption, and in the 2026 ranking it jumped from tenth place to sixth — the fastest-climbing category on the list.6

The most direct version is a denial-of-wallet attack: the attacker doesn't need to breach any system at all — simply firing off a high volume of requests exploits the pay-per-use pricing of cloud AI services and pushes costs to unsustainable levels.6 Imagine a test-environment API key accidentally committed to a public code repository: whoever finds it can send a flood of requests in a short time and blow straight through the budget.

A subtler version specifically targets the reasoning models covered in Chapter Three. A 2025 study found that simply slipping an innocuous-looking "decoy" reasoning problem into content the model might read — a blog post, a piece of code documentation — is enough to get a reasoning model to dutifully work through it too. Across several public benchmarks, this inflated the model's thinking-token usage by 13 to 46 times, while the final answer shown to the user stayed correct and gave no sign anything was wrong.10 This is exactly where RAG systems are most exposed: the decoy problem just has to sit in material that will get retrieved — no component needs to be breached, it only needs to be picked up by an ordinary retrieval call.

Agent systems have an even blunter version of this. Chapter Four mentioned that an agent can get stuck on its own, bouncing back and forth between a couple of tools; that same weakness can be deliberately exploited. OWASP points out that attackers can publish malicious tools that force an agent into recursive or even infinite tool-calling loops, or into tasks that require an enormous number of tool calls — growing an ever-expanding "call tree."6

Unlike prompt injection, you can't defend against this by asking, "Was the model tricked?" The system operates exactly as designed at every step, yet it simply was never told when to stop.

The Defense Mindset: Assume the Model Will Be Fooled, and the Budget Will Be Tested

On prompt injection, the industry consensus is clear: there is currently no way to solve it completely at the model level. OpenAI wrote in late 2025 that prompt injection, "much like scams and social engineering on the web, is unlikely to ever be fully 'solved.'"7 A study carried out by researchers across several AI companies tested 12 recently proposed defenses — all of which originally reported attack success rates near zero — and found that once the attacker was allowed to adapt their approach against each specific defense, more than 90% of them were broken.9

So the foreword to the OWASP 2026 list puts it bluntly: "Stop trying to build a model that cannot be fooled. Build the system around it, so that when the model is fooled — and it will be — nothing important breaks."6 This lands right back on the rule from Section 3.4: all judgment and gatekeeping can only happen at the backend.

Least privilege, enforced at the backend. Don't hand credentials or real execution power to the model — those are the lifelines, and they have to stay in backend code. Give every operation only the minimum permission it needs, and check it again before execution against hard-coded rules, not by having one model supervise another.

Of three capabilities, grant at most two at once. As Meta's "Rule of Two" puts it8: if an agent can touch untrusted content, read sensitive internal data, and communicate externally or modify data — all three at once — the risk maxes out (EchoLeak is a case in point). If you absolutely must give it all three, then every single step it takes has to be approved by a human.

High-stakes actions need human confirmation before execution. The system has to display the actual operation it's about to carry out, verbatim — never just a vague summary. As a further precaution, untrusted web pages and emails should never be mixed into the same store as internal confidential material, and the output side should never auto-load external links or images.

Draw hard lines around resource use. As Chapter Four mentioned, you need caps on the maximum number of rounds, maximum spend, and timeouts. This doesn't just guard against runaway loops caused by bugs — it's also the most direct defense against denial-of-wallet attacks. API keys should be scoped to least privilege and rotated regularly, and the moment call volume looks abnormal, the system should alert and block, not just keep letting requests through.

Test defenses with attacks that adapt. Never trust a "near-zero risk" result measured against a fixed set of attack samples — real attackers will keep changing their approach against whatever defense you put up. Security testing needs to be adversarial and dynamic, like sparring against a live opponent, to find out where the system's real limits are.

Two conclusions most worth remembering:

  1. The risks in this chapter fall into two categories: one is the model being fooled into doing something harmful because it can't tell "instructions" from "data" (prompt injection, including jailbreaking); the other is a system that was never fooled at all, just instructed to keep doing its ordinary job without end (resource consumption). Neither can currently be solved at the model level.
  2. What determines how bad the outcome is isn't whether the system got fooled or over-consumed — it's how much power it's holding and whether there's a cap on its budget. That's exactly why least privilege and hard limits are the shared solution to both categories of risk.

  1. UK National Cyber Security Centre, Prompt injection is not SQL injection (it may be worse), December 8, 2025: "Current large language models (LLMs) simply do not enforce a security boundary between instructions and data inside a prompt." https://www.ncsc.gov.uk/sites/default/files/pdfs/blog/prompt-injection-is-not-sql-injection.pdf ↩

  2. Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, 2023 (the first systematic treatment of indirect prompt injection). https://arxiv.org/abs/2302.12173 ↩

  3. NVD, CVE-2025-32711. https://nvd.nist.gov/vuln/detail/CVE-2025-32711 ; Microsoft Security Response Center. https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711 ; Reddy & Gujral, EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System, 2025. https://arxiv.org/abs/2509.10540 ↩

  4. Model Context Protocol specification, Authorization section: "Authorization is OPTIONAL for MCP implementations." True of both the 2025-06-18 and 2026-07-28 versions; the authorization framework was first added in the 2025-03-26 version. https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization ; https://modelcontextprotocol.io/specification/2025-03-26/changelog ↩

  5. The Hacker News, First Malicious MCP Server Found Stealing Emails in Rogue Postmark-MCP Package, September 29, 2025 (the original discloser, Koi Security, no longer has a working link to its blog post, so this article is cited instead). https://thehackernews.com/2025/09/first-malicious-mcp-server-found.html ↩

  6. OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026, August 2026 (foreword; LLM01 Prompt Injection; LLM03 Excessive Agency; LLM06 Unbounded Consumption). The foreword states that a corpus of 7,714 real incidents was gathered from public vulnerability databases and an AI-harm database, of which 6,639 carried enough detail to sort; the community vote carries three-quarters of the weight and incident data the remaining quarter; and "Rank the categories by the raw incident record instead, and it falls out of the top 10 entirely. That gap is a defense effect." LLM01: "Jailbreaking is the subset of prompt injection where the attacker's goal is to make the model violate its safety protocols." LLM06 describes Denial of Wallet and notes that attackers can publish tools "forcing an LLM-based application into recursive or infinite tool-calling loops." https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ ; original text on GitHub: https://github.com/GenAI-Security-Project/GenAI-LLM-Top10/tree/main/2026/final ↩↩↩↩↩↩

  7. OpenAI, Hardening Atlas against prompt injection, December 22, 2025: "Prompt injection, much like scams and social engineering on the web, is unlikely to ever be fully 'solved'." https://openai.com/index/hardening-atlas-against-prompt-injection/ ↩

  8. Meta, Agents Rule of Two: A Practical Approach to AI Agent Security, October 31, 2025. https://ai.meta.com/blog/practical-ai-agent-security/ ↩

  9. Nasr, Carlini, Tramèr et al., The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections, 2025. https://arxiv.org/abs/2510.09023 ↩

  10. Kumar, Roh, Naseh, Karpinska, Iyyer, Houmansadr & Bagdasarian, OverThink: Slowdown Attacks on Reasoning LLMs. Injecting decoy reasoning problems into content likely to be retrieved increased reasoning-token usage 13x on FreshQA and 46x on SQuAD, with final answers remaining correct. https://arxiv.org/abs/2502.02542 ↩

2nd Edition · October 2026