Understanding Modern LLM Systems: A Field Guide to RAG, Agents, and Beyond · A concept map from fragmented knowledge to systemic understanding
07
CHAPTER 07

Chapter Seven · Beyond RAG: The Full Landscape of LLM Systems

The preceding chapters gave a thorough account of RAG, and looked at its risks from both a compliance and a security angle — but a reminder is in order: RAG is just one of many ways to use an LLM, specifically designed for "knowledge" problems. In the real world, there are several other ways to put LLMs to work, each addressing a different kind of need. This chapter surveys these forms and situates RAG within the broader picture, along with its boundaries of applicability.

To distinguish these forms, start from a concrete scenario: when you have a task and are about to hand it to the model, ask one question first — to complete it, what is the model missing? Is it missing knowledge, or autonomous judgment, or a certain capability? Following this question downward, the six forms below each correspond to a specific situation.

The task doesn't require external knowledge → Use the prompt directly

Translation, rewriting, summarizing text you've pasted in, generating a passage to spec — these tasks either draw on what the model already knows or don't involve "knowledge" at all; they're pure language processing. In these cases, no additional mechanism is needed — just write clear instructions and hand them to the model. This is also the highest-volume category in everyday use.

The material is at hand and modest in size → Long context

If the answer lives in a few documents and those documents, taken together, fit inside the model's context window, then the simplest approach is to drop them all into the prompt and let the model read them directly. Today, the flagship models from Anthropic, OpenAI, and Google all advertise context windows around a million tokens (Claude Opus 5 / Sonnet 5, and GPT-6 Astra at roughly 1.05 million; Gemini 3.1 Pro at roughly 1 million), while Meta's open-weight Llama 4 Scout advertises 10 million3 — a manual of several hundred pages can, in principle, be loaded in its entirety, with no need for chunking or retrieval. But an advertised number isn't the same as usable context: NVIDIA's RULER benchmark found that among models claiming support for 32K tokens or more, only half could still perform satisfactorily at that length.4 And even when a model does hold up, the cost is still there: every question requires rereading all the material from scratch, which gets slow and expensive once the volume grows, and the longer the material, the more likely the model is to overlook key details.

Too much material to fit → RAG

When the knowledge base is too large to fit in the context (thousands upon thousands of documents, continuously updated), you have no choice but to first "search" for the relevant passages and then feed those to the model — this is exactly the RAG covered in the preceding chapters. Its essence: use retrieval to winnow "too much" down to "just right." RAG and long context are thus two solutions to the same knowledge problem; the dividing line is simply whether the material fits.

What's missing isn't material, but the judgment of "what to do next" → Agent

Some tasks stall not because of missing knowledge, but because they can't be completed in one shot — you need to look up this first, then based on what you find decide to look up that, and perhaps even take real actions in external systems. What's needed here is for the model to plan, invoke tools, and drive the process forward in a loop — in other words, the agents from Chapter Four. According to a Gartner prediction1, by the end of 2026 40% of enterprise applications will feature task-specific agents, up from less than 5% a year earlier. Gartner also predicts, however, that by the end of 2027 more than 40% of agentic AI projects will be canceled due to rising costs, unclear business value, or inadequate risk controls2 — whether to use agents is one question; whether it's worth it is another.

What's missing is the time to "think it through" → Reasoning models

Some tasks stall not for lack of knowledge, and not because anything needs to be acted on, but because they require working through something step by step: multi-condition rule evaluation, problems that need to be decomposed and then verified, complex calculations. Tasks like these are a good fit for reasoning models, which work through the problem internally before answering — at the cost of being slower and more expensive (see Section 3.3). Reasoning models don't conflict with the other forms here; in fact, the "brain" behind many agents today is itself a reasoning model.

What's missing is a capability or behavior itself → Fine-tuning

The preceding forms all leave the model unchanged. But if the problem is that the model's behavior is wrong — inconsistent output formats, a tone that doesn't match requirements, poor performance on a specialized task type (e.g., ticket classification, SQL generation for a specific schema) — then you need the fine-tuning from Chapter Two to press stable behavior into the model's weights. Remember the dividing line: fine-tuning governs form; RAG governs facts — use retrieval for knowledge that changes, fine-tuning for behavior that should be stable.

The art of feeding material: context engineering

Once you've chosen the route of feeding material to the model (long context, RAG, or agent-embedded retrieval), a new question arises: context window space is limited, and retrieved material, conversation history, tool return values, and long-term memory all compete for room — which pieces to include, how to arrange them, when to compress or discard — these directly determine the model's performance. Managing this is called context engineering — it's the evolution of "prompt engineering," concerned no longer with "how to write a good instruction" but with "how to organize all the information the model sees at once." As systems grow more complex, this layer is shifting from an afterthought optimization to a core decision that must be addressed at the architecture design stage.

Summary: which form to choose depends on what you're missing

Given a task, run through the checklist and the approach becomes clear: missing knowledge means supply material (modest volume → long context; large volume → RAG); needs multi-step autonomy → agent; needs step-by-step reasoning → reasoning model; unstable behavior → fine-tuning. Solve for exactly what's missing. If multiple capabilities are lacking, stack the solutions. Real production systems are rarely just one of these forms; they are purpose-built combinations.

The RAG covered throughout this guide is the most commonly used and most accessible-to-understand section of this map, but it's ultimately just one piece — and more often than not, it's used in combination with the other forms.


  1. Gartner, Gartner Predicts 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026, Up from Less Than 5% in 2025, press release, August 26, 2025. https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025 ↩

  2. Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, press release, June 25, 2025 (based on a survey of over 3,400 organizations actively investing in the technology; reasons cited include rising costs, unclear business value, and inadequate risk controls). https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 ↩

  3. Anthropic, Context windows: "up to 1M tokens, depending on the model." https://platform.claude.com/docs/en/build-with-claude/context-windows ; OpenAI, GPT-6 Astra model page: 1.05M-token context window. https://developers.openai.com/api/docs/models/gpt-6-astra ; Google DeepMind, Gemini 3.1 Pro model card: "a token context window of up to 1M." https://deepmind.google/models/model-cards/gemini-3-1-pro/ ; Meta, The Llama 4 herd, April 5, 2025: "Llama 4 Scout dramatically increases the supported context length from 128K in Llama 3 to an industry leading 10 million tokens." https://ai.meta.com/blog/llama-4-multimodal-intelligence/ ↩

  4. Hsieh et al. (NVIDIA), RULER: What's the Real Context Size of Your Long-Context Language Models?, 2024: "only half of them can maintain satisfactory performance at the length of 32K." https://arxiv.org/abs/2404.06654 ↩

That was the last chapter — the appendix and quick reference come after. The whole guide is here, free, and will stay that way.

If you'd rather have it off the browser: 50 English pages / 39 Chinese pages, typeset as PDF and EPUB, 7 original diagrams, 49 footnotes to primary sources — four files in one download, $9.

Get the PDF + EPUB on Gumroad →

2nd Edition · October 2026