When You Need an Agent, and When You Don't
The RAG described earlier is a fixed pipeline: retrieve material, assemble, generate — one pass through. For many questions, that's enough.
But some tasks can't be completed in a single pass. Take the AI customer support example again: a customer reports "after upgrading to the new version, a certain feature stopped working." To answer this, the system needs to check whether there are known issues with that version, then look up which configuration the customer is using, and possibly also check whether there are relevant change notes for that specific configuration — and what to look up in the next step depends on what was found in the previous one. If the first step already turns up the answer in the known issues list, the remaining steps don't need to happen at all.
This mode of operation — "break the task down on its own, decide the next step on its own, loop until done" — is what an agent is. The core distinction from a fixed pipeline is: how many steps to take, and what to do in each step, is decided by the model on the fly — not hard-coded in advance.
The flip side also needs to be stated clearly: simple tasks should not use agents. For questions that can be answered in a single step, calling the model directly or running a fixed RAG pipeline is better — faster, cheaper, and easier to debug when something goes wrong. The prevailing advice is to start with the simplest fixed pipeline and add complexity only when genuinely necessary. Autonomy has costs, which we'll cover at the end of this chapter.
ReAct: The Most Common Operating Pattern
The most fundamental and widely used agent pattern is called ReAct1, named after the paper's title combining Reasoning and Acting. In the original paper, the model's execution trace alternates between three parts:
- Thought: based on the information it has so far, the model judges what to do next
- Action: the model requests a tool call — querying a database, running a calculation, hitting an API endpoint
- Observation: the content returned after the tool executes, fed back to the model by the harness
Note that the first two are produced by the model; the third is fed in from outside. Later references often write these three steps as Reason / Act / Observe to align with the ReAct name — they mean the same thing.
It's worth noting that in the original paper, the Thought step was something the model wrote out in text. Today's reasoning models have largely internalized this step: they can "think" between tool calls and decide what to do next without a developer needing to prompt them into writing out an explicit Thought.2 So ReAct today functions more as a mental model for understanding how agents operate than as a template that must be followed literally.
The three steps cycle repeatedly until the model judges the task complete.
ReAct is the default starting point, but not the only pattern. When a task is complex enough that a single loop can't keep up, the model can be made to first decompose the entire task into a plan and then execute each item; when the same type of error recurs, a layer can be added for the model to review its failure reasons. The industry's experience is: start with ReAct, establish baselines for success rate, tool-call accuracy, latency, and cost; only upgrade once you've confirmed that a single agent genuinely can't solve the problem — premature complexity is a common and expensive mistake.
Four Supporting Concepts
Harness
The model itself is simply an inference engine that generates output given input context. It doesn't maintain state on its own, doesn't loop on its own, and can't directly execute external operations. What actually organizes the model into an agent is the surrounding runtime layer, typically called the harness. The harness drives the model through repeated cycles of thinking, calling tools, receiving results, and continuing to reason, enabling the system to carry tasks through to completion.
Beyond driving this loop, the harness is also responsible for: maintaining the roster of available tools, parsing the model's tool-call requests and actually executing them, deciding what goes into the context window each round (whether old content gets compressed or dropped), retrying on errors, intercepting high-risk actions for human approval, and logging the entire process.
In the language of Section 3.4, the harness is that "sole active initiator" — the backend. An agent doesn't change the communication diagram; it just has the backend run some of those steps repeatedly.
The term is borrowed from software testing (a "test harness" is scaffolding that runs code under controlled conditions). It deserves its own name because the quality of the harness matters as much as the quality of the model: swap the harness around the same model and task success rates can differ substantially; in practice, the bottleneck in a good number of enterprise agent projects turns out to be harness design, not the model's underlying capability. Every component in a harness encodes an assumption about "what the model can't do on its own" — and those assumptions are worth revisiting regularly, because they may have been wrong to begin with, and they tend to go stale quickly as the model improves.3
Function Calling (Tool Use)
The model does not actually execute anything. All it can do is output a structured piece of text meaning "I'd like to use the 'search product documentation' tool, with the argument 'legacy data migration'" — and then the harness carries out the execution and relays the result back.
In other words, the model can only express intent; it can't even initiate an action on its own — this is the flip side of the rule from Section 3.4: the initiative always stays with the backend. (This capability is now more commonly called "tool use" or "tool calling"; "function calling" is the earlier name.)
MCP (Model Context Protocol)
An open standard that specifies the interface tools and data sources should use to connect with an agent. It was originally released by Anthropic, and since December 2025 has been stewarded by the Agentic AI Foundation under the Linux Foundation.4 Before it existed, each new tool integration required writing bespoke adapter code; with a unified standard, it's like going from "every appliance with its own proprietary plug" to a USB port.
That said, using MCP comes with two costs that shouldn't be ignored. The first is a steep context cost: because MCP tool descriptions have to stay resident in context, connecting too many servers can eat up a large share of the context window before the conversation even starts (the industry currently mitigates this with "load tools on demand"). The second is a security cost: MCP lets an agent connect to any third-party tool it likes, which also means inheriting whatever vulnerabilities those tools carry — we'll look at this risk in detail in Chapter Six.
Skill
A pre-packaged workflow description: a document spelling out how a certain type of task should be done — what phrasing to use, which steps to follow, what counts as done — which the agent follows when needed. In practice it's just a folder containing an instruction file, optionally with templates and reference materials. This format was originally developed by Anthropic and released as an open standard in December 2025, and has since been adopted by a number of agent products.5
The key design is called progressive disclosure: normally only the skill's name and a one-sentence description are loaded; only when a task matches does the full content get read into context. So you can install dozens of skills without blowing up the context window.
Skills and MCP are often confused, but they're actually complementary: MCP solves "what it can reach" — connecting the agent to external systems; skills solve "how it should work" — teaching it a set of procedures. A skill doesn't connect to anything or execute any calls; it's simply a written set of instructions. Most production-grade agents need both.
Memory
The model has no memory — Chapter One said as much when explaining Loop C: multi-turn conversation works by stuffing the history back into the context. Agents face the same problem, and it's worse: a task that runs for dozens of rounds produces thinking, tool calls, and return results that quickly exceed what the context window can hold.
So agent memory comes in two tiers:
- Short-term memory: what's currently in the context window — the entire process of this task up to the present moment. It has a hard capacity limit, and the harness must continually decide what to keep, what to compress, and what to drop.
- Long-term memory: an external store outside the context that persists past experiences, conclusions, and user preferences, retrieving them when needed. Its implementation is essentially RAG — except the retrieval target isn't company documents, but the agent's own history.
Memory went from being essentially a synonym for "context window" around 2024 to being recognized as a core architectural layer on par with reasoning, orchestration, and tools — no longer optional.
Agentic RAG: Handing Retrieval Decisions to the Model
Agentic RAG is the third generation of RAG mentioned in Chapter One.
Its approach applies this chapter's mechanisms to the retrieval step: retrieval is no longer a fixed step three in a pipeline, but becomes a tool the agent can invoke on demand. The model judges for itself whether the material gathered so far is sufficient to support an answer; if not, it adjusts its angle and retrieves another round, repeating until satisfied — this is exactly Loop A from Chapter One.
To be clear: the components don't change, only the commander does. The vector database, embedding service, reranking service, and generation model all remain exactly the same, as do their communication patterns with the backend. The sole difference is that "how many retrieval rounds to run, and how to retrieve" shifts from preset, hard-coded rules to the model's real-time judgment during execution.
The cost of this is reduced predictability. Under fixed rules, how many steps a Q&A exchange will take, how many model calls it requires, and how much it will cost are all known in advance. Once the model decides on the fly, these depend on its judgment in the moment, and the model's output is inherently stochastic — the same question might take one retrieval round this time and four the next. Cost, latency, and even the answer itself go from being fixed values to being a range.
Costs and Risks
Context snowballs. With each loop iteration, the harness must resend all preceding thoughts, tool calls, and return results to the model — the input keeps growing; Section 3.3 explained that prefill cost grows quadratically with input length, so this part of the cost rises faster than the number of rounds.
This is exactly the scenario prompt caching is best suited for. Because each round of an agent loop typically only appends new content at the end of the conversation, the large block of background information earlier on stays unchanged, so the system can reuse results it's already computed. To get the most out of this, stable content — system instructions, reference material — should be placed together at the very beginning, with everything that changes pushed to the end. Caching typically works by matching a fixed prefix; even a small change in the middle, like reordering tools or inserting a timestamp, changes everything after it and invalidates the cache for that portion.
Prompt caching can cut usage costs substantially. Writing to the cache for the first time may cost 25% to 100% more with some providers, but cache-hit portions afterward typically drop to a tenth of the original input price or less.6 That said, the discount only applies to the cached input portion — the thinking and tool calls an agent generates each round are still billed at full output price. The cache also has an expiration window, and if the system automatically compresses the context mid-task, that can change the prefix too and invalidate the cache.
It may not stop. The model's ability to judge "the task is complete" is unreliable — it might bounce back and forth between two tools, or fall into an infinite loop. So the harness must enforce hard limits: a maximum number of rounds, a maximum spend, and what to do on timeout.
Tool failures cascade. If a tool returns an error or anomalous data, the model may continue reasoning on that basis and veer further and further off course. Production systems need to validate inputs before execution, retry on transient failures, and explicitly tell the model about failures rather than silently skipping them.
High-stakes actions need human approval. The most fundamental difference between an agent and a Q&A system is that an agent actually takes action — sending emails, modifying data, placing orders. For actions with irreversible consequences, the standard practice is to set up human-approval checkpoints and maintain complete operational audit trails. This should be treated as a first-class component at the architecture design stage, not patched in after an incident. Chapters Five and Six will revisit this from the compliance and security angles, respectively.
Agents amplify the damage of prompt injection. If a Q&A system is injected, the worst outcome is a wrong answer; if an agent is injected, it may actually carry out the attacker's desired operation. Chapter Six covers this risk in detail.
-
Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, ICLR 2023. https://arxiv.org/abs/2210.03629 ↩
-
Anthropic, Extended thinking (Interleaved thinking section): "Interleaved thinking lets Claude think between tool calls within a single assistant turn…" https://platform.claude.com/docs/en/build-with-claude/extended-thinking ↩
-
Prithvi Rajasekaran, Harness design for long-running application development, Anthropic Engineering, March 24, 2026. https://www.anthropic.com/engineering/harness-design-long-running-apps ↩
-
Model Context Protocol Blog, MCP joins the Agentic AI Foundation, December 9, 2025. https://blog.modelcontextprotocol.io/posts/2025-12-09-mcp-joins-agentic-ai-foundation/ ↩
-
Agent Skills official repository README: "The Agent Skills format was originally developed by Anthropic, released as an open standard…" (repository created 2025-12-16). https://github.com/agentskills/agentskills ↩
-
Anthropic, Prompt caching: cache-hit price is 0.1x the base input price (as low as 0.05x or 0.025x for some newer models), and cache-write price is 1.25x (5-minute cache) or 2x (1-hour cache). https://platform.claude.com/docs/en/build-with-claude/prompt-caching ; OpenAI, Prompt caching: cache-hit price is 0.1x, and cache-write price is 1.25x starting with GPT-5.6. https://developers.openai.com/api/docs/guides/prompt-caching ↩