Understanding Modern LLM Systems: A Field Guide to RAG, Agents, and Beyond · A concept map from fragmented knowledge to systemic understanding
03
CHAPTER 03

Chapter Three · Deep Dive: Four Core Technical Mechanisms

3.1 Embedding: How Machines "Understand" the Meaning of Language

First, why it's needed: at bottom, computers can only do math with numbers — they don't understand the relationship between the words "cat" and "dog." But if we can turn every word and every sentence into a string of numbers (say, a "coordinate" made up of 1,024 numbers), and arrange things so that "phrases with similar meaning" end up "close together" in the space these numbers represent while "phrases with unrelated meaning" end up far apart, then the computer can indirectly judge whether "these two sentences are about the same thing" simply by calculating the distance between two coordinates. This string of numbers is called an "embedding vector," and the conversion process is called "embedding."

An analogy: imagine a vast library where the librarian doesn't shelve books alphabetically, but instead places books on similar topics on neighboring shelves — a book about cats and a book about dogs both belong to the "pets" category, so they'd be placed close together; a book about cats and a book about auto repair would be placed far apart. To find "books similar to this one," you just look at what's sitting next to it. Embedding does exactly this, except instead of placing books on three-dimensional shelves, it places them in an abstract space with hundreds or thousands of dimensions (the human brain can't directly visualize what "1,024-dimensional space" looks like, but mathematically it's perfectly computable).

How this string of numbers is "learned": nobody manually decrees that "cat = [0.2, 0.8, ...]." Instead, it's done through a training method called contrastive learning — the model is shown vast quantities of "these two sentences are related" and "these two sentences are unrelated" examples, and every time it gets one wrong, its parameters are nudged slightly, pulling "related" pairs closer in the space and pushing "unrelated" pairs further apart. After repeating this millions of times, the model organically "figures out" a coordinate system in which "close together = similar in meaning" holds true. No single dimension is deliberately designed to represent a specific meaning. Instead, meaning emerges from the combined "position" of all dimensions—much like a person's character isn't defined by a single trait, but by an overall impression.

The technical process in concrete terms (two steps):

  1. Tokenization: first, the text is cut into small units the model can process, called "tokens." Note that a token is neither a "word" nor a "character" — it's a segment divided by frequency of occurrence: common algorithms (BPE, SentencePiece) start by splitting text into its smallest units, then repeatedly merge the pair that "appears together most frequently," until a vocabulary of the desired size is reached (today's mainstream models mostly fall between 100,000 and 262,000 entries: Llama 3 at roughly 128K, Qwen3 at roughly 152K, GPT-4o at roughly 200K, Gemma 3 at roughly 262K2). So high-frequency words may be a single token, while rare words get split into several subword pieces.

This design has a practical benefit: no matter how rare or newly coined a word is, it can always be decomposed into known subwords or even individual bytes — there's never a situation where "this word is unrecognized and can't be processed." That was a common failure point with earlier tokenization methods.

There's also something practically relevant to cost here: the same passage can consume very different numbers of tokens depending on the language and the model. Since both pricing and "how much the model can process at once" are measured in tokens rather than characters, and early tokenizers weren't especially kind to Chinese, the same meaning often used to cost noticeably more tokens in Chinese than in English. Newer tokenizers have narrowed this gap substantially: when OpenAI announced GPT-4o, it reported that the same sample passage dropped from 34 tokens in Chinese to 24 — on par with English.3 So exactly how much more a given language costs depends on which model's tokenizer you're using.

  1. Encoding (Transformer neural network): the sequence of tokens is fed into a neural network called a "transformer," which computes a vector for each token, then combines them via "pooling" (merging many token vectors into a single vector representing the whole sentence) and "normalization" (scaling the vector to unit length, so similarity can be compared using a standardized method). The result is a single vector that represents the meaning of the entire sentence. The number of values this vector contains (i.e., its "dimensionality") typically falls between 384 and 4,096 as of 2026, with general-purpose scenarios usually starting at 768 or 1,024, then adjusting based on retrieval quality and storage cost.

A note: there's more than one way to build an embedding model, and the approaches have kept evolving. The earliest traditional route learned word vectors first and then averaged or weighted them into a sentence vector — word2vec (2013) and GloVe (2014) are the classic examples; FastText (2016) added subword information on top. Around 2019, Siamese networks trained specifically for sentence similarity, such as Sentence-BERT, took off, built on BERT-style encoders. Since late 2023, a new wave of embedding models built on top of large language models — adding a pooling layer and contrastive fine-tuning (e.g., E5-mistral, GritLM, NV-Embed) — has taken turns at the top of the public MTEB leaderboard.4 For a user, though, the conclusion is the same either way: no matter which method was used to train it under the hood, what you get is a sentence vector you can plug straight into a distance calculation — you don't need to worry about how it was produced.

One limitation you must know about: embedding models have a maximum input length per call. Every embedding model specifies a maximum number of tokens for a single input; anything beyond that limit is silently truncated — with no error or warning. A chunk that's too large may have its second half simply not represented in the vector, making it permanently invisible to retrieval. This is one of the hard reasons why the "offline indexing" step in Chapter One must chunk long documents first: chunk sizes must fall within the length limit of the embedding model being used.

A rule that's easy to trip on but critically important: when you convert documents from your knowledge base into vectors ahead of time (this is called "indexing"), and when you convert the user's question into a vector at query time (this is called "querying"), you must use the same embedding model for both. The reason is simple: different models "learn" different coordinate systems, just like two maps that use different scales and different origin points — you can't take coordinates from Map A and look them up on Map B. So if you want to switch to a different embedding model, you can't just swap out the query side — you have to recompute every vector in the entire database (called re-embedding), and the dimensionality of the stored vectors and the query vectors must match.

The "dimensionality must match" above means the vectors stored in the database and the vectors computed at query time must be the same length — otherwise they simply can't be compared. But this doesn't mean the dimensionality itself is permanently fixed. Current mainstream models widely support a technique called Matryoshka (nesting-doll) embeddings: a long vector computed by the same model can simply have its trailing dimensions chopped off without retraining — for example, keeping only the first 256 out of 1,024 dimensions typically degrades retrieval quality by only a few percentage points, while dramatically reducing storage and search costs. If you want to save costs this way, that's perfectly fine — just make sure both the stored vectors and the query vectors are truncated to the same length.

3.2 Retrieval: How to Find Everything That Should Be Found

Chapter One mentioned that retrieval in a real system is a "wide-to-narrow funnel." This section explains why the funnel is designed this way — what deficiency each layer compensates for in the layer above it.

Start with the base capability: what vector search can do

Section 3.1 explained that text with similar meaning produces vectors that are close together. So the most basic form of "retrieval" is: convert the question into a vector too, then find the entries in the database whose positions are closest. As for how "closest" is concretely calculated, in practice it usually yields a score between 0 and 1, where closer to 1 means more closely matched in meaning — the "score 0.82" you see in the system is this number. Under the hood it's computing the angle between two vectors: the more aligned the directions, the higher the score. No need to dig into the details.

The greatest strength of this approach is that it doesn't depend on literal wording. A user asks "how do I move old data to the new version"; the documentation says "historical record migration procedure." The two sentences share not a single word in common, yet vector search connects them anyway. This is its fundamental advantage over traditional keyword search.

First-layer deficiency: at scale, comparing every single entry is unacceptably slow

If the database holds ten million vectors, comparing against every one of them and sorting the results for each question is something real-time Q&A simply can't sustain. The solution is ANN (Approximate Nearest Neighbor): rather than insisting on finding the absolute most similar entries with 100% accuracy, ANN uses pre-built index structures to compare against only a small fraction, trading a tiny loss in precision for a speedup of several hundred times.

This "precision" has a standard measure, called recall@k (top-k recall): out of the truly most similar top k entries, how many were actually found and returned. It's an adjustable knob, not a fixed setting — and the cost curve gets steep the higher you push it: pushing recall closer to 100% requires comparing against more and more candidates, and latency grows much faster than recall itself does. So most production systems don't chase near-perfect recall; instead they look for a balance point between "accurate enough" and "fast enough."

The specific index structures come in three main types, and which one to use depends on data scale:

  • HNSW (Hierarchical Navigable Small World graph): pre-connects all vectors into a "neighbor relationship graph"; at query time, it works like playing "six degrees of separation" — starting from one point, at each hop jumping to a neighbor closer to the target, reaching the target region in just a few steps. The default choice for most production systems in 2026 — high recall, supports adding new data at any time. Its one hard constraint is that the entire graph must fit in memory; once the data volume exceeds what a single machine's memory can hold, you need to switch to something else.
  • IVF (Inverted File Index): uses k-means to roughly cluster all vectors into groups, then at query time only searches the most relevant few clusters in detail. Lower memory usage than HNSW, but also lower recall under equivalent conditions — suited for large datasets with mostly static content and tight memory budgets.
  • DiskANN: keeps the full index on SSD, with only compressed vectors and navigation structures in memory. The standard answer at the billion-vector scale: on a workstation with just 64GB of memory, the paper achieved over 95% recall on a billion vectors with average latency under 3 milliseconds — and the same machine can hold 5 to 10 times as much data as a pure in-memory graph index like HNSW.5

Implementations of these methods can be found in FAISS (an open-source vector search library) and various commercial vector databases.

Second-layer deficiency: vector search misses things that require exact literal matches

Vectors' strength is understanding semantics, and that's also their weakness — they're insensitive to literal text. A user asks "does model X-200 support this feature?"; vector search is likely to return a bunch of passages about product specifications in general, yet miss the specific line that actually mentions X-200. Semantic similarity is virtually useless for product models, error codes, or version numbers — situations where a single character changes the entire meaning."

So production systems typically run two retrieval tracks in parallel: vector search for "meaning matches" and keyword search (the industry-standard algorithm is called BM25) for "literal matches," then merge the two result sets. This is hybrid search, and what Chapter One called "searching through two channels at once." When merging, because the two scoring systems aren't on the same scale, a conversion rule is needed. The common approach is called RRF (Reciprocal Rank Fusion) — rather than comparing scores directly, it only looks at what rank each result occupies in its respective list, and entries that rank highly in both get a higher combined score.

Third-layer deficiency: some of what's retrieved, the user shouldn't be allowed to see

Feeding retrieval results directly into the prompt is dangerous: different users have access to different documents; a customer shouldn't see internal-version documentation, and an external partner shouldn't see internal pricing details. So at the time of chunking and indexing, every chunk must carry metadata about which document it came from and who is allowed to see that document (i.e., ACL — Access Control List), and retrieval must filter in real time based on the querying user's identity. The same applies to timeliness filtering — superseded old-version documents shouldn't be dug up and used as evidence.

This step is more engineering-intensive than it sounds: filter first, then search breaks the ANN index structure (the neighbor-relationship graph was built on the full dataset — remove some of the points and the paths break); search first, then filter risks retrieving 50 entries but being left with only 2, or even zero, after filtering. Mature vector databases provide specialized filtered-search mechanisms for this, but it remains one of the most frequent sources of performance issues in real deployments.

This ordering also has a security dimension. The OWASP 2026 list specifically warns that if similarity search runs over the full dataset first and permission filtering happens afterward, an attacker — even one who can't see a single one of someone else's documents — can infer that those documents exist, and roughly what they're about, just from the number of results returned, the distribution of scores, and how fast the response comes back. In a product where multiple customers share one vector store, that's a cross-tenant leak. So it recommends writing the permission scope directly into the retrieval query itself, and in sensitive scenarios, building separate indexes per customer altogether.6 Performance and security point to the same conclusion here: permission filtering belongs inside retrieval, not after it.

Fourth layer: the dozens of rough-screened results still need further selection

The goal of the preceding layers is "don't miss anything" — better to bring back too many than too few. However, you can't simply stuff all these results into the prompt. Context windows have strict length limits, and irrelevant filler distracts the model. So a final round of reranking is needed: a more precise but slower model takes the few dozen rough-screened results, compares each one against the question individually, rescores and reorders them, and keeps only the most on-point handful.

A reranking model works differently from an embedding model: the embedding model converts the question and the document separately into vectors and then compares distances — fast but coarse; the reranking model reads the question and the document together in one pass and then scores — slow but much more accurate. Precisely because it's slow, it can only be used after the candidate set has been narrowed to a few dozen — which also explains why the entire funnel has to be "wide first, then narrow": use fast-and-coarse methods to shrink the scope from millions to dozens, then use slow-and-precise methods to pick a handful out of those dozens.

3.3 LLM Serving: How the Model Handles Many Users at Once

The previous two sections covered "how the material gets found." This section covers what happens after the material is handed to the model. Note the keyword here is serving — the focus isn't only on how the model computes internally, but on how a server organizes things when it needs to serve hundreds or thousands of users simultaneously.

A single inference happens in two phases

The way a model generates text is called autoregressive generation: it produces one token at a time, appends it to the existing text, then predicts the next token based on "everything so far," repeating until it produces a stop token.

This creates two phases with opposite computational characteristics:

  • Prefill: reads the entire input (system instructions + retrieved material + user question) all at once in parallel, computing two things (K and V, explained below) for every token, and storing them in a GPU memory buffer called the KV cache. This phase is compute-bound.
  • Decode: starting from the first output token, generates one at a time, sequentially. Each new token written requires looking back at all preceding content. This phase is memory-bandwidth-bound — the bottleneck isn't how fast you can compute, but how fast you can read the KV cache into the processor.

Why the KV cache must exist: decode requires looking back at all preceding text for every new token. If every token required recomputing all the preceding tokens from scratch, then by the 100th token you'd be recomputing the first 99 — this redundant computation grows quadratically with text length, making long outputs completely impractical. Storing and reusing previously computed results brings the computation down to near-linear. Trading GPU memory for eliminated redundant computation — that's the entire point of the KV cache.

So what are K and V? When the model processes each token, it computes three vectors: Q (the question this token carries: "which preceding tokens are relevant to me?"), K (the sign each token holds up, reading "here's what I'm about"), and V (the actual information this token carries). The current token takes its Q and compares it against the K of every preceding token; whichever matches most strongly gets a higher weight; then the V values from all tokens are combined via weighted average — this is "attention." The reason only K and V are cached, not Q, is that K and V get reused by every subsequent new token — worth storing; Q is only used once in the current step, then discarded.

Two metrics for measuring speed

These two phases each correspond to a standard industry metric; understanding them lets you follow most discussions about inference performance:

  • TTFT (Time To First Token): the time from sending the request to seeing the first output token, determined by prefill.
  • TPOT / ITL (Time Per Output Token / Inter-Token Latency): after output begins, the rhythm between successive tokens, determined by decode.

It's worth noting that total wait time is usually dominated by decode. As an illustration, suppose a regular (non-reasoning) model outputs around 60 tokens per second — about 17 milliseconds per token. A roughly 500-token response would then take about 8 seconds for decode alone, while the time to first token might be only a few tenths of a second. So "the longer the answer, the slower it feels" is a linear accumulation — the user experience is very direct.

Reasoning models have made decode an even bigger share of the cost. Starting around 2024, "reasoning models" emerged that generate a large block of internal thinking before producing an answer. That thinking is decoded token by token too: it's usually not shown to the user, but it still occupies the context window and is billed as output tokens just the same.7 From the user's side, this shows up as a longer time to first token — before the first visible character appears, the model may already have quietly written a large amount of thinking. When OpenAI released its first reasoning model, o1, it noted that performance kept improving the longer the model was allowed to think8 — effectively trading inference-time compute and wait time for answer quality.

How a server handles many users at once: continuous batching

If a server processes one request at a time, finishing it before accepting the next, the GPU sits idle most of the time — because during decode, compute capacity goes underutilized, resulting in severe waste.

Modern inference frameworks (vLLM, SGLang, etc.) use an approach called continuous batching: in each iteration, all currently active requests each advance by one token, then the next iteration begins; when a request finishes it exits the batch, and when a new request arrives it joins immediately — no need to wait for the entire batch to complete. This is the key to keeping the GPU fully utilized, and the reason it can serve dozens to hundreds of users simultaneously.

Early continuous batching also produced a phenomenon everyone has run into: the response stream suddenly stutters mid-sentence. When a new request's prefill squeezes into the batch, it demands a large burst of compute, forcing requests that are currently streaming output to wait — what the user sees is the text suddenly freezing. Mainstream frameworks now handle this directly: vLLM enables "chunked prefill" by default, splitting a long prefill into small pieces and prioritizing requests that are already streaming output, which the official docs say improves inter-token latency.9 A further step is to split prefill and decode across separate machines entirely (prefill-decode disaggregation).

A few practical cost rules

  • When inputs get very long, prefill costs rise faster than you'd expect. Only the attention portion of the computation scales quadratically with input length; everything else is linear. Kaplan et al. estimate that as long as the input length stays under about 12 times the model's width (an internal sizing parameter), the length-dependent portion of the computation remains a relatively small share of the total.10 For large models, that threshold sits somewhere above tens of thousands of tokens: at a few thousand tokens, doubling the material roughly doubles the cost; once you're past the threshold, the quadratic term starts to dominate and growth becomes noticeably faster than linear. This is the hidden price of "stuffing more material into the prompt."
  • Longer answers increase decode time linearly, and decode usually dominates the user's total wait time.
  • The KV cache is a significant consumer of GPU memory, and it grows as the context gets longer — it's the main constraint on how many users a single machine can serve concurrently. To address this, the industry developed optimizations like PagedAttention1: managing KV cache the way an operating system manages memory pages, preventing large blocks of GPU memory from sitting idle. The open-source inference framework vLLM uses exactly this approach.

An important optimization for prefill: prompt caching. In a RAG system, the beginning of every request is often identical — the same system instructions, sometimes the same long document. Since the content is the same, the computed KV cache is also the same, so there's no need to recompute it every time. Cloud providers therefore widely offer cross-request prefix caching: the KV cache for this shared prefix is retained for a period — anywhere from a few minutes up to 24 hours, depending on the provider and configuration11 — and subsequent requests that hit the same prefix reuse it directly, saving precisely that quadratic prefill cost.

This mechanism has an easily overlooked implication: the KV cache is not, as commonly assumed, "computed and discarded, existing only within the current request." With caching enabled, a portion of the content resides briefly in the provider's infrastructure — Chapter Five will return to this point when discussing compliance.

3.4 Who Calls Whom: How the System's Components Divide the Work in a Single Query

The core rule: across the entire system, only the "backend / orchestration layer" actively initiates actions. The vector database, the model services — these components are all passive: they respond only when asked, and they never communicate directly with each other.

Using the secretariat analogy from Chapter One, the backend is the secretary who owns this task. Which archive room to pull files from, what to do if they can't be found, whether the material is sufficient, whether to go back for another round, whether the draft is fit for submission — every one of these decisions is made by the secretary, and if something goes wrong, the secretary bears responsibility. The archive room only hands over files when asked; the writer only drafts a response when material arrives. The two never interact directly and may not even know the other exists. So the backend is not a messenger; it's the sole owner of this task.

The complete communication sequence:

  1. Frontend → Backend (corresponds to Chapter One ❶ Receive the question): the user's question is sent to the backend.
  2. Backend → Embedding service (corresponds to ❷ Embed & understand): sends the question text over, requesting it be converted into a query vector.
  3. Backend → Vector database (corresponds to the first half of ❸ Retrieve): the backend uses the query vector to find the closest matching information. At the exact same time, it checks the user's permissions (ACL — Access Control List) to make sure it only pulls documents the user is allowed to see. It then returns the most relevant results (known as the "top-k").

The keyword search mentioned in Section 3.2 usually happens right here too. Because most modern vector databases have keyword search built in, they can run both searches and merge the results in a single trip. The only exception is if your keyword index lives on a completely separate system (like a standalone Elasticsearch setup) — in that case, the backend makes two separate requests and merges the results itself. 4. Backend → Reranking service (corresponds to the second half of ❸ Retrieve): sends the few dozen rough-screened results along with the question, requesting the service to score and reorder each one, keeping only the most relevant handful. The reranking model is the same type of component as the embedding model — both are standalone model services running on GPUs, which can be self-hosted or called via a ready-made API. 5. Backend (completed internally, no external call) (corresponds to ❹ Augment): assembles the system instructions, retrieved original text, conversation history, and user question into the complete prompt. 6. Backend → Generation model service (corresponds to ❺ Generate): sends the complete prompt over, requesting it to generate an answer. 7. Backend → Frontend (corresponds to ❻ Guardrails and ❼ Respond with citations): receives the model's output stream while checking it; the portions that pass are pushed to the interface along with their source citations. 8. Backend → Logging system (corresponds to ❽ Log): writes a record of this entire Q&A exchange for the permanent record.

Why the two "eight-step" breakdowns don't align neatly: Chapter One divided by "what action is performed"; this section divides by "who sends a message to whom." The two framings were never meant to map one-to-one. ❸ Retrieve is a single action but requires two to three separate communications; ❻ Guardrails and ❼ Respond with citations are two actions but share a single return transmission. Same system, different angle, naturally different slicing.

About step 7: the answer is "streamed" over, not "sent" after it's fully written

Section 3.3 explained that the model generates token by token. The backend doesn't wait for the entire answer to be complete before sending it back — it pushes tokens to the frontend as they arrive, which is why you see text appearing one character at a time (technically this usually uses SSE or WebSocket).

This creates a real engineering conflict: streaming output and guardrails are at odds. If tokens have already been pushed to the user's screen, how can you still "check before releasing"? The practical compromise is to detect problems in-flight while streaming, and the moment something is flagged, immediately interrupt and retract or replace what's already been displayed. This is why you sometimes see an AI's response vanish mid-sentence, replaced by "I'm sorry, I'm unable to answer that question."

Four points worth remembering:

  1. The downstream components are mutually unaware of each other's existence. The vector database doesn't know there's a model service; the reranking service doesn't know there's a logging system — it's the backend that takes the output of one component, translates it into the format the next component expects, and passes it along. The more components there are, the more pronounced this becomes: in the eight steps above, the backend communicates with six different parties, while the number of direct connections among those parties is zero.

  2. All judgment and gatekeeping can only happen at the backend. It's the only component with full visibility, and it sits within the company's own sphere of control (the model service may well belong to another company entirely). This is why compliance and security review logic must live in this layer — Chapters Five and Six will cover this in detail.

  3. The backend is also responsible for all "what if something goes wrong" handling. If the model service is rate-limited, should it queue and retry? If retrieval times out, should it degrade gracefully or return an error outright? If a step fails, where should it fall back to? Only the backend can make these calls. In real systems, the code for handling exceptions is often more extensive than the code for the happy path — this is precisely the difference between an "owner" and a "messenger."

  4. Agentic RAG doesn't change this diagram (covered in detail in Chapter Four). It simply has the backend run steps 2 through 4 for additional rounds, with the number of rounds decided by the model on the fly. The division of labor, who can talk to whom — none of that changes.



  1. Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023 (the core paper behind vLLM). https://arxiv.org/abs/2309.06180 ↩

  2. Meta, Introducing Meta Llama 3, 2024: "a tokenizer with a vocabulary of 128K tokens." https://ai.meta.com/blog/meta-llama-3/ ; Qwen Team, Qwen3 Technical Report, 2025: "vocabulary size of 151,669." https://arxiv.org/abs/2505.09388 ; OpenAI's tiktoken source code, o200k_base encoding (used by GPT-4o), end-of-text token id 199999. https://github.com/openai/tiktoken ; Gemma Team, Gemma 3 Technical Report, 2025: "262k entries." https://arxiv.org/abs/2503.19786 ↩

  3. OpenAI, Hello GPT-4o, May 2024 (Language tokenization section). https://openai.com/index/hello-gpt-4o/ ↩

  4. Mikolov et al., Efficient Estimation of Word Representations in Vector Space, 2013 (word2vec). https://arxiv.org/abs/1301.3781 ; Pennington et al., GloVe: Global Vectors for Word Representation, EMNLP 2014. https://aclanthology.org/D14-1162/ ; Bojanowski et al., Enriching Word Vectors with Subword Information, 2016 (FastText). https://arxiv.org/abs/1607.04606 ; Reimers & Gurevych, Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, EMNLP 2019. https://arxiv.org/abs/1908.10084 ; Wang et al., Improving Text Embeddings with Large Language Models, 2023 (E5-mistral; the abstract reports new state-of-the-art results on BEIR and MTEB). https://arxiv.org/abs/2401.00368 ; Muennighoff et al., Generative Representational Instruction Tuning, 2024 (GritLM; the abstract reports a new state of the art on MTEB). https://arxiv.org/abs/2402.09906 ; Lee et al., NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, 2024 (the abstract reports NV-Embed-v1 and v2 at No. 1 on the MTEB leaderboard as of May 24 and August 30, 2024). https://arxiv.org/abs/2405.17428 ↩

  5. Subramanya et al., DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node, NeurIPS 2019. https://www.microsoft.com/en-us/research/publication/diskann-fast-accurate-billion-point-nearest-neighbor-search-on-a-single-node/ ↩

  6. OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, LLM09:2026 Vector and Embedding Weaknesses. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (original text in the GenAI Security Project's GenAI-LLM-Top10 repository, 2026/final directory) ↩

  7. OpenAI, Reasoning models (official documentation): "While reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens." https://developers.openai.com/api/docs/guides/reasoning ↩

  8. OpenAI, Learning to reason with LLMs, September 2024: "performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)." https://openai.com/index/learning-to-reason-with-llms/ ↩

  9. vLLM official documentation, Optimization and Tuning: "In V1, chunked prefill is enabled by default whenever possible… It improves inter-token latency (ITL)…" https://docs.vllm.ai/en/stable/configuration/optimization/ ↩

  10. Kaplan et al., Scaling Laws for Neural Language Models, 2020, Section 2.1: "For contexts and models with d_model > n_ctx/12, the context-dependent computational cost per token is a relatively small fraction of the total compute." https://arxiv.org/abs/2001.08361 ↩

  11. OpenAI, Prompt caching. https://developers.openai.com/api/docs/guides/prompt-caching ; Anthropic, Prompt caching. https://platform.claude.com/docs/en/build-with-claude/prompt-caching ↩

2nd Edition · October 2026