Understanding Modern LLM Systems: A Field Guide to RAG, Agents, and Beyond · A concept map from fragmented knowledge to systemic understanding
APPENDIX

Appendix · Toolkit Overview

Below, organized by pipeline stage, are common tools and platforms as of 2026, listed for reference. This field moves fast — the maintenance status, naming, and pricing of specific tools can change at any time, so verify before making production decisions.

By pipeline stage (common tools/platforms for each):

Stage What This Step Does Common Tools / Platforms
Document parsing / chunking Reading formats like PDF and Word into text, and splitting into small segments Unstructured, LlamaParse, LlamaIndex's loader ecosystem, Docling
Embedding models Converting text into vectors OpenAI embeddings, Cohere, BGE, Voyage (part of MongoDB since 20251), sentence-transformers
Vector databases Storing vectors and performing similarity search Pinecone, Weaviate, Milvus, Chroma, Qdrant, pgvector
Reranking models Precision-ranking rough-screened results, keeping only the most relevant few Cohere Rerank, BGE-reranker, Jina Reranker (Jina AI part of Elastic since 20252)
Orchestration / RAG frameworks Development frameworks that wire the entire pipeline together LangChain, LlamaIndex, Haystack, RAGFlow
Agent frameworks Development frameworks supporting ReAct loops and multi-agent collaboration LangGraph, CrewAI, Microsoft Agent Framework; vendor-specific SDKs: OpenAI Agents SDK, Claude Agent SDK, Google ADK; AutoGen (now in maintenance mode, succeeded by Microsoft Agent Framework3)
Inference serving Running large models and exposing them as a service vLLM, SGLang, TGI (Text Generation Inference, archived March 2026; the project recommends migrating to vLLM or SGLang4)
Evaluation Measuring how well a RAG/Agent system performs RAGAS, DeepEval, TruLens, promptfoo
Observability Logging, tracing, and debugging the entire system's operation LangSmith, Langfuse, Arize Phoenix

("Reranking models" correspond to the precision-ranking stage of the Section 3.2 funnel; "evaluation" and "observability" correspond to the log-feedback mechanism described in Chapter One's Loop D — the tools are simply off-the-shelf products that implement these mechanisms.)

By brand ecosystem (product suites from the same company/community):

  • LangChain family: LangChain (base framework) + LangGraph (purpose-built for agent orchestration) + LangSmith (observability / debugging) — three products within the same ecosystem designed to work together.
  • Hugging Face family: the Model Hub (the world's largest open-source model repository) + Transformers (a library for conveniently loading various models) + TGI (inference serving, now archived) + a full suite of open-source tools.
  • Cloud-provider MaaS: AWS Bedrock, Microsoft Foundry (formerly Azure AI Foundry, renamed at Microsoft Ignite in November 20255), Google Gemini Enterprise Agent Platform (evolved from Vertex AI, April 20266), and similar services package embedding and generation models as managed cloud services, allowing enterprises to make API calls without building their own servers.


  1. MongoDB press release, February 24, 2025: acquisition of Voyage AI. https://investors.mongodb.com/news-releases/news-release-details/mongodb-announces-acquisition-voyage-ai-enable-organizations ↩

  2. Elastic press release (Business Wire), October 9, 2025: Elastic Completes Acquisition of Jina AI. https://www.businesswire.com/news/home/20251009619654/en/Elastic-Completes-Acquisition-of-Jina-AI-a-Leader-in-Frontier-Models-for-Multimodal-and-Multilingual-Search ↩

  3. Microsoft, autogen repository README: "AutoGen is now in maintenance mode... community managed going forward," with Microsoft Agent Framework named as the successor. https://github.com/microsoft/autogen ; Microsoft, agent-framework repository. https://github.com/microsoft/agent-framework ↩

  4. Hugging Face, text-generation-inference repository README: "text-generation-inference is now in maintenance mode"; the repository was archived on 2026-03-21. https://github.com/huggingface/text-generation-inference ↩

  5. Microsoft Learn, What is Microsoft Foundry? (the "Evolution of Foundry" section maps old names to new: Azure AI Studio / Azure AI Foundry → Microsoft Foundry). The rename was announced at Microsoft Ignite on November 18, 2025. https://learn.microsoft.com/en-us/azure/foundry/what-is-foundry ↩

  6. Google Cloud, Introducing Gemini Enterprise Agent Platform, April 22, 2026 (Vertex AI evolved into this platform, with prior functionality folded in). https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-agent-platform ↩

One page left after this: the quick reference card. The whole guide is here, free, and will stay that way.

If you'd rather have it off the browser: 50 English pages / 39 Chinese pages, typeset as PDF and EPUB, 7 original diagrams, 49 footnotes to primary sources — four files in one download, $9.

Get the PDF + EPUB on Gumroad →

2nd Edition · October 2026