Understanding Context and RAG
Before talking about RAG, it’s important to understand context in large language models (LLMs).
When you send a prompt, the LLM first breaks it into tokens — small units of text, roughly 3–4 characters of English on average. These tokens, along with any chat history and documents you provide, make up the context. The context window is the maximum number of tokens the model can handle at once, and it includes both your input and the model’s output. Different models have different context sizes.

Once the LLM has your context, it doesn’t “search” its training data like a database. Instead, it predicts the next token based on probabilities learned during training, using the information in the current context to guide that prediction.
Note: Google, OpenAI, Anthropic, and others have powerful closed-source and open-source LLMs with varying costs and max context windows. There is a 3rd party service from OpenRouter.ai which lets you see the specs of the varying models and also lets you invoke any of them through their service. For models that I don't host myself, I now connect to them through OpenRouter.ai instead of going directly through the OpenAIs of the world. This way all my prompts and chats are measured for cost and context and I can make the best decisions about what models I'll use and whether I'll host them locally or not. Furthermore, I get one bill even when I use different model providers. Microsoft Azure AI foundry includes a similar service.
If a model already knows something from its training data (e.g., the text of the U.S. Constitution), you can just ask about it directly. But for new material — like your own SOPs or a proprietary dataset — you must feed it into the context yourself. The catch is that if the total size of your prompt, documents, and chat history exceeds the model’s context limit, some of it won’t fit.
This is where Retrieval Augmented Generation (RAG) comes in. RAG works outside the LLM to select only the most relevant chunks of your documents to inject into the context, so the model can work with them efficiently and within limits.

Why RAG Matters Beyond the Usual Sales Pitch
Most of the help you’ll find online for RAG focuses on a business need: How can I teach an LLM about my proprietary or personal information so it gains additional “intelligence”? The truth is, you’ve already been doing that with every prompt — every input adds temporary knowledge to the model’s working memory. People often frame RAG as the method for feeding an LLM large volumes of new information, and that’s true. But it’s equally important to understand RAG as a tool for managing context — the LLM’s working memory.
Understanding why you need RAG is essential. Even if context limits eventually expand to 5M or 100M tokens, the need for relevance filtering won’t go away. Bigger windows can still waste processing power and attention on irrelevant details.
Some see LLMs as limitless. But those pushing boundaries already know the reality: large contexts and large models require significant hardware and cost. This is why running LLMs locally is appealing — you control the expense. Yet what you can host depends on your infrastructure. Bigger contexts require more powerful, pricier systems.
The good news? LLM technology is no longer locked up by a handful of companies. Anyone can download open-source models or even train their own. The real barrier is hardware cost — which is why learning RAG and context management is a practical skill. With it, you can deliver business value through LLMs even with tight budgets and constrained resources, enabling smaller organizations to benefit from the same AI capabilities as the big players.
People often frame RAG as the method for feeding an LLM large volumes of new information, and that’s true. But it’s equally important to understand RAG as a tool for managing context — the LLM’s working memory.



