✅ Lesson 5.3: Implementing Contextual Retrieval with RAG (Retrieval-Augmented Generation)
🎯 Lesson Objectives
By the end of this lesson, you will:
-
Understand what Retrieval-Augmented Generation (RAG) is and why it’s useful.
-
Set up a basic RAG pipeline using your own documents (PDFs, CSVs, text, or web data).
-
Combine user queries with vector search results to generate highly relevant responses.
-
Implement RAG using LangChain or a custom method with your preferred vector DB (e.g., ChromaDB, Qdrant).
-
Integrate RAG with your LLM backend for an enhanced chat experience.
🧠 1. What is RAG and Why Does It Matter?
Retrieval-Augmented Generation (RAG) enhances LLMs by injecting external knowledge at runtime from a custom document base.
🔍 Without RAG:
LLMs answer based on pretraining data, which may be outdated or inaccurate.
📚 With RAG:
You embed a document retrieval step:
-
User submits a query.
-
The system searches a vector database of your documents.
-
Top-k relevant chunks are appended to the prompt.
-
LLM generates a grounded, accurate answer.
RAG is ideal for:
-
Internal knowledge bases
-
Legal, medical, and policy documents
-
Customer support chatbots
-
Academic research assistants
🧱 2. Components of a RAG Pipeline
| Component | Description | Example Tool |
|---|---|---|
| Embeddings | Convert documents and queries into vectors | SentenceTransformers, OpenAI |
| Vector Store | Store and search document embeddings | ChromaDB, Weaviate, Qdrant |
| Retriever | Find relevant chunks for a query | LangChain Retriever, manual |
| Prompt Builder | Construct prompt with retrieved context + user input | LangChain, custom code |
| LLM | Generate final response | GPT, Mistral, LLaMA |
🛠️ 3. Building a RAG Pipeline (Basic Version)
Step 1: Prepare and Embed Your Documents
from langchain.document_loaders import PyPDFLoader
from langchain.embeddings import HuggingFaceEmbeddings
from langchain.vectorstores import Chroma
loader = PyPDFLoader("policy.pdf")
documents = loader.load_and_split()
embedding_model = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
vectorstore = Chroma.from_documents(documents, embedding_model, persist_directory="./db")
vectorstore.persist()
Step 2: Retrieve Relevant Chunks Based on a Query
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})
query = "What is the return policy?"
relevant_docs = retriever.get_relevant_documents(query)
context = "nn".join([doc.page_content for doc in relevant_docs])
Step 3: Construct a Prompt + Call the LLM
from openai import OpenAI
import openai
openai.api_key = "your-api-key"
prompt = f"""
You are a helpful assistant. Use the following context to answer the question.
Context:
{context}
Question: {query}
Answer:
"""
response = openai.ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": prompt}]
)
print(response.choices[0].message.content)
🧠 4. Using LangChain’s RAG Chain (Optional)
If you’re using LangChain, you can do this in one go with a RetrievalQA chain:
from langchain.chains import RetrievalQA
from langchain.llms import OpenAI
llm = OpenAI(temperature=0.7)
qa_chain = RetrievalQA.from_chain_type(
llm=llm,
retriever=vectorstore.as_retriever(),
return_source_documents=True
)
query = "How do I request a refund?"
result = qa_chain.run(query)
print(result)
🔄 5. Integrating RAG into Your Backend API
Example: Flask + RAG
@app.route("/rag-chat", methods=["POST"])
def rag_chat():
data = request.json
query = data.get("query")
# Step 1: Retrieve relevant context
context_docs = retriever.get_relevant_documents(query)
context = "nn".join([doc.page_content for doc in context_docs])
# Step 2: Build prompt
prompt = f"Answer the question using the following context:nn{context}nnQuestion: {query}"
# Step 3: Get LLM response
res = openai.ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": prompt}]
)
return jsonify({"reply": res.choices[0].message.content})
🧪 6. Practice Activity
🔧 Assignment:
Choose one or more documents (PDF, CSV, or web content).
Chunk and embed them using SentenceTransformers or OpenAI embeddings.
Store embeddings in a vector database (Chroma/Qdrant).
Create a basic API or script that:
Retrieves top 3 relevant chunks
Builds a context-augmented prompt
Generates an answer via your LLM
Integrate this into your Streamlit or React chat UI.
🔐 7. Best Practices
| Practice | Why It Matters |
|---|---|
| Chunk by meaning | Use sentence boundaries or semantic chunkers |
| Attach metadata | Track source, page, or filename for transparency |
| Limit token length | Avoid context overflow in LLM prompt |
| Filter junk data | Remove irrelevant or duplicated content |
| Log retrieval hits | Track which documents were used to answer each query |
❓ 8. Comprehension Check
-
What is the main advantage of RAG over traditional LLM use?
-
Why is chunking important before indexing documents?
-
How would you improve retrieval precision for a legal document?
📘 9. Further Resources
82
