Lesson 5.3: Implementing Contextual Retrieval with RAG (Retrieval-Augmented Generation)

 


✅ Lesson 5.3: Implementing Contextual Retrieval with RAG (Retrieval-Augmented Generation)


🎯 Lesson Objectives

By the end of this lesson, you will:

  • Understand what Retrieval-Augmented Generation (RAG) is and why it’s useful.

  • Set up a basic RAG pipeline using your own documents (PDFs, CSVs, text, or web data).

  • Combine user queries with vector search results to generate highly relevant responses.

  • Implement RAG using LangChain or a custom method with your preferred vector DB (e.g., ChromaDB, Qdrant).

  • Integrate RAG with your LLM backend for an enhanced chat experience.


🧠 1. What is RAG and Why Does It Matter?

Retrieval-Augmented Generation (RAG) enhances LLMs by injecting external knowledge at runtime from a custom document base.

🔍 Without RAG:

LLMs answer based on pretraining data, which may be outdated or inaccurate.

📚 With RAG:

You embed a document retrieval step:

  1. User submits a query.

  2. The system searches a vector database of your documents.

  3. Top-k relevant chunks are appended to the prompt.

  4. LLM generates a grounded, accurate answer.

RAG is ideal for:

  • Internal knowledge bases

  • Legal, medical, and policy documents

  • Customer support chatbots

  • Academic research assistants


🧱 2. Components of a RAG Pipeline

Component Description Example Tool
Embeddings Convert documents and queries into vectors SentenceTransformers, OpenAI
Vector Store Store and search document embeddings ChromaDB, Weaviate, Qdrant
Retriever Find relevant chunks for a query LangChain Retriever, manual
Prompt Builder Construct prompt with retrieved context + user input LangChain, custom code
LLM Generate final response GPT, Mistral, LLaMA

🛠️ 3. Building a RAG Pipeline (Basic Version)


Step 1: Prepare and Embed Your Documents

from langchain.document_loaders import PyPDFLoader
from langchain.embeddings import HuggingFaceEmbeddings
from langchain.vectorstores import Chroma

loader = PyPDFLoader("policy.pdf")
documents = loader.load_and_split()

embedding_model = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")

vectorstore = Chroma.from_documents(documents, embedding_model, persist_directory="./db")
vectorstore.persist()

Step 2: Retrieve Relevant Chunks Based on a Query

retriever = vectorstore.as_retriever(search_kwargs={"k": 3})

query = "What is the return policy?"
relevant_docs = retriever.get_relevant_documents(query)

context = "nn".join([doc.page_content for doc in relevant_docs])

Step 3: Construct a Prompt + Call the LLM

from openai import OpenAI
import openai

openai.api_key = "your-api-key"

prompt = f"""
You are a helpful assistant. Use the following context to answer the question.

Context:
{context}

Question: {query}
Answer:
"""

response = openai.ChatCompletion.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": prompt}]
)

print(response.choices[0].message.content)

🧠 4. Using LangChain’s RAG Chain (Optional)

If you’re using LangChain, you can do this in one go with a RetrievalQA chain:

from langchain.chains import RetrievalQA
from langchain.llms import OpenAI

llm = OpenAI(temperature=0.7)
qa_chain = RetrievalQA.from_chain_type(
    llm=llm,
    retriever=vectorstore.as_retriever(),
    return_source_documents=True
)

query = "How do I request a refund?"
result = qa_chain.run(query)
print(result)

🔄 5. Integrating RAG into Your Backend API

Example: Flask + RAG

@app.route("/rag-chat", methods=["POST"])
def rag_chat():
    data = request.json
    query = data.get("query")

    # Step 1: Retrieve relevant context
    context_docs = retriever.get_relevant_documents(query)
    context = "nn".join([doc.page_content for doc in context_docs])

    # Step 2: Build prompt
    prompt = f"Answer the question using the following context:nn{context}nnQuestion: {query}"

    # Step 3: Get LLM response
    res = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": prompt}]
    )

    return jsonify({"reply": res.choices[0].message.content})

🧪 6. Practice Activity

🔧 Assignment:

  1. Choose one or more documents (PDF, CSV, or web content).

  2. Chunk and embed them using SentenceTransformers or OpenAI embeddings.

  3. Store embeddings in a vector database (Chroma/Qdrant).

  4. Create a basic API or script that:

    • Retrieves top 3 relevant chunks

    • Builds a context-augmented prompt

    • Generates an answer via your LLM

  5. Integrate this into your Streamlit or React chat UI.


🔐 7. Best Practices

Practice Why It Matters
Chunk by meaning Use sentence boundaries or semantic chunkers
Attach metadata Track source, page, or filename for transparency
Limit token length Avoid context overflow in LLM prompt
Filter junk data Remove irrelevant or duplicated content
Log retrieval hits Track which documents were used to answer each query

❓ 8. Comprehension Check

  1. What is the main advantage of RAG over traditional LLM use?

  2. Why is chunking important before indexing documents?

  3. How would you improve retrieval precision for a legal document?


📘 9. Further Resources


 

82