✅ Lesson 4.2: Adding PDFs, CSVs, Text, and URLs to a Vector Database (e.g., ChromaDB, Weaviate, Qdrant)
🎯 Lesson Objectives
By the end of this lesson, learners will be able to:
-
Understand what vector databases are and why they’re used in LLM workflows.
-
Convert unstructured and structured data (PDF, CSV, text, and web content) into embeddings.
-
Store these embeddings in a vector database like ChromaDB, Weaviate, or Qdrant.
-
Implement indexing, retrieval, and basic search queries.
-
Integrate these vector stores with LLM-powered search or RAG (Retrieval-Augmented Generation).
🧠 1. Why Vector Databases Are Essential for LLMs
Large Language Models are stateless by default. If you want them to recall knowledge from your data (PDFs, CSVs, websites, etc.), you must:
-
Convert your content into vectors (embeddings).
-
Store and retrieve them efficiently using a vector database.
-
Feed relevant results into the LLM during a prompt (RAG).
Use cases:
-
AI customer service with your documents
-
LLM-powered internal search engines
-
Q&A bots based on your own PDFs or CSVs
📦 2. Supported Data Types
| Data Type | Examples | Challenges |
|---|---|---|
| Manuals, research papers, policies | Layout, multi-column text | |
| CSV | Product catalogs, FAQs, customer records | Cleaning, column selection |
| Text | .txt files, logs, markdown notes |
Chunking |
| URLs | Blog posts, websites, Wikipedia entries | Scraping and summarizing |
🧰 3. Tools You’ll Need
| Purpose | Tool |
|---|---|
| Embeddings | OpenAI, Hugging Face, or sentence-transformers |
| Vector Store | ChromaDB, Weaviate, Qdrant |
| Text Extraction | PyMuPDF (fitz), pandas, BeautifulSoup |
| Chunking | LangChain, LlamaIndex, Haystack |
🟣 4. ChromaDB Quickstart (Local)
✅ Install:
pip install chromadb langchain pypdf sentence-transformers
✅ Add a PDF:
from langchain.document_loaders import PyPDFLoader
from langchain.embeddings import SentenceTransformerEmbeddings
from langchain.vectorstores import Chroma
# Load PDF
loader = PyPDFLoader("sample.pdf")
pages = loader.load_and_split()
# Create embedding function
embeddings = SentenceTransformerEmbeddings(model_name="all-MiniLM-L6-v2")
# Store in ChromaDB
db = Chroma.from_documents(pages, embedding=embeddings, persist_directory="db")
db.persist()
✅ Query:
query = "What is the refund policy?"
results = db.similarity_search(query, k=3)
for r in results:
print(r.page_content)
🟡 5. Weaviate (Cloud or Docker)
✅ Install:
pip install weaviate-client sentence-transformers
✅ Create Weaviate Schema:
import weaviate
client = weaviate.Client("http://localhost:8080")
schema = {
"class": "Document",
"vectorizer": "none",
"properties": [{"name": "content", "dataType": ["text"]}]
}
client.schema.create_class(schema)
✅ Add CSV Content:
import pandas as pd
from sentence_transformers import SentenceTransformer
df = pd.read_csv("products.csv")
model = SentenceTransformer("all-MiniLM-L6-v2")
for row in df.itertuples():
vector = model.encode(row.Description)
client.data_object.create({
"content": row.Description
}, "Document", vector=vector.tolist())
✅ Query Weaviate:
res = client.query.get("Document", ["content"]).with_near_text({"concepts": ["refund policy"]}).with_limit(3).do()
print(res)
🔴 6. Qdrant (Powerful and Scalable)
✅ Install:
pip install qdrant-client sentence-transformers
✅ Start Qdrant (Docker):
docker run -p 6333:6333 -v $(pwd)/qdrant_data:/qdrant/storage qdrant/qdrant
✅ Upload Texts:
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct, VectorParams
from sentence_transformers import SentenceTransformer
client = QdrantClient(host="localhost", port=6333)
model = SentenceTransformer("all-MiniLM-L6-v2")
texts = ["Return policy is valid for 30 days.", "Shipping takes 3–5 business days."]
vectors = model.encode(texts)
client.recreate_collection(
collection_name="documents",
vectors_config=VectorParams(size=len(vectors[0]), distance="Cosine")
)
client.upsert(
collection_name="documents",
points=[PointStruct(id=i, vector=vectors[i], payload={"text": texts[i]}) for i in range(len(texts))]
)
✅ Query:
query = model.encode("How long is the return window?")
results = client.search("documents", query_vector=query, limit=2)
for r in results:
print(r.payload["text"])
🌐 7. Adding URLs as Data
Use BeautifulSoup or newspaper3k to scrape website content:
from bs4 import BeautifulSoup
import requests
url = "https://example.com/refund-policy"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")
text = soup.get_text().replace("n", " ")
chunks = [text[i:i+500] for i in range(0, len(text), 500)]
Then embed and store chunks just like text documents.
🔄 8. Best Practices: Chunking, Tagging, and Metadata
| Practice | Description |
|---|---|
| Chunking | Divide content into 300–500 token pieces for better recall |
| Metadata | Attach fields like filename, date, or tags to documents |
| Preprocessing | Remove headers, footers, tables of contents |
| Version control | Version each document upload for traceability |
🧪 9. Practice Activity
🔧 Assignment:
Choose one database (ChromaDB, Weaviate, or Qdrant).
Upload:
One PDF file
One CSV file
One block of plain text
One URL
Chunk and store the data with proper metadata.
Run a search query like:
"How to request a refund?"Log and compare top 3 results.
❓ 10. Comprehension Check
-
What is the role of embeddings in a vector database?
-
Why is chunking necessary before indexing text?
-
Which vector database would you choose for scaling to millions of documents, and why?
📘 11. Further Resources
81
