Excellent! Letβs expand on π Scraping multiple URLs in a batch using your AI agent. This allows your agent to gather and process data from several web pages at once, summarize or compare them, and perform large-scale research or monitoring tasks.
π Scraping Multiple URLs in a Batch with Your AI Agent
π― Objective
By the end of this lesson, you will:
-
Enable your AI agent to accept and process a list of URLs
-
Scrape content from each page and aggregate or analyze the results
-
Build use cases like bulk blog summaries, product comparisons, news monitoring, and more
π Updated Tools and Code
β Step 1: Accept a List of URLs
Modify your scraper function to loop through multiple URLs.
def batch_scrape(urls: list, max_per_url=5):
results = []
for url in urls:
try:
response = requests.get(url, headers={"User-Agent": "Mozilla/5.0"})
soup = BeautifulSoup(response.content, "html.parser")
paragraphs = [p.get_text() for p in soup.find_all("p")]
content = "n".join(paragraphs[:max_per_url])
results.append(f"--- Content from {url} ---n{content}n")
except Exception as e:
results.append(f"--- Error with {url} ---n{str(e)}n")
return "n".join(results)
β Step 2: Wrap into LangChain Tool
def scrape_multiple_urls(input_str):
url_list = [url.strip() for url in input_str.split(',') if url.strip().startswith('http')]
return batch_scrape(url_list)
from langchain.tools import Tool
tools = [
Tool(
name="BatchWebScraper",
func=scrape_multiple_urls,
description="Scrapes multiple URLs given as a comma-separated string"
)
]
β Step 3: Use with Agent
from langchain.agents import initialize_agent, AgentType
from langchain.llms import Ollama # or OpenAI
llm = Ollama(model="mistral")
agent = initialize_agent(
tools=tools,
llm=llm,
agent=AgentType.ZERO_SHOT_REACT_DESCRIPTION,
verbose=True
)
# Test
agent.run("Scrape and summarize these URLs: https://en.wikipedia.org/wiki/Solar_power, https://en.wikipedia.org/wiki/Wind_power")
β Optional: Clean HTML with Trafilatura
Swap in for cleaner results:
import trafilatura
def smart_batch_scrape(urls):
results = []
for url in urls:
try:
downloaded = trafilatura.fetch_url(url)
extracted = trafilatura.extract(downloaded)
summary = extracted[:1500] if extracted else "No readable content."
results.append(f"--- {url} ---n{summary}n")
except Exception as e:
results.append(f"--- Error at {url} ---n{str(e)}")
return "n".join(results)
π‘ Example Use Cases
| Use Case | Description |
|---|---|
| π° News Monitoring Agent | Scrape 5β10 major news URLs daily and summarize |
| ποΈ Product Comparison Agent | Scrape 3 eCommerce sites and compare pricing/features |
| βοΈ Blog Digest Generator | Feed in 5 blog post URLs and generate one-paragraph summaries |
| π§Ύ Content Checker Bot | Scan several competitor websites for specific keywords or offers |
π§ Prompt Examples
"Scrape and compare the key points from these 3 blog posts:
https://blog1.com/post, https://blog2.com/post, https://blog3.com/post"
"Scrape the latest info about solar power from these 5 articles and give me a summary of trends."
π§βπ» Streamlit Batch Scraper UI (Optional)
import streamlit as st
st.title("π Batch Web Scraper Agent")
urls = st.text_area("Enter URLs (comma-separated)")
if st.button("Scrape and Summarize"):
result = agent.run(f"Scrape and summarize: {urls}")
st.markdown("### π€ AI Response")
st.write(result)
β Summary
| What You Did | Outcome |
|---|---|
| π§ Built multi-URL logic | AI now handles a list of web pages |
| π Connected to LangChain tool | Agent uses it for summarizing, comparing, analyzing |
| π Enabled real-world automation | Blog digests, competitive tracking, multi-source research |
π Want to Go Further?
Would you like help with:
-
π Scheduling periodic scrapes (e.g., daily summary agent)?
-
π Saving scraped content to a database or file?
-
π Training a custom model with scraped content (RAG setup)?
Just tell me your goal and Iβll guide the next step!
55
