Excellent! Let’s expand on 🌐 Scraping multiple URLs in a batch using your AI agent. This allows your agent to gather and process data from several web pages at once, summarize or compare them, and perform large-scale research or monitoring tasks.


🌐 Scraping Multiple URLs in a Batch with Your AI Agent


🎯 Objective

By the end of this lesson, you will:

  • Enable your AI agent to accept and process a list of URLs

  • Scrape content from each page and aggregate or analyze the results

  • Build use cases like bulk blog summaries, product comparisons, news monitoring, and more


πŸ›  Updated Tools and Code


βœ… Step 1: Accept a List of URLs

Modify your scraper function to loop through multiple URLs.

def batch_scrape(urls: list, max_per_url=5):
    results = []
    for url in urls:
        try:
            response = requests.get(url, headers={"User-Agent": "Mozilla/5.0"})
            soup = BeautifulSoup(response.content, "html.parser")
            paragraphs = [p.get_text() for p in soup.find_all("p")]
            content = "n".join(paragraphs[:max_per_url])
            results.append(f"--- Content from {url} ---n{content}n")
        except Exception as e:
            results.append(f"--- Error with {url} ---n{str(e)}n")
    return "n".join(results)

βœ… Step 2: Wrap into LangChain Tool

def scrape_multiple_urls(input_str):
    url_list = [url.strip() for url in input_str.split(',') if url.strip().startswith('http')]
    return batch_scrape(url_list)

from langchain.tools import Tool

tools = [
    Tool(
        name="BatchWebScraper",
        func=scrape_multiple_urls,
        description="Scrapes multiple URLs given as a comma-separated string"
    )
]

βœ… Step 3: Use with Agent

from langchain.agents import initialize_agent, AgentType
from langchain.llms import Ollama  # or OpenAI

llm = Ollama(model="mistral")

agent = initialize_agent(
    tools=tools,
    llm=llm,
    agent=AgentType.ZERO_SHOT_REACT_DESCRIPTION,
    verbose=True
)

# Test
agent.run("Scrape and summarize these URLs: https://en.wikipedia.org/wiki/Solar_power, https://en.wikipedia.org/wiki/Wind_power")

βœ… Optional: Clean HTML with Trafilatura

Swap in for cleaner results:

import trafilatura

def smart_batch_scrape(urls):
    results = []
    for url in urls:
        try:
            downloaded = trafilatura.fetch_url(url)
            extracted = trafilatura.extract(downloaded)
            summary = extracted[:1500] if extracted else "No readable content."
            results.append(f"--- {url} ---n{summary}n")
        except Exception as e:
            results.append(f"--- Error at {url} ---n{str(e)}")
    return "n".join(results)

πŸ’‘ Example Use Cases

Use Case Description
πŸ“° News Monitoring Agent Scrape 5–10 major news URLs daily and summarize
πŸ›οΈ Product Comparison Agent Scrape 3 eCommerce sites and compare pricing/features
✍️ Blog Digest Generator Feed in 5 blog post URLs and generate one-paragraph summaries
🧾 Content Checker Bot Scan several competitor websites for specific keywords or offers

🧠 Prompt Examples

"Scrape and compare the key points from these 3 blog posts:
https://blog1.com/post, https://blog2.com/post, https://blog3.com/post"
"Scrape the latest info about solar power from these 5 articles and give me a summary of trends."

πŸ§‘β€πŸ’» Streamlit Batch Scraper UI (Optional)

import streamlit as st

st.title("🌐 Batch Web Scraper Agent")

urls = st.text_area("Enter URLs (comma-separated)")
if st.button("Scrape and Summarize"):
    result = agent.run(f"Scrape and summarize: {urls}")
    st.markdown("### πŸ€– AI Response")
    st.write(result)

βœ… Summary

What You Did Outcome
🧠 Built multi-URL logic AI now handles a list of web pages
πŸ”— Connected to LangChain tool Agent uses it for summarizing, comparing, analyzing
🌍 Enabled real-world automation Blog digests, competitive tracking, multi-source research

πŸ”œ Want to Go Further?

Would you like help with:

  1. πŸ” Scheduling periodic scrapes (e.g., daily summary agent)?

  2. πŸ“‚ Saving scraped content to a database or file?

  3. πŸ“Š Training a custom model with scraped content (RAG setup)?

Just tell me your goal and I’ll guide the next step!

55