✅ Lesson 6.4: Adding Image Upload or Multimodal Features
🎯 Lesson Objectives
By the end of this lesson, you will be able to:
-
Understand what multimodal capabilities are and why they matter.
-
Implement image upload functionality in your chatbot interface.
-
Integrate image analysis or captioning using models like OpenAI GPT-4 Vision, CLIP, or BLIP.
-
Extend chatbot conversations to include both visual and text input/output.
-
Use frameworks like Streamlit, Gradio, or React to handle image-based interactions.
🧠 1. What Are Multimodal Features?
A multimodal chatbot can accept and process more than just text—for example:
-
🖼️ Images
-
🔊 Audio
-
📹 Video
-
📍 Location data
Why it matters:
-
More expressive communication
-
Visual problem-solving (e.g., “What’s wrong with this code screenshot?”)
-
Educational and accessibility use cases
-
E-commerce (e.g., “Find similar items to this photo”)
🛠️ 2. Uploading Images in Your UI
✅ In Streamlit:
import streamlit as st
from PIL import Image
st.title("Upload an Image")
uploaded_file = st.file_uploader("Choose an image", type=["jpg", "jpeg", "png"])
if uploaded_file:
image = Image.open(uploaded_file)
st.image(image, caption="Uploaded Image", use_column_width=True)
✅ In Gradio:
import gradio as gr
def greet_image(img):
return "Image uploaded successfully!"
gr.Interface(fn=greet_image, inputs="image", outputs="text").launch()
🤖 3. Using AI to Analyze Images
Once an image is uploaded, use AI to:
-
Describe it (captioning)
-
Answer questions about it (VQA)
-
Extract text (OCR)
-
Classify it (image recognition)
✅ Option A: OpenAI GPT-4 Vision
Requires GPT-4 with image input capability.
import openai
with open("myimage.jpg", "rb") as img:
response = openai.ChatCompletion.create(
model="gpt-4-vision-preview",
messages=[
{"role": "system", "content": "You're a helpful assistant who explains images."},
{"role": "user", "content": [
{"type": "text", "text": "What do you see in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
]}
],
max_tokens=500
)
print(response['choices'][0]['message']['content'])
You’ll need to base64-encode the image for API use.
✅ Option B: Using BLIP (Salesforce’s Image Captioning)
from transformers import BlipProcessor, BlipForConditionalGeneration
from PIL import Image
import torch
processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-base")
model = BlipForConditionalGeneration.from_pretrained("Salesforce/blip-image-captioning-base")
image = Image.open("myimage.jpg").convert('RGB')
inputs = processor(image, return_tensors="pt")
out = model.generate(**inputs)
caption = processor.decode(out[0], skip_special_tokens=True)
print("Caption:", caption)
✅ Option C: CLIP for Zero-Shot Classification
import torch
import clip
from PIL import Image
model, preprocess = clip.load("ViT-B/32")
image = preprocess(Image.open("cat.jpg")).unsqueeze(0)
text = clip.tokenize(["a photo of a cat", "a photo of a dog"])
with torch.no_grad():
image_features = model.encode_image(image)
text_features = model.encode_text(text)
logits = (image_features @ text_features.T).softmax(dim=-1)
print("Probabilities:", logits)
🔄 4. Connecting Image + Text in a Chat Flow
Let the user ask:
-
“What’s in this image?”
-
“Is there a bird in this photo?”
-
“Can you describe this screenshot?”
Combine with chatbot logic:
user_query = "What do you see?"
image_caption = generate_caption(uploaded_file)
combined_prompt = f"User uploaded an image and asked: '{user_query}'. The image shows: '{image_caption}'"
response = query_llm(combined_prompt)
🔗 5. Use Cases for Multimodal Chatbots
| Use Case | Example |
|---|---|
| 🧑🎓 Learning Assistant | “Explain this diagram” |
| 🛍️ Product Finder | “Find me shoes like this” |
| 🖥️ Debugging Assistant | “What’s wrong with this error screenshot?” |
| 📖 Accessibility Reader | “Read and explain the image text” (OCR + LLM) |
| 🎨 Art Feedback | “Give me feedback on my painting” |
🧪 6. Practice Activity
🔧 Assignment:
Create a chatbot interface (Streamlit or Gradio) with an image uploader.
Use an image captioning model (e.g., BLIP or GPT-4 Vision) to generate a description.
Let the user ask a question related to the uploaded image.
Feed the image caption and user question into your LLM to generate a multimodal response.
Bonus: Show image classification or OCR features alongside captions.
❓ 7. Comprehension Check
-
What is the difference between BLIP and CLIP?
-
How can GPT-4 Vision enhance a chatbot?
-
Why is image captioning important before LLM response?
-
What types of user tasks benefit from multimodal input?
📘 8. Further Resources
Look at
-
A Streamlit app that lets users upload an image and chat about it?
-
A Gradio interface with both image and voice input?
-
A React + Flask multimodal chatbot template?
108
