✅ Lesson 6.4: Adding Image Upload or Multimodal Features

 

✅ Lesson 6.4: Adding Image Upload or Multimodal Features


🎯 Lesson Objectives

By the end of this lesson, you will be able to:

  • Understand what multimodal capabilities are and why they matter.

  • Implement image upload functionality in your chatbot interface.

  • Integrate image analysis or captioning using models like OpenAI GPT-4 Vision, CLIP, or BLIP.

  • Extend chatbot conversations to include both visual and text input/output.

  • Use frameworks like Streamlit, Gradio, or React to handle image-based interactions.


🧠 1. What Are Multimodal Features?

A multimodal chatbot can accept and process more than just text—for example:

  • 🖼️ Images

  • 🔊 Audio

  • 📹 Video

  • 📍 Location data

Why it matters:

  • More expressive communication

  • Visual problem-solving (e.g., “What’s wrong with this code screenshot?”)

  • Educational and accessibility use cases

  • E-commerce (e.g., “Find similar items to this photo”)


🛠️ 2. Uploading Images in Your UI

✅ In Streamlit:

import streamlit as st
from PIL import Image

st.title("Upload an Image")
uploaded_file = st.file_uploader("Choose an image", type=["jpg", "jpeg", "png"])

if uploaded_file:
    image = Image.open(uploaded_file)
    st.image(image, caption="Uploaded Image", use_column_width=True)

✅ In Gradio:

import gradio as gr

def greet_image(img):
    return "Image uploaded successfully!"

gr.Interface(fn=greet_image, inputs="image", outputs="text").launch()

🤖 3. Using AI to Analyze Images

Once an image is uploaded, use AI to:

  • Describe it (captioning)

  • Answer questions about it (VQA)

  • Extract text (OCR)

  • Classify it (image recognition)

✅ Option A: OpenAI GPT-4 Vision

Requires GPT-4 with image input capability.

import openai

with open("myimage.jpg", "rb") as img:
    response = openai.ChatCompletion.create(
        model="gpt-4-vision-preview",
        messages=[
            {"role": "system", "content": "You're a helpful assistant who explains images."},
            {"role": "user", "content": [
                {"type": "text", "text": "What do you see in this image?"},
                {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
            ]}
        ],
        max_tokens=500
    )
    print(response['choices'][0]['message']['content'])

You’ll need to base64-encode the image for API use.


✅ Option B: Using BLIP (Salesforce’s Image Captioning)

from transformers import BlipProcessor, BlipForConditionalGeneration
from PIL import Image
import torch

processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-base")
model = BlipForConditionalGeneration.from_pretrained("Salesforce/blip-image-captioning-base")

image = Image.open("myimage.jpg").convert('RGB')
inputs = processor(image, return_tensors="pt")
out = model.generate(**inputs)
caption = processor.decode(out[0], skip_special_tokens=True)

print("Caption:", caption)

✅ Option C: CLIP for Zero-Shot Classification

import torch
import clip
from PIL import Image

model, preprocess = clip.load("ViT-B/32")
image = preprocess(Image.open("cat.jpg")).unsqueeze(0)
text = clip.tokenize(["a photo of a cat", "a photo of a dog"])

with torch.no_grad():
    image_features = model.encode_image(image)
    text_features = model.encode_text(text)

    logits = (image_features @ text_features.T).softmax(dim=-1)
    print("Probabilities:", logits)

🔄 4. Connecting Image + Text in a Chat Flow

Let the user ask:

  • “What’s in this image?”

  • “Is there a bird in this photo?”

  • “Can you describe this screenshot?”

Combine with chatbot logic:

user_query = "What do you see?"
image_caption = generate_caption(uploaded_file)
combined_prompt = f"User uploaded an image and asked: '{user_query}'. The image shows: '{image_caption}'"
response = query_llm(combined_prompt)

🔗 5. Use Cases for Multimodal Chatbots

Use Case Example
🧑‍🎓 Learning Assistant “Explain this diagram”
🛍️ Product Finder “Find me shoes like this”
🖥️ Debugging Assistant “What’s wrong with this error screenshot?”
📖 Accessibility Reader “Read and explain the image text” (OCR + LLM)
🎨 Art Feedback “Give me feedback on my painting”

🧪 6. Practice Activity

🔧 Assignment:

  1. Create a chatbot interface (Streamlit or Gradio) with an image uploader.

  2. Use an image captioning model (e.g., BLIP or GPT-4 Vision) to generate a description.

  3. Let the user ask a question related to the uploaded image.

  4. Feed the image caption and user question into your LLM to generate a multimodal response.

  5. Bonus: Show image classification or OCR features alongside captions.


❓ 7. Comprehension Check

  1. What is the difference between BLIP and CLIP?

  2. How can GPT-4 Vision enhance a chatbot?

  3. Why is image captioning important before LLM response?

  4. What types of user tasks benefit from multimodal input?


📘 8. Further Resources


Look at

  • A Streamlit app that lets users upload an image and chat about it?

  • A Gradio interface with both image and voice input?

  • A React + Flask multimodal chatbot template?

 

108