Lesson 6.3: Adding Voice Input and Output (TTS/STT with Coqui.ai, Whisper, etc.)

 


✅ Lesson 6.3: Adding Voice Input and Output (TTS/STT with Coqui.ai, Whisper, etc.)


🎯 Lesson Objectives

By the end of this lesson, you will be able to:

  • Understand how voice interfaces enhance user experience in chat apps.

  • Convert speech to text (STT) using Whisper or browser-based tools.

  • Convert chatbot responses to speech (TTS) using Coqui.ai, pyttsx3, or Google TTS.

  • Integrate voice features into Streamlit, Gradio, or React frontends.

  • Build a simple voice-enabled chatbot interface that listens and speaks.


🧠 1. Why Add Voice Capabilities?

Voice is a natural interface—especially useful for:

  • Accessibility (users with visual impairments)

  • Hands-free usage (e.g., on mobile)

  • More human-like interaction

Adding speech-to-text (STT) and text-to-speech (TTS) creates a conversational AI experience.


🗣️ 2. Speech-to-Text (STT): From Voice to Text

✅ Option A: Using Whisper by OpenAI (Highly Accurate)

Whisper is a deep learning model that converts spoken audio to text.

🔧 Install Whisper:

pip install openai-whisper

🔧 Convert Audio File to Text:

import whisper

model = whisper.load_model("base")
result = model.transcribe("user_audio.wav")
print(result["text"])

⚠️ You need to first record audio, e.g., from browser or with Python microphone.


✅ Option B: Using Gradio or Streamlit Audio Input

In Gradio:

import gradio as gr
import whisper

model = whisper.load_model("base")

def transcribe(audio):
    audio_path = audio  # audio is a file path
    result = model.transcribe(audio_path)
    return result["text"]

gr.Interface(fn=transcribe, inputs="microphone", outputs="text").launch()

In Streamlit:

import streamlit as st
import soundfile as sf
import whisper

audio_bytes = st.file_uploader("Upload audio", type=["wav", "mp3"])
if audio_bytes:
    with open("temp.wav", "wb") as f:
        f.write(audio_bytes.read())
    model = whisper.load_model("base")
    result = model.transcribe("temp.wav")
    st.write("You said:", result["text"])

🔊 3. Text-to-Speech (TTS): From Bot Response to Voice

✅ Option A: Using Coqui TTS (Open Source, Natural Voice)

🔧 Install:

pip install TTS

🔧 Generate Voice:

from TTS.api import TTS

tts = TTS(model_name="tts_models/en/ljspeech/tacotron2-DDC", progress_bar=False, gpu=False)
tts.tts_to_file(text="Hello there! How can I help you?", file_path="output.wav")

You can play the audio in Streamlit:

import streamlit as st
audio_file = open("output.wav", "rb")
st.audio(audio_file.read())

✅ Option B: Using pyttsx3 (Offline TTS for Quick Use)

import pyttsx3

engine = pyttsx3.init()
engine.say("Hello! I'm your chatbot.")
engine.runAndWait()

✅ Works offline
⚠️ Robotic-sounding; best for simple prototypes


✅ Option C: Using gTTS (Google TTS API)

from gtts import gTTS

tts = gTTS("Hi there! How can I assist you today?", lang="en")
tts.save("response.mp3")

Then play it:

import streamlit as st
st.audio("response.mp3")

🔗 4. Full Workflow: Voice-In Voice-Out Chat

🧱 Architecture:

[User speaks] → [STT → Text] → [Chatbot LLM] → [Text → TTS] → [Bot speaks]

✅ Streamlit Integration (Simplified):

import streamlit as st
import whisper
from TTS.api import TTS
import openai

st.title("Voice Chatbot")

# Upload or record audio
audio_file = st.file_uploader("Upload your voice", type=["wav", "mp3"])

if audio_file:
    with open("temp.wav", "wb") as f:
        f.write(audio_file.read())

    model = whisper.load_model("base")
    result = model.transcribe("temp.wav")
    st.write("You said:", result["text"])

    # Send to LLM
    openai.api_key = "your-api-key"
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": result["text"]}
        ]
    )
    reply = response.choices[0].message.content
    st.write("Bot:", reply)

    # TTS
    tts = TTS("tts_models/en/ljspeech/tacotron2-DDC")
    tts.tts_to_file(text=reply, file_path="response.wav")
    st.audio("response.wav")

🔐 5. Privacy & Security Considerations

Concern Recommendation
Sensitive voice data Avoid storing raw audio unless necessary
User consent Ask before using microphone
Offline STT/TTS Use Whisper and Coqui locally for privacy
Access control Secure TTS/STT endpoints if exposed via API

🧪 6. Practice Activity

🔧 Assignment:

  1. Build a Streamlit or Gradio voice chat interface.

  2. Use Whisper to convert speech to text.

  3. Send the query to your LLM backend.

  4. Convert the response to speech using Coqui or gTTS.

  5. Add a play button to hear the bot’s response.


❓ 7. Comprehension Check

  1. What’s the difference between Whisper and pyttsx3?

  2. Which TTS method would you use for an offline app?

  3. How does voice improve accessibility?

  4. What precautions should you take when storing voice data?


📘 8. Further Resources


 

95