✅ Lesson 6.3: Adding Voice Input and Output (TTS/STT with Coqui.ai, Whisper, etc.)
🎯 Lesson Objectives
By the end of this lesson, you will be able to:
-
Understand how voice interfaces enhance user experience in chat apps.
-
Convert speech to text (STT) using Whisper or browser-based tools.
-
Convert chatbot responses to speech (TTS) using Coqui.ai, pyttsx3, or Google TTS.
-
Integrate voice features into Streamlit, Gradio, or React frontends.
-
Build a simple voice-enabled chatbot interface that listens and speaks.
🧠 1. Why Add Voice Capabilities?
Voice is a natural interface—especially useful for:
-
Accessibility (users with visual impairments)
-
Hands-free usage (e.g., on mobile)
-
More human-like interaction
Adding speech-to-text (STT) and text-to-speech (TTS) creates a conversational AI experience.
🗣️ 2. Speech-to-Text (STT): From Voice to Text
✅ Option A: Using Whisper by OpenAI (Highly Accurate)
Whisper is a deep learning model that converts spoken audio to text.
🔧 Install Whisper:
pip install openai-whisper
🔧 Convert Audio File to Text:
import whisper
model = whisper.load_model("base")
result = model.transcribe("user_audio.wav")
print(result["text"])
⚠️ You need to first record audio, e.g., from browser or with Python microphone.
✅ Option B: Using Gradio or Streamlit Audio Input
In Gradio:
import gradio as gr
import whisper
model = whisper.load_model("base")
def transcribe(audio):
audio_path = audio # audio is a file path
result = model.transcribe(audio_path)
return result["text"]
gr.Interface(fn=transcribe, inputs="microphone", outputs="text").launch()
In Streamlit:
import streamlit as st
import soundfile as sf
import whisper
audio_bytes = st.file_uploader("Upload audio", type=["wav", "mp3"])
if audio_bytes:
with open("temp.wav", "wb") as f:
f.write(audio_bytes.read())
model = whisper.load_model("base")
result = model.transcribe("temp.wav")
st.write("You said:", result["text"])
🔊 3. Text-to-Speech (TTS): From Bot Response to Voice
✅ Option A: Using Coqui TTS (Open Source, Natural Voice)
🔧 Install:
pip install TTS
🔧 Generate Voice:
from TTS.api import TTS
tts = TTS(model_name="tts_models/en/ljspeech/tacotron2-DDC", progress_bar=False, gpu=False)
tts.tts_to_file(text="Hello there! How can I help you?", file_path="output.wav")
You can play the audio in Streamlit:
import streamlit as st
audio_file = open("output.wav", "rb")
st.audio(audio_file.read())
✅ Option B: Using pyttsx3 (Offline TTS for Quick Use)
import pyttsx3
engine = pyttsx3.init()
engine.say("Hello! I'm your chatbot.")
engine.runAndWait()
✅ Works offline
⚠️ Robotic-sounding; best for simple prototypes
✅ Option C: Using gTTS (Google TTS API)
from gtts import gTTS
tts = gTTS("Hi there! How can I assist you today?", lang="en")
tts.save("response.mp3")
Then play it:
import streamlit as st
st.audio("response.mp3")
🔗 4. Full Workflow: Voice-In Voice-Out Chat
🧱 Architecture:
[User speaks] → [STT → Text] → [Chatbot LLM] → [Text → TTS] → [Bot speaks]
✅ Streamlit Integration (Simplified):
import streamlit as st
import whisper
from TTS.api import TTS
import openai
st.title("Voice Chatbot")
# Upload or record audio
audio_file = st.file_uploader("Upload your voice", type=["wav", "mp3"])
if audio_file:
with open("temp.wav", "wb") as f:
f.write(audio_file.read())
model = whisper.load_model("base")
result = model.transcribe("temp.wav")
st.write("You said:", result["text"])
# Send to LLM
openai.api_key = "your-api-key"
response = openai.ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": result["text"]}
]
)
reply = response.choices[0].message.content
st.write("Bot:", reply)
# TTS
tts = TTS("tts_models/en/ljspeech/tacotron2-DDC")
tts.tts_to_file(text=reply, file_path="response.wav")
st.audio("response.wav")
🔐 5. Privacy & Security Considerations
| Concern | Recommendation |
|---|---|
| Sensitive voice data | Avoid storing raw audio unless necessary |
| User consent | Ask before using microphone |
| Offline STT/TTS | Use Whisper and Coqui locally for privacy |
| Access control | Secure TTS/STT endpoints if exposed via API |
🧪 6. Practice Activity
🔧 Assignment:
Build a Streamlit or Gradio voice chat interface.
Use Whisper to convert speech to text.
Send the query to your LLM backend.
Convert the response to speech using Coqui or gTTS.
Add a play button to hear the bot’s response.
❓ 7. Comprehension Check
-
What’s the difference between Whisper and pyttsx3?
-
Which TTS method would you use for an offline app?
-
How does voice improve accessibility?
-
What precautions should you take when storing voice data?
📘 8. Further Resources
95
