🎙️ Part 1: Add Voice Input + Text-to-Speech Output to the Streamlit Web Chat
🧰 Tools Needed
| Task | Tool |
|---|---|
| 🎤 Voice Input | Browser’s Web Speech API (via Streamlit component) |
| 🔊 Voice Output | gTTS (Google Text-to-Speech) + streamlit_audio |
| 📦 Hosting-ready | Everything runs in browser or Streamlit, so compatible with Hugging Face/Render |
✅ Step-by-Step Integration
🔹 1. Install required packages
Add to requirements.txt:
gTTS
streamlit
streamlit_audio_recorder
Install:
pip install gTTS streamlit_audio_recorder
(Note:
streamlit_audio_recorderis a custom component — if not available, we’ll usestreamlit_webrtcor just start with browser-side input)
🔹 2. Add Voice Input (via browser + Streamlit input)
We’ll use browser’s voice typing or audio upload for now.
Add this to your app (app.py), replacing the user_query line:
import tempfile
from gtts import gTTS
import base64
def tts_speak(text):
tts = gTTS(text=text, lang='en')
tts_file = tempfile.NamedTemporaryFile(delete=False, suffix=".mp3")
tts.save(tts_file.name)
return tts_file.name
# Input section
user_query = st.text_input("Ask a question (or use microphone):")
# Audio response
def play_audio(audio_path):
with open(audio_path, "rb") as f:
audio_bytes = f.read()
b64 = base64.b64encode(audio_bytes).decode()
audio_html = f"""
<audio autoplay controls>
<source src="data:audio/mp3;base64,{b64}" type="audio/mp3">
</audio>
"""
st.markdown(audio_html, unsafe_allow_html=True)
# Main interaction
if user_query:
response = qa_chain.run({"question": user_query})
st.session_state.chat_history.append(("You", user_query))
st.session_state.chat_history.append(("Bot", response))
# Play response
audio_file = tts_speak(response)
play_audio(audio_file)
You can now type or dictate using browser mic, and the AI will speak back the response.
🔗 Bonus: Let Users Upload Their Voice as Audio Files
Just add this in UI section:
audio = st.file_uploader("🎙️ Upload voice (.wav/.mp3)", type=["mp3", "wav"])
if audio:
st.audio(audio, format="audio/mp3")
st.warning("Voice-to-text not yet added here. You can type instead or wait for Whisper integration.")
Later we’ll integrate Whisper for full STT.
☁️ Part 2: Hosting on Hugging Face Spaces or Render
✅ Hugging Face Spaces
Requirements:
-
Free account at huggingface.co
-
GitHub repo of your app
-
A
requirements.txtandapp.py -
Add
README.mdand optionalSpace metadata
Folder:
app/
├── app.py
├── documents/
├── requirements.txt
├── README.md
Add README.md
# 🧠 AI Chatbot with Voice
Chat with your uploaded documents using a local LLM (via Ollama), with voice input and output.
## Features
- PDF, DOCX, TXT file support
- Text + voice chat
- Memory-enabled LLM
## Instructions
1. Upload documents
2. Ask questions or speak
3. Listen to AI reply
Push to GitHub, then create a Space:
-
Go to huggingface.co/spaces
-
Click “Create new Space”
-
Choose SDK: Streamlit
-
Connect your GitHub repo
Spaces will auto-install from your
requirements.txtand runapp.py
✅ Render or Railway (Alternative Hosts)
For Render:
-
Create
render.yamlfor config -
Connect repo
-
Choose Python, set entrypoint:
streamlit run app.py
For Railway:
-
Use template: Python + Streamlit
-
Paste repo or zip
-
Configure deployment & resources
✅ Summary
| Feature | Status |
|---|---|
| ✅ Streamlit Web UI | ✅ Done |
| ✅ Voice Input (typed + browser mic) | ✅ Done |
| ✅ TTS Output (gTTS) | ✅ Done |
| ✅ Hugging Face Hosting | Ready |
| ✅ Optional Audio Upload | Supported |
55
