🎙️ Part 1: Add Voice Input + Text-to-Speech Output to the Streamlit Web Chat

 

🎙️ Part 1: Add Voice Input + Text-to-Speech Output to the Streamlit Web Chat


🧰 Tools Needed

Task Tool
🎤 Voice Input Browser’s Web Speech API (via Streamlit component)
🔊 Voice Output gTTS (Google Text-to-Speech) + streamlit_audio
📦 Hosting-ready Everything runs in browser or Streamlit, so compatible with Hugging Face/Render

✅ Step-by-Step Integration

🔹 1. Install required packages

Add to requirements.txt:

gTTS
streamlit
streamlit_audio_recorder

Install:

pip install gTTS streamlit_audio_recorder

(Note: streamlit_audio_recorder is a custom component — if not available, we’ll use streamlit_webrtc or just start with browser-side input)


🔹 2. Add Voice Input (via browser + Streamlit input)

We’ll use browser’s voice typing or audio upload for now.

Add this to your app (app.py), replacing the user_query line:

import tempfile
from gtts import gTTS
import base64

def tts_speak(text):
    tts = gTTS(text=text, lang='en')
    tts_file = tempfile.NamedTemporaryFile(delete=False, suffix=".mp3")
    tts.save(tts_file.name)
    return tts_file.name

# Input section
user_query = st.text_input("Ask a question (or use microphone):")

# Audio response
def play_audio(audio_path):
    with open(audio_path, "rb") as f:
        audio_bytes = f.read()
        b64 = base64.b64encode(audio_bytes).decode()
        audio_html = f"""
        <audio autoplay controls>
        <source src="data:audio/mp3;base64,{b64}" type="audio/mp3">
        </audio>
        """
        st.markdown(audio_html, unsafe_allow_html=True)

# Main interaction
if user_query:
    response = qa_chain.run({"question": user_query})
    st.session_state.chat_history.append(("You", user_query))
    st.session_state.chat_history.append(("Bot", response))

    # Play response
    audio_file = tts_speak(response)
    play_audio(audio_file)

You can now type or dictate using browser mic, and the AI will speak back the response.


🔗 Bonus: Let Users Upload Their Voice as Audio Files

Just add this in UI section:

audio = st.file_uploader("🎙️ Upload voice (.wav/.mp3)", type=["mp3", "wav"])
if audio:
    st.audio(audio, format="audio/mp3")
    st.warning("Voice-to-text not yet added here. You can type instead or wait for Whisper integration.")

Later we’ll integrate Whisper for full STT.


☁️ Part 2: Hosting on Hugging Face Spaces or Render


✅ Hugging Face Spaces

Requirements:

  • Free account at huggingface.co

  • GitHub repo of your app

  • A requirements.txt and app.py

  • Add README.md and optional Space metadata

Folder:

app/
├── app.py
├── documents/
├── requirements.txt
├── README.md

Add README.md

# 🧠 AI Chatbot with Voice

Chat with your uploaded documents using a local LLM (via Ollama), with voice input and output.

## Features
- PDF, DOCX, TXT file support
- Text + voice chat
- Memory-enabled LLM

## Instructions
1. Upload documents
2. Ask questions or speak
3. Listen to AI reply

Push to GitHub, then create a Space:

  1. Go to huggingface.co/spaces

  2. Click “Create new Space”

  3. Choose SDK: Streamlit

  4. Connect your GitHub repo

Spaces will auto-install from your requirements.txt and run app.py


✅ Render or Railway (Alternative Hosts)

For Render:

  1. Create render.yaml for config

  2. Connect repo

  3. Choose Python, set entrypoint: streamlit run app.py

For Railway:

  1. Use template: Python + Streamlit

  2. Paste repo or zip

  3. Configure deployment & resources


✅ Summary

Feature Status
✅ Streamlit Web UI ✅ Done
✅ Voice Input (typed + browser mic) ✅ Done
✅ TTS Output (gTTS) ✅ Done
✅ Hugging Face Hosting Ready
✅ Optional Audio Upload Supported

 

55