🔄 Lesson 8.3: Handling API Failures and Timeouts


🎯 Lesson Objective

By the end of this lesson, you will:

  • Understand the common causes of API errors and timeouts in LLM-powered apps

  • Learn strategies to gracefully handle failures from OpenAI, Ollama, HuggingFace, or custom tools

  • Implement retry logic, user feedback, logging, and fallback mechanisms

  • Build resilient agents that continue functioning under stress, downtime, or latency spikes


💥 1. Why Do API Failures Happen?

LLMs and agent systems depend heavily on:

  • External APIs (e.g., OpenAI, Claude, Cohere)

  • Internal tools or services (e.g., Whisper, vector DBs, file parsers)

Common reasons for failure:

Type of Failure Example
🔌 Network Issues API gateway unreachable, server offline
⏱️ Timeout Long file or query execution (especially PDFs, long chains)
🚫 Rate Limit Exceeded Too many calls per second or minute
🔐 Auth Errors Expired API key, invalid token
💬 Model Overload Too many tokens, big document chunks
⚠️ Model Error LLM misfires or returns 500/502 errors

🧰 2. Handling API Failures Gracefully

Instead of your agent crashing or freezing, use:

✅ Try/Except Blocks (Python Example)

import openai

try:
    response = openai.ChatCompletion.create(
        model="gpt-4",
        messages=[{"role": "user", "content": "Hello"}]
    )
    result = response['choices'][0]['message']['content']
except openai.error.OpenAIError as e:
    result = "⚠️ Sorry, something went wrong. Please try again later."

✅ Handling Specific Exceptions

from openai.error import RateLimitError, Timeout, APIConnectionError

try:
    # Agent logic here
except RateLimitError:
    result = "⏳ Too many requests. Please wait a few seconds."
except Timeout:
    result = "⏱️ This is taking too long. Try again shortly."
except APIConnectionError:
    result = "🔌 Connection error. Please check your network."

🔁 3. Retry Logic with Backoff

Automatically retry failed requests (once or twice) before failing.

import time

def call_with_retry(func, retries=3, delay=2):
    for i in range(retries):
        try:
            return func()
        except Exception as e:
            if i == retries - 1:
                raise e
            time.sleep(delay * (i + 1))

Example use:

response = call_with_retry(lambda: openai.ChatCompletion.create(...))

🧠 Use libraries like tenacity for more advanced retry mechanisms.


🔙 4. Fallback to Local or Simpler Models

If OpenAI API fails, fallback to:

  • A local model via Ollama (mistral, llama3, phi)

  • A simpler response (e.g., “I couldn’t process that. Try a different question.”)

  • Cached answers (frequent questions)

try:
    response = openai.ChatCompletion.create(...)
except:
    response = ollama_chat("What would you like to know?")

📋 5. Display User-Friendly Errors

Do not expose raw errors or stack traces to users.

Instead, show clean messages:

Scenario Message to User
Timeout “⏳ Still thinking… please try again.”
Rate limit hit “⚠️ You’re asking too quickly. Pause a moment.”
Server error “🚧 We’re having trouble. Try again soon.”
Auth error “🔒 Login or API error. Please contact admin.”

🛠️ 6. Streamlit/Gradio Safe Handling Example

Streamlit Chat Example with Failure Catching:

user_input = st.chat_input("Ask me anything")
if user_input:
    try:
        response = agent.run(user_input)
    except Exception as e:
        st.error("🤖 The assistant is temporarily unavailable.")
        log_error(e)
    else:
        st.chat_message("assistant").write(response)

🧾 7. Logging & Alerting for Failures

Always log critical failures and monitor them:

Tool Purpose
Log files Debugging backend failures
Supabase / SQLite Store errors and retry metadata
Slack/Email alerts Get notified on persistent failures
Langfuse Track model errors and performance issues
Sentry, LogRocket Monitor frontend JS/API crashes

🧠 8. Timeout Configuration and Optimization

For OpenAI:

Set a timeout explicitly to avoid agent freezing:

openai.ChatCompletion.create(..., timeout=10)

For LangChain Agents:

Wrap tools with timeout decorators or handlers.

from func_timeout import func_timeout, FunctionTimedOut

try:
    result = func_timeout(15, agent.run, args=(user_input,))
except FunctionTimedOut:
    result = "⏱️ This is taking too long. Please rephrase or try again later."

✅ Summary

Strategy Purpose
Try/Except Handling Catch errors and show fallback
Retry with Backoff Automatically retry failed API calls
Model Fallbacks Switch to local models or cached answers
User-Friendly Messaging Keep UI clear, don’t expose backend errors
Logging + Monitoring Detect failure trends, alert admins
Timeout Configs Prevent app freezes or overload

 

67