📊 Lesson 8.2: Logging, Monitoring, and Failover


🎯 Lesson Objective

By the end of this lesson, you will:

  • Understand the importance of logging and monitoring for tracking and diagnosing agent behavior

  • Learn how to log user inputs, outputs, tool usage, and errors

  • Set up real-time monitoring and alerting for performance and safety

  • Implement failover strategies so your AI agent continues working smoothly under failure or load


📋 1. Why Logging and Monitoring Matter

As your AI agents scale and handle critical business functions, it’s essential to:

  • Detect failures early (timeouts, API issues, prompt bugs)

  • Audit interactions for safety, accuracy, or user concerns

  • Improve performance by analyzing slow chains, large responses, or token spikes

  • Monitor for abuse, misuse, or prompt injection attempts

Without logging and monitoring:

❌ You won’t know why the agent failed, what the user experienced, or how to fix it.


🧾 2. What to Log in AI Agent Systems

Category Example Fields
🧠 Prompt Logs Input text, final prompt, system instructions
💬 Chat History Timestamps, user messages, AI replies
🧰 Tool Usage Tool name, input params, output, status
⚠️ Errors Exception type, traceback, failed component
⏱️ Performance Response time, latency, token usage
🧑 User Session Data User ID, role, browser, session length

🔧 Example: Python Log Entry (JSON Format)

{
  "timestamp": "2025-07-01T15:22:11",
  "user_id": "ronny123",
  "input": "Summarize this PDF",
  "output": "This document explains...",
  "tools_used": ["PDFLoader", "summarizer_chain"],
  "latency_ms": 1540,
  "token_usage": 820,
  "error": null
}

🛠️ 3. Logging Tools and Frameworks

📁 Local or Lightweight Options

  • Python logging module (writes to file or console)

  • SQLite/PostgreSQL table (good for audits)

  • Streamlit session_state for in-app logging

☁️ Cloud-Ready Solutions

Tool Use Case
Langfuse Logs prompts, responses, chains, errors, and latency metrics visually
OpenPanel Visual debugger and log viewer for LLM agents
Supabase Realtime database for logging user-agent interactions
Logtail / LogRocket Logs frontend + backend events in cloud apps
Sentry Full error tracking and alerts

🧠 4. Monitoring Agent Performance and Health

Metric Description
Request volume How many chats, tools, or calls per hour/day
Token usage Track if users overconsume LLM resources
Error rate Spike in failed tool calls or API errors
Latency Track slow prompts or overloaded chains
User satisfaction Track feedback buttons or thumbs up/down
Uptime Monitor whether LLMs, vector DBs, and APIs are online

✅ Tools to Help:

  • Langfuse dashboard → Full trace explorer

  • Prometheus + Grafana → Self-hosted performance dashboards

  • Healthcheck URLs → Add /health endpoints to agents


🚨 5. Real-Time Alerts and Notifications

Set up alerts for:

  • ⚠️ >10 failed prompts in 5 minutes

  • ⏳ Agent taking over 5 seconds for >50% requests

  • 🚫 Tool returning “null” or “Invalid input” repeatedly

  • 🧨 Prompt injection attempt detected

Send alerts to:

  • Slack

  • Email

  • SMS

  • Telegram bots

  • Logging dashboards (Langfuse / Datadog)


🔁 6. Failover Mechanisms

Failover is the ability to automatically switch to a backup system when something fails.

🤖 Examples of Failover Strategies:

Situation Failover Strategy
OpenAI API fails Switch to local Ollama model
PDF loader breaks Use fallback file summarizer
User exceeds limits Return cached answer or suggest retry
HuggingFace model down Use another hosted API or simpler LLM

🧪 Example in Code:

try:
    response = openai.ChatCompletion.create(...)
except openai.error.OpenAIError:
    response = run_local_agent(prompt)

🗃️ 7. Centralized Log and Replay System

Create a dashboard where you can:

  • Search all chats or prompts

  • Replay user-agent conversations

  • Inspect tool calls and failures

  • Tag conversations for follow-up

Use:

  • Langfuse with LangChain agent integrations

  • Custom React dashboard backed by Supabase

  • Django admin panel with logging table


✅ Summary

Feature Benefit
Logging Know what happened, when, and why
Monitoring Detect spikes, failures, or latency issues
Real-time Alerts Act fast on critical failures
Failover Ensure continuity even when models or tools fail
Audit & Replay Analyze user behavior, troubleshoot, or improve UX

 

77