📊 Lesson 8.2: Logging, Monitoring, and Failover
🎯 Lesson Objective
By the end of this lesson, you will:
-
Understand the importance of logging and monitoring for tracking and diagnosing agent behavior
-
Learn how to log user inputs, outputs, tool usage, and errors
-
Set up real-time monitoring and alerting for performance and safety
-
Implement failover strategies so your AI agent continues working smoothly under failure or load
📋 1. Why Logging and Monitoring Matter
As your AI agents scale and handle critical business functions, it’s essential to:
-
Detect failures early (timeouts, API issues, prompt bugs)
-
Audit interactions for safety, accuracy, or user concerns
-
Improve performance by analyzing slow chains, large responses, or token spikes
-
Monitor for abuse, misuse, or prompt injection attempts
Without logging and monitoring:
❌ You won’t know why the agent failed, what the user experienced, or how to fix it.
🧾 2. What to Log in AI Agent Systems
| Category | Example Fields |
|---|---|
| 🧠 Prompt Logs | Input text, final prompt, system instructions |
| 💬 Chat History | Timestamps, user messages, AI replies |
| 🧰 Tool Usage | Tool name, input params, output, status |
| ⚠️ Errors | Exception type, traceback, failed component |
| ⏱️ Performance | Response time, latency, token usage |
| 🧑 User Session Data | User ID, role, browser, session length |
🔧 Example: Python Log Entry (JSON Format)
{
"timestamp": "2025-07-01T15:22:11",
"user_id": "ronny123",
"input": "Summarize this PDF",
"output": "This document explains...",
"tools_used": ["PDFLoader", "summarizer_chain"],
"latency_ms": 1540,
"token_usage": 820,
"error": null
}
🛠️ 3. Logging Tools and Frameworks
📁 Local or Lightweight Options
-
Python
loggingmodule (writes to file or console) -
SQLite/PostgreSQL table (good for audits)
-
Streamlit
session_statefor in-app logging
☁️ Cloud-Ready Solutions
| Tool | Use Case |
|---|---|
| Langfuse | Logs prompts, responses, chains, errors, and latency metrics visually |
| OpenPanel | Visual debugger and log viewer for LLM agents |
| Supabase | Realtime database for logging user-agent interactions |
| Logtail / LogRocket | Logs frontend + backend events in cloud apps |
| Sentry | Full error tracking and alerts |
🧠 4. Monitoring Agent Performance and Health
| Metric | Description |
|---|---|
| Request volume | How many chats, tools, or calls per hour/day |
| Token usage | Track if users overconsume LLM resources |
| Error rate | Spike in failed tool calls or API errors |
| Latency | Track slow prompts or overloaded chains |
| User satisfaction | Track feedback buttons or thumbs up/down |
| Uptime | Monitor whether LLMs, vector DBs, and APIs are online |
✅ Tools to Help:
-
Langfuse dashboard → Full trace explorer
-
Prometheus + Grafana → Self-hosted performance dashboards
-
Healthcheck URLs → Add
/healthendpoints to agents
🚨 5. Real-Time Alerts and Notifications
Set up alerts for:
-
⚠️ >10 failed prompts in 5 minutes
-
⏳ Agent taking over 5 seconds for >50% requests
-
🚫 Tool returning “null” or “Invalid input” repeatedly
-
🧨 Prompt injection attempt detected
Send alerts to:
-
Slack
-
Email
-
SMS
-
Telegram bots
-
Logging dashboards (Langfuse / Datadog)
🔁 6. Failover Mechanisms
Failover is the ability to automatically switch to a backup system when something fails.
🤖 Examples of Failover Strategies:
| Situation | Failover Strategy |
|---|---|
| OpenAI API fails | Switch to local Ollama model |
| PDF loader breaks | Use fallback file summarizer |
| User exceeds limits | Return cached answer or suggest retry |
| HuggingFace model down | Use another hosted API or simpler LLM |
🧪 Example in Code:
try:
response = openai.ChatCompletion.create(...)
except openai.error.OpenAIError:
response = run_local_agent(prompt)
🗃️ 7. Centralized Log and Replay System
Create a dashboard where you can:
-
Search all chats or prompts
-
Replay user-agent conversations
-
Inspect tool calls and failures
-
Tag conversations for follow-up
Use:
-
Langfuse with LangChain agent integrations
-
Custom React dashboard backed by Supabase
-
Django admin panel with logging table
✅ Summary
| Feature | Benefit |
|---|---|
| Logging | Know what happened, when, and why |
| Monitoring | Detect spikes, failures, or latency issues |
| Real-time Alerts | Act fast on critical failures |
| Failover | Ensure continuity even when models or tools fail |
| Audit & Replay | Analyze user behavior, troubleshoot, or improve UX |
77
