Retry Logic & Errors
LangGraph workflows often rely on external tools (e.g., APIs, databases) that can fail due to network issues, rate limits, or invalid inputs. Robust error handling and retry logic are critical to ensure reliability. This section covers strategies for managing tool call failures, including retries, timeouts, and recovery patterns.
Core Concepts¶
Retry Strategies¶
Retrying failed tool calls is a common pattern to handle transient errors (e.g., network glitches). LangGraph integrates with libraries like tenacity to implement retries with exponential backoff and jitter.
from tenacity import retry, stop_after_attempt, wait_exponential_jitter
@retry(stop=stop_after_attempt(3), wait=wait_exponential_jitter(0.1, 0.5))
def call_tool_with_retry():
result = tool_call()
if result.status == "failed":
raise Exception("Tool call failed")
return result
Key parameters:
- stop_after_attempt: Maximum retry attempts.
- wait_exponential_jitter: Randomized delays to avoid thundering herd problems.
Timeouts and Circuit Breakers¶
Timeouts prevent workflows from hanging indefinitely. Circuit breakers stop retrying after a certain number of failures, avoiding infinite loops.
from langgraph import tool_call
import asyncio
async def safe_tool_call():
try:
result = await tool_call(timeout=5.0) # 5-second timeout
return result
except asyncio.TimeoutError:
logger.error("Tool call timed out")
return {"status": "timeout", "message": "Operation exceeded allowed duration"}
Circuit breaker example (stateful):
failure_count = 0
max_failures = 3
cooldown_period = 60 # seconds
def check_circuit_breaker():
global failure_count
if failure_count >= max_failures:
if time.time() - last_failure_time > cooldown_period:
failure_count = 0
return False
return True
return False
Failure Recovery Patterns¶
Fallback Responses¶
Use fallback logic to provide default responses when tool calls fail:
def handle_tool_failure(error):
if "rate limit" in error:
return "Error: API rate limit exceeded. Please try again later."
elif "network" in error:
return "Error: Temporary network issue. Retrying..."
else:
return "Error: Unknown failure. Please check the tool configuration."
State Persistence¶
Persist workflow state to retry from a known checkpoint. For example, save intermediate results to a database or file system.
def save_checkpoint(state):
with open("workflow_checkpoint.json", "w") as f:
json.dump(state, f)
def load_checkpoint():
try:
with open("workflow_checkpoint.json", "r") as f:
return json.load(f)
except FileNotFoundError:
return {}
Monitoring and Logging¶
Structured Logging¶
Log detailed error metadata for debugging:
import logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
def log_tool_error(error, tool_name):
logger.error(f"[Tool: {tool_name}] Error: {error}")
logger.debug("Stack trace: %s", error.__traceback__)
Metrics Collection¶
Track retry counts and failure types using Prometheus or custom metrics:
from prometheus_client import Counter
tool_failure_counter = Counter("tool_failure_total", "Total tool failures by type")
def record_tool_failure(error_type):
tool_failure_counter.labels(error_type).inc()
Diagram: Retry and Recovery Flow¶
graph TD
A[Tool Call] --> B{Success?}
B -->|Yes| C[Return Result]
B -->|No| D[Check Retry Policy]
D -->|Retry| E[Retry Tool Call]
D -->|No Retry| F[Log Failure]
F --> G[Trigger Recovery (Fallback/Checkpoint)]
Key takeaways¶
- Use exponential backoff with jitter for retries to avoid cascading failures.
- Combine timeouts with circuit breakers to prevent infinite loops.
- Implement fallback responses and state persistence for graceful degradation.
- Log structured error metadata and track metrics for post-mortem analysis.
- Prioritize idempotency in tool calls to safely retry without side effects.