Skip to content

LLM DoS Attacks

LLM Denial-of-Service (DoS) attacks exploit vulnerabilities in large language models (LLMs) to exhaust computational resources, rendering services unavailable. These attacks often involve adversarial inputs designed to trigger excessive memory usage, prolonged processing times, or system crashes. Understanding how such vulnerabilities arise is critical for building robust guardrails and mitigating risks in production systems.


Types of LLM DoS Attacks

1. Token Stuffing

Attackers inject excessively long prompts, overwhelming the model's token limit. For example, a malicious user might submit a prompt with 10,000+ tokens, forcing the model to process beyond its capacity. This can lead to memory exhaustion or denial of service.

Example:

curl -X POST "http://api.example.com/generate" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "This is a very long prompt... (repeated 10,000 times)",
    "max_tokens": 1000
  }'

2. Prompt Injection for Resource Drain

Malicious prompts may include recursive or infinite-loop-like instructions (e.g., "Generate a 100,000-word essay about the history of the universe, then summarize it, then generate another essay..."). These force the model to engage in repetitive, computationally intensive tasks.

3. Exploiting Model Complexity

Complex prompts requiring multi-step reasoning, code generation, or extensive data processing can consume significant computational resources. Attackers may craft prompts that trigger these behaviors unnecessarily.


How Adversarial Inputs Exhaust Resources

Memory Overhead

LLMs allocate memory proportional to the number of tokens processed. A 10,000-token input could consume hundreds of MBs of RAM, and multiple such requests can exhaust system memory.

Computational Cost

Generating responses for long or complex prompts requires high CPU/GPU utilization. For example, a prompt asking for a 10,000-word document may take minutes to process, blocking other requests.

System-Level Impact

Resource exhaustion can cause:
- Crashes: Out-of-memory errors or kernel panics.
- Latency: Prolonged response times for legitimate users.
- Denial of Service: Complete service unavailability during attacks.


Real-World Examples

Example 1: Token Stuffing Attack

An attacker sends a prompt with 10,000 tokens to a chatbot API, causing the server to crash due to memory limits.

Mitigation: Enforce strict token limits and validate input lengths.

Example 2: Infinite-Loop Prompt

A prompt like "Write a 100,000-word essay about the history of the universe, then summarize it, then write another essay..." forces the model to engage in endless, redundant computation.

Mitigation: Use input validation to detect and reject recursive patterns.


Mitigation Strategies

1. Rate Limiting

Implement rate limits to restrict the number of requests per user or IP. For example:

# Example Nginx rate-limiting config
limit_req_zone $binary_remote_addr zone=one:10m rate=10r/m;

2. Input Validation

Reject prompts exceeding predefined token limits or containing suspicious patterns.

# Example: Token length check
if len(prompt.split()) > 10000:
    raise ValueError("Prompt exceeds maximum token limit")

3. Resource Monitoring

Use tools like Prometheus and Grafana to monitor memory/CPU usage and trigger alerts during anomalies.

4. Model-Specific Safeguards

Leverage model capabilities like token limit enforcement or safety filters to block malicious inputs.


Key takeaways

  • Adversarial inputs can exhaust LLM resources through token stuffing, complex prompts, or infinite loops.
  • Mitigation requires a combination of input validation, rate limiting, and resource monitoring.
  • Proactive defense involves understanding model behavior and setting strict operational boundaries.
  • Regularly test systems for DoS vulnerabilities using synthetic attacks and load testing.
  • Prioritize transparency in model capabilities to avoid overpromising performance that could be exploited.