Data Validation
APIs often handle untrusted data, making them susceptible to vulnerabilities like JSON injection and insecure deserialization. These flaws arise when input validation is inadequate, allowing attackers to manipulate data payloads to execute arbitrary code, bypass authentication, or corrupt application state. This section explores how to identify and exploit these risks during penetration testing.
JSON Injection¶
What is JSON Injection?¶
JSON injection occurs when an attacker injects malicious JSON content into an API request, exploiting weaknesses in how the server parses or processes the data. This can lead to code execution, data tampering, or bypassing input validation mechanisms.
Common Attack Vectors¶
- Unescaped special characters: Attackers inject JSON comments (
//) or malformed syntax to alter the structure of the payload. - Malformed payloads: Exploiting parsing errors in JSON libraries to execute arbitrary code.
- Context-dependent injection: Leveraging specific API endpoints that dynamically evaluate JSON data (e.g., using
eval()or similar functions).
Example Test Case¶
curl -X POST https://api.example.com/data \
-H "Content-Type: application/json" \
-d '{"name":"Alice//","preferences":"[\"malicious\"]"}'
// as a comment, it might ignore the rest of the payload, potentially bypassing validation logic.
Detection Tools¶
- Burp Suite: Intercept and modify requests to test for JSON injection.
- OWASP ZAP: Use the "JSON Injection" scanner rule to automate detection.
Insecure Deserialization¶
What is Insecure Deserialization?¶
Insecure deserialization occurs when an API deserializes untrusted data (e.g., from user input or external sources) without proper validation. This can allow attackers to execute arbitrary code, manipulate objects, or trigger denial-of-service conditions.
Common Attack Vectors¶
- Remote code execution (RCE): Sending a serialized object that, when deserialized, executes malicious code (e.g., via Java's
ObjectInputStreamor Python'spickle). - Object graph traversal: Exploiting deserialization to access sensitive data or modify internal application state.
- Type confusion: Leveraging deserialization to bypass type checks and access privileged functionality.
Example Test Case (Python)¶
import pickle
malicious_data = pickle.dumps({'__globals__': {'__builtins__': __builtins__}, '__reduce__': (lambda *a: (lambda *b: __import__('os').system('id'))(),)})
# Send malicious_data via an API endpoint that deserializes input
id command, exposing system information.
Detection Tools¶
- Deserialization fuzzers: Tools like
dumb-fuzzerorpickle-toolscan test for deserialization vulnerabilities. - Static analysis: Use tools like
bandit(for Python) to detect unsafe deserialization patterns.
Mitigation Strategies¶
- Input validation: Sanitize all user inputs using whitelists (e.g., allow only alphanumeric characters for names).
- Avoid dynamic evaluation: Never use
eval(),exec(), or similar functions on untrusted data. - Use safe serialization formats: Prefer JSON over binary formats like
pickleorJava's Serializablewhen possible. - Content-type enforcement: Ensure APIs strictly enforce
Content-Typeheaders to prevent malformed payloads. - Library updates: Keep dependencies updated to patch known deserialization vulnerabilities (e.g., Apache Commons Collections in Java).
Key takeaways¶
- Validate all input: Use whitelists and escape special characters to prevent JSON injection.
- Avoid unsafe deserialization: Never deserialize untrusted data; use safe formats and libraries.
- Leverage tools: Use Burp Suite, OWASP ZAP, and fuzzer tools to automate detection of these vulnerabilities.
- Prioritize context: Understand how the API processes data (e.g., dynamic evaluation, serialization) to identify attack surfaces.