Skip to content

SystemTap Tracing

SystemTap is a powerful dynamic tracing tool for Linux that allows administrators to monitor kernel events, system calls, and resource contention without modifying the kernel source. By writing scripts in its domain-specific language (DSL), you can gather detailed insights into system behavior, diagnose performance bottlenecks, and optimize resource usage. This guide covers writing SystemTap scripts to trace kernel events, analyze system calls, and detect resource contention.


Getting Started with SystemTap

Before writing scripts, ensure SystemTap is installed and enabled on your system. Most Linux distributions include it via packages like systemtap (Debian/Ubuntu) or systemtap (RHEL/CentOS). Verify installation with:

stap -V
SystemTap requires kernel support for tracing. If not already enabled, rebuild the kernel with CONFIG_SYSTEMTAP enabled. CONFIG_KPROBES may be optional depending on the kernel configuration and the specific tracing features required.

A basic script to list available probes:

probe kernel.function("sys_call_table") {  
    printf("Syscall table probe active\n")  
}
Save this as test.stp and run:
stap test.stp


Monitoring Kernel Events

SystemTap scripts use probe statements to track kernel events. For example, trace process creation:

probe process("init").mark("fork") {  
    printf("Process %d forked\n", pid())  
}
This script logs when the init process forks. To trace scheduling events:
probe sched.sched_switch {  
    printf("CPU %d: switched from %d to %d\n", cpu(), prev_pid(), next_pid())  
}
Compile and run with:
stap -e 'probe sched.sched_switch { ... }'


Analyzing System Calls

System calls are critical for understanding application-kernel interactions. Trace read() calls:

probe syscall.read {  
    printf("Read %d bytes from fd %d by process %d\n", user_stack(), fd(), pid())  
}
Filter by process name:
probe syscall.write where pid() == execname("nginx") {  
    printf("Nginx wrote %d bytes\n", user_stack())  
}
Use user_stack() to capture data sizes and fd() to identify file descriptors.


Detecting Resource Contention

SystemTap can identify CPU, memory, or I/O bottlenecks. For CPU contention:

probe sched.sched_wakeup {  
    printf("CPU %d: Waking up task %d\n", cpu(), tid())  
}
For lock contention:
probe lock.contention {  
    printf("Lock contention detected by %d\n", tid())  
}
Monitor I/O waits:
probe block.io_wait {  
    printf("I/O wait event on device %d\n", dev())  
}


Best Practices

  • Use filters: Limit probes to specific processes or events to reduce noise.
  • Avoid over-probing: Excessive probes can degrade performance.
  • Test in safe environments: Validate scripts on non-critical systems first.
  • Leverage logging: Use --log to debug script syntax or runtime errors.
  • Profile selectively: Focus on high-impact events (e.g., syscalls, scheduling) rather than broad tracing.

Key takeaways

  • SystemTap scripts use probe statements to trace kernel events, syscalls, and resource contention.
  • Filter probes by process name, PID, or event type to isolate critical metrics.
  • Monitor syscalls to understand application-kernel interactions and data flow.
  • Detect CPU, memory, and I/O bottlenecks with targeted probes for scheduling, locks, and block I/O.
  • Prioritize efficiency by avoiding over-probing and testing scripts in controlled environments.