Skip to content

Tracepoints

The Linux tracepoint mechanism is a kernel-level infrastructure for capturing structured events during system operation. Unlike dynamic tracing mechanisms like kprobes, tracepoints are static, pre-defined event points embedded in the kernel source code. They provide a standardized way to instrument kernel code without requiring runtime modifications, making them ideal for performance analysis and debugging. Tracepoints are managed via the trace subsystem and are accessible through user-space tools like perf, BCC (Berkeley Packet Filter Compiler), and other tracing frameworks.


Kernel Tracepoint Definitions

Tracepoints are defined in the Linux kernel source code, typically in the kernel/trace directory. Each tracepoint has a unique name and a set of parameters that describe the event. For example, the sched_switch tracepoint captures context-switching events between processes, including details like the previous and next tasks.

Tracepoint definitions are stored in .c files (e.g., trace_sched.c) and are compiled into the kernel. These definitions are also exported via the trace_events directory in the kernel source, enabling user-space tools to access metadata about available tracepoints.


Registration and User-Space Interaction

Tracepoints are registered with the kernel's tracing subsystem during boot. User-space tools interact with them through the /sys/kernel/tracing virtual file system or via the perf utility. For example:

# List all available tracepoints
perf list | grep 'sched'

This command might output entries like sched:sched_switch, indicating a tracepoint named sched_switch in the sched subsystem.

In BCC, tracepoints are accessed using the Tracepoint class in the bcc Python library. For instance:

from bcc import BPF

bpf = BPF(src_file="tracepoint.c")
bpf.attach_tracepoint(tp="sched:sched_switch", fn="handle_switch")

This code attaches a user-defined handler to the sched_switch tracepoint.


Example Usage

To collect data from a tracepoint, you can use perf to record events:

perf record -e sched:sched_switch -a

This command records all sched_switch events system-wide. The collected data can then be analyzed with perf report.

For BCC, a simple script might look like this:

from bcc import BPF
import time

bpf = BPF(src_file="tracepoint.c")
print("Tracing... Hit Ctrl-C to stop.")

def print_event(cpu, data, size):
    event = bpf["events"].event(data)
    print(f"Switched from {event.prev_pid} to {event.next_pid}")

bpf["events"].open_perf_buffer(print_event)
while 1:
    bpf.poll()

This script prints context-switch events in real time.


Key takeaways

  • Tracepoints are static, pre-defined event points in the Linux kernel for structured event collection.
  • They are managed via the trace subsystem and accessible through /sys/kernel/tracing or perf.
  • BCC and other tools use tracepoint metadata to attach handlers and collect data efficiently.
  • Tracepoints offer lower overhead compared to dynamic tracing mechanisms like kprobes.