Curated summary
How we wrote a Python profiler
Datadog built a Python continuous profiler because Python lacked Java-style, always-on production profiling tools. The post argues that deterministic profilers such as cProfile impose too much overhead for continuous use, while statistical profiling can provide representative performance data with minimal disruption. Datadog’s profiler addresses this through modular collectors, recording, scheduling, and data export.
Profiling Versus Tracing
- Profiling measures resource consumption such as CPU time and memory allocation to reveal performance problems.
- Tracing records individual operations—such as SQL queries or HTTP requests—within a request timeline.
- Tracing explains request latency, but profiling provides deeper insight into code-level execution and operating-system resource usage.
Limitations of Deterministic Python Profilers
- Python’s
cProfile, available since CPython 2.5, records every function call and the time spent in each call. - It can provide a complete execution flow, but its usefulness depends heavily on code structure:
- A program built around a few large functions produces little actionable detail.
- A program containing thousands of functions can incur two- or three-times runtime overhead.
- This overhead makes deterministic profiling unsuitable for always-on production environments.
Why Profile in Production?
- Optimizing without profiling is essentially guessing; real workloads often differ from development environments.
- Production systems vary from developer machines in hardware, concurrency, input data, and workload behavior.
- Continuous profiling captures how an application actually consumes resources under authentic conditions.
- These requirements lead to statistical rather than deterministic profiling.
Statistical Profiling in Python
- Statistical profilers sample program activity periodically instead of recording every function call.
- Individual short-lived calls may be missed, but repeated sampling over hours produces a reliable picture of resource consumption.
- Lower overhead allows the application to run closer to its normal, unprofiled behavior.
- Datadog evaluated numerous open-source Python profilers but found limitations involving platform support, collected data, or presentation-focused designs.
- The team therefore developed its own statistical profiler, incorporating ideas from the tools it studied.
Datadog Python Profiler Design
The profiler was designed around three constraints:
- Keep runtime overhead as low as possible.
- Make deployment simple.
- Support common operating systems and environments.
Its architecture, inspired by the JDK Flight Recorder, consists of:
- Collectors: Gather data such as CPU usage and memory allocation.
- Recorder: Stores events produced by collectors.
- Exporter: Sends profiling data outside the application.
- Scheduler: Invokes components at appropriate intervals, such as exporting data every 60 seconds.
- Profiler: Provides the high-level interface used by applications.
Stack Collection
- The stack collector is the primary built-in collector.
- It wakes 100 times per second and captures the execution stack of every Python thread.
- For each thread, it gathers information including:
- The currently executing function
- CPU time consumed
- Exceptions being handled
- The collector monitors the time required to inspect the application so it can control and limit its own CPU overhead.
A statistical profiler with low overhead is the appropriate foundation for continuous production profiling, giving teams evidence about real application behavior without substantially changing that behavior.
Related reading
Continue with another curated summary.
Our journey taking Kubernetes state metrics to the next level
Read originalHow we optimized our Akka application using Datadog’s Continuous Profiler
Read originalSecure publication of Datadog Agent integrations with TUF and in-toto
Read originalHackathon project: Viewing Datadog metrics in Minecraft
Read original