TracekitTracekit

How to Read a Flame Graph and Find CPU Bottlenecks

Learn how to read a flame graph, separate CPU samples from trace timing, find wide hotspots, and verify each performance bottleneck.

Terry Osayawe9 min read
How to Read a Flame Graph and Find CPU Bottlenecks

If you need to visualize CPU profiling data, start with a flame graph. It compresses thousands of sampled call stacks into one view.

This guide shows how to read a flame graph and find a performance bottleneck. It also explains a common source of confusion: CPU flame graphs and trace flame graphs answer different questions.

Flame graph quick reference

PartCPU flame graph meaningFirst question to ask
One boxOne stack frame or functionWhich function is this?
Box widthShare of samples containing that frameIs it wide because of itself or its children?
Vertical positionCall-stack depthWhich caller led to this function?
Horizontal positionGrouped stack order, not timeDo not read left to right as a timeline
Top edgeCode running when samples were takenWhich wide top block consumes CPU directly?
ColorUsually visual separationCheck the tool before assigning meaning

The original CPU flame graph guide defines these rules. Its key warning is simple: the horizontal axis does not show elapsed time.

What a CPU flame graph shows

A CPU profiler samples running stacks at a fixed interval. Each sample records the current function and its callers.

The flame graph groups identical stacks and draws one box for each frame. Wider boxes appear in more samples.

That width represents sampled CPU use. It does not represent call count. A wide function may run slowly, run often, or both.

The bottom frame is usually the process entry point. Child calls stack above their parents. Some tools invert this layout into an icicle graph.

The width includes child work

A box has two useful costs:

  • Total or cumulative cost includes the function and every child above it.
  • Self cost includes only samples where that function sits on the top edge.

This difference prevents a common mistake. A wide parent can look expensive even when one child performs almost all work.

For example, renderInvoice() may span 40% of the profile. If compressPdf() fills most of its top edge, optimize compressPdf() first.

Grafana's flame graph documentation also separates self value from cumulative value. Use both when your profiler exposes them.

How to read a flame graph in five steps

1. Confirm what the profile measures

Check whether the graph measures CPU, allocation, memory, lock time, or off-CPU time. The same shape can represent different resources.

Do not use an on-CPU profile to explain a request that only waits for a network response. That wait may use little CPU.

2. Check the capture scope

Record the process, host, release, workload, and capture window. A flame graph without this context can send you toward irrelevant code.

Use a representative slow period. Avoid mixing unrelated workloads unless you can filter them later.

3. Scan the top edge for wide plateaus

Start at the top edge. Wide, flat blocks show functions that were directly running during many samples.

These blocks are strong CPU hotspot candidates. Narrow top blocks may still matter, but they contribute less to this profile.

4. Follow each hotspot down to its caller

Move down from a wide top block. The frames below it show the call path that reached the hotspot.

This path explains ownership. The same library function may appear under several application paths.

5. Validate with a second measurement

Make one focused change. Repeat the same workload and capture another profile.

Compare CPU time, throughput, latency, and the new flame shape. A narrower block alone does not prove that users see an improvement.

Flame graph patterns that matter

PatternLikely meaningNext check
Wide block on the top edgeHigh self CPU useInspect loops, parsing, serialization, or compression
Wide parent with one wide childCost sits mainly in one calleeFollow the child before changing the parent
Wide parent with many narrow childrenWork spreads across helpersCheck total cost and repeated calls
Same function in several towersMany callers reach one functionUse caller or sandwich views
Large unknown or hexadecimal framesSymbols are missingFix symbol resolution and capture again
Mostly scheduler or wait framesWrong profile type for the problemCapture off-CPU data or use tracing
Shape changes only under loadContention or workload sensitivityCompare equal traffic and release conditions

A tall stack is not automatically bad

Height shows depth, not cost. A tall, narrow tower may contribute very little CPU time.

Optimize width before height. Only inspect depth when recursion, wrappers, or excessive indirection creates measurable cost.

Colors are not universal

Traditional CPU flame graphs often use warm colors without semantic meaning. Other tools color frames by package, service, or value.

Read the legend first. Do not assume blue means a database or red means an error.

CPU flame graph versus trace flame graph

Many observability tools use a flame-shaped layout for distributed traces. The layout looks similar, but the data has a different meaning.

QuestionCPU profileDistributed trace
What is measured?Sampled stack activityTimed spans
Typical unitSamples or CPU timeWall-clock duration
Best useFind hot functionsFind slow requests, services, queries, or dependencies
Shows waiting?Not in a normal on-CPU profileYes, when a span covers the wait
Shows every function?Only sampled and resolved framesOnly instrumented operations
Horizontal orderUsually not chronologicalDepends on the trace renderer

Use a trace first when an API request is slow. It can show whether the delay comes from a database, another service, or an external API.

Use a CPU profile when the trace points to application work and the process is CPU-bound. The profile can then identify the hot function.

The latency investigation guide explains this trace-first workflow. The event loop guide covers Node.js saturation signals.

How Tracekit handles flame graphs

Tracekit's current flame graph is trace-based. It builds a hierarchy from OpenTelemetry spans and their parent identifiers.

Each block uses span duration in milliseconds. Use this view to find slow service calls, database spans, and instrumented application steps.

The view supports span search, hover details, and zoom. You can also use the Trace Visualizer with OTLP, Jaeger, or Zipkin trace data.

Tracekit does not currently collect continuous CPU profiles. A Tracekit span flame graph cannot identify an unsampled function inside a wide span.

When a trace isolates a slow code path, add a bounded capture point. Tracekit dynamic logs can capture runtime state without a redeploy.

Keep normal application logs in your logger. Dynamic logs are targeted capture points, not general log ingestion.

How to generate CPU flame graph data

Choose the profiler that matches your runtime and operating system. Follow its current documentation because capture flags can change.

Linux perf

Brendan Gregg describes a three-stage workflow: capture stacks, fold identical stacks, and render the folded data.

perf record -F 99 -a -g -- sleep 60
perf script > out.perf

You can then use the maintained FlameGraph tools to fold and render the data.

Node.js

The official Node.js flame graph guide documents Linux perf with V8 profiling options. It also explains broken symbol labels.

Start with the documented command for your Node.js release. Confirm that JavaScript frames resolve before you trust the result.

Go

Go includes CPU profiling through runtime/pprof and go tool pprof. The official Go profiling guide shows the capture and analysis flow.

Profile the same binary that serves the tested workload. Keep build symbols available for readable function names.

Common mistakes

Reading the x-axis as time

The horizontal order usually groups similar stacks. It does not show what ran first.

Use a trace waterfall when you need chronology. Use a flame graph when you need aggregated stack cost.

Treating width as one slow call

Width comes from the profile's value, such as sample count. One function may be wide because it runs very often.

Check call counts or request traces before you describe one invocation as slow.

Optimizing a library frame without its caller

A shared library can appear under many code paths. Follow the stack downward to find the application path that owns the work.

Ignoring missing symbols

Unknown frames can hide the true hotspot. Fix debug symbols, JIT mappings, frame pointers, or unwinding before optimization.

Profiling the wrong workload

A fast test endpoint cannot explain a production batch job. Capture the workload that produces the actual performance problem.

Confusing CPU time with request latency

A slow database call can dominate wall-clock latency while using little application CPU. Pair the profile with a distributed trace.

If the trace shows repeated database spans, use the N+1 query guide.

A practical bottleneck checklist

Before you change code, answer these questions:

  1. Does the graph measure the resource I need to explain?
  2. Does the capture include the affected release and workload?
  3. Which wide frame reaches the top edge?
  4. Which caller path owns that frame?
  5. Are function names and symbols complete?
  6. Does tracing show CPU work or waiting?
  7. Can I reproduce the hotspot with the same workload?
  8. Did the fix improve an external result, such as latency or throughput?

Frequently asked questions

How do I read a flame graph?

Find wide blocks on the top edge, then follow each block downward through its callers. Width shows relative value, not chronology.

What does a wide block mean?

It means the frame appears in a large share of the measured value. Inspect self and cumulative cost before choosing a fix.

What does the top of a CPU flame graph show?

The top edge shows functions running when samples were taken. Wide plateaus on this edge are direct CPU hotspot candidates.

Does a flame graph show time from left to right?

No, a standard CPU flame graph groups similar stack frames. Use a timeline or trace waterfall for execution order.

Can a flame graph show an N+1 query?

A trace view can show repeated database spans. A CPU flame graph may not show network waiting or query count clearly.

Should I start with tracing or CPU profiling?

Start with tracing for slow requests. Use CPU profiling after traces and CPU metrics point to local application work.

Final rule

Read the data source before you read the shape. Then find a wide top-edge frame, trace its caller path, and verify the fix.

Share this post

Related Posts