# Node.js Response Time Monitoring by Route

> Measure Node.js response time by route, read p95 latency with traces, and separate event loop delays from database and API waits.

- Published: 2026-10-06T00:00:00.000Z
- Updated: 2026-10-06T00:00:00.000Z
- Author: Terry Osayawe
- Tags: nodejs, performance-monitoring, distributed-tracing, alerting, production-debugging
- Canonical: https://tracekit.dev/blog/nodejs-response-time-monitoring

**Node.js response time monitoring** starts with one duration distribution for each important route. Record the request method, stable route template, status, and total server duration. Compare p50 and p95 with request volume and errors. Then open a slow trace to find where the time went.

A single average hides slow requests. A service-wide number also hides the route that users need. This guide shows a small measurement example, the metrics to keep, and a response-time investigation that works in production.

## What should Node.js response time monitoring measure?

Measure elapsed time from the server receiving a request to the server completing its response. OpenTelemetry names the recommended HTTP server histogram [`http.server.request.duration`](https://opentelemetry.io/docs/specs/semconv/http/http-metrics/) and uses seconds as its unit. Instrumentation may still emit an older metric name, so check the version and output you run.

Keep these dimensions with each observation:

| Dimension | Why it matters | Safe example |
| --- | --- | --- |
| Method | Separates reads from writes | `GET` |
| Route template | Groups requests without a unique series for every ID | `/orders/:id` |
| Status code or class | Separates fast failures from successful work | `200`, `5xx` |
| Service and environment | Separates applications and test traffic | `checkout-api`, `production` |

Use a route template, not a full URL or customer ID. OpenTelemetry requires [`http.route` to have low cardinality](https://opentelemetry.io/docs/specs/semconv/http/http-spans/). Full paths can create too many metric series and expose sensitive values.

Also record request count and server errors. A faster p95 can be bad news if slow requests now fail before the work finishes. Compare latency, traffic, and errors for the same route and time window.

## Start with one Express measurement

This small probe measures one route before you add a full telemetry pipeline. It uses a monotonic clock and a fixed route name. Replace `console.log` with your metrics exporter for sustained monitoring.

```js
const route = '/orders/:id'

app.get(route, (req, res, next) => {
  const start = process.hrtime.bigint()

  res.once('finish', () => {
    const durationMs = Number(process.hrtime.bigint() - start) / 1e6

    console.log(JSON.stringify({
      method: req.method,
      route,
      status: res.statusCode,
      duration_ms: durationMs,
    }))
  })

  next()
}, loadOrder)
```

The [`finish` event](https://nodejs.org/api/http.html) means Node.js handed the last response bytes to the operating system. It does not prove the client received them. Add a `close` handler if you also need to count responses that end early. Keep client-side timing when network time matters.

Do not build a long-term dashboard from raw log lines alone. Export request duration as a histogram. Histograms keep the distribution needed for p95 and p99. Use a counter for request and error rates. If you already have HTTP instrumentation, check its output before adding a second duration metric.

## Read p50, p95, and p99 together

Percentiles answer different questions:

- **p50** shows the typical request.
- **p95** shows a slow tail that affects roughly one request in twenty.
- **p99** helps inspect rarer delays when enough requests exist.

Suppose `GET /orders/:id` has p50 of 90 ms and p95 of 820 ms. The typical request looks fine, but some users wait much longer. A trace from the slow tail can show whether the extra time sits in SQL, an outbound API, or the application itself.

Compare percentiles only with their request count and time window. A route with five calls in a window cannot support a stable p99 decision. Also split traffic by route, status, and release before you treat one service-wide percentile as a diagnosis.

## Find the slow part with a trace

A response-time metric says **when** a route slowed down. A trace can show **where** the request spent time. Use this order:

1. Select the route and time window with higher p95.
2. Check volume and error rate in that same window.
3. Open a slow server span from the route.
4. Compare database and outbound HTTP child spans.
5. Check repeated calls, retries, and gaps without child spans.
6. Compare the first bad interval with a recent release.

Do not add every child span duration and call the sum the response time. Child spans can overlap, and the server span can include work without child spans. Read the timeline and dependent path instead.

The [OpenTelemetry JavaScript instrumentation guide](https://opentelemetry.io/docs/languages/js/libraries/) documents HTTP and Express instrumentation. It also explains why instrumentation must load before the libraries it observes. Verify one request with a database call and an outbound call before you trust the trace.

For a broader setup, use the [Node.js application performance monitoring guide](/blog/node-js-monitoring-complete-guide-for-production-apps). For repeated SQL calls, follow the [N+1 query investigation](/blog/how-to-detect-and-fix-n1-query-problems-complete-guide).

## Check Node.js runtime delay before blaming a dependency

A slow server span with no slow child span can point to work inside the Node.js process. Check event loop delay, event loop utilization, CPU, and memory for the same interval. A long synchronous task can delay several routes at once.

Node.js provides [`monitorEventLoopDelay()` and `eventLoopUtilization()`](https://nodejs.org/api/perf_hooks.html). Event loop delay samples use nanoseconds, so convert them before you compare them with milliseconds of request time. See the [event loop monitoring guide](/blog/nodejs-monitoreventloopdelay-eventlooputilization) for a practical measurement example.

A wide server span alone does not prove CPU use. If the trace points to local work, use a CPU profile to find the function. Tracekit does not collect continuous CPU profiles. Its [distributed tracing feature](/features/distributed-tracing) helps locate the slow request path.

## Set an alert that leads to an action

A useful response-time alert combines a route, sustained latency, and enough traffic to make the signal credible. Check errors beside latency. Set the threshold from the route's normal range and user expectation, not from one universal number.

Include these details in the alert:

- affected route and service;
- p95 value, baseline, and time window;
- request count and error rate;
- a link to representative slow traces;
- recent release, when available.

Use the [Tracekit alert documentation](/docs/alerts) to set up alert rules. A slow request trace can lead to a database span, an outbound request, or a local runtime gap. If one runtime value remains unknown, use [dynamic logs](/docs/code-monitoring) to capture targeted runtime state at a bounded capture point. Keep normal application logs in your logger.

## Where Tracekit fits

The current [Tracekit Node.js integration](/docs/languages/nodejs) creates incoming HTTP server spans with route, status, and duration. It also documents outgoing HTTP and supported database instrumentation. Check those spans with a known request after setup.

Start with one route that users care about. Confirm its duration distribution, errors, and trace shape. Then add another route. This gives the team a reliable path from a slow response to the work that caused it.
