Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start by separating a true non-terminating loop from any other event-loop stall. A synchronous JavaScript loop that never yields monopolizes Node.js’s single execution thread, so callbacks, timers and incoming requests cannot run. Sustained CPU use supports that hypothesis; low CPU with requests waiting points more toward blocked or slow asynchronous work. Capture evidence safely, profile the hot path, inspect the termination logic, then mitigate the affected workload and roll out a bounded fix.

What an “infinite loop” looks like in production

Node.js executes JavaScript on a single event-loop thread. As Clinic.js puts it, “The event loop is single-threaded: only one operation is processed at a time.” A synchronous loop that does not terminate or yield prevents that thread from processing other callbacks. The same symptom can come from a loop that is merely extremely long, runaway recursion, repeated synchronous work per request, or traversal of unexpectedly large input.

  • CPU high and latency rising: consistent with a hot synchronous path, although native work and garbage collection can also consume CPU.
  • CPU low while requests wait: investigate slow or unavailable asynchronous dependencies, connection pools and timers before assuming a loop.
  • One route or job affected: correlate the incident with the input, tenant, queue message or deployment that triggers it.
  • All routes affected: a process-wide event-loop blockage is more likely than a single downstream call.

A CPU profile samples activity during a time window; it does not, by itself, prove that code can never finish. Confirm non-termination by reading the source and reproducing with the same class of input.

First response: establish scope without destroying evidence

  1. Record the boundary of the incident. Note the process or instance, start time, affected routes or jobs, request identifiers, input characteristics, recent deployments, configuration changes and whether replicas are also affected.
  2. Check telemetry before attaching tools. Compare CPU, event-loop delay, request latency, error rate, memory, garbage-collection activity and dependency timings. Use the service’s existing dashboards and incident procedure.
  3. Decide whether capture is safe. A busy process may have little headroom. Prefer a canary or representative reproduction when possible. Follow your platform’s policy for diagnostic files because reports can contain request-adjacent data, paths, environment details and heap information.
  4. Preserve a comparison point. Save the deployed Node.js version, operating-system image, package lockfile and the exact build identifier. Profiling tools and report options vary by runtime and operating system.

Capture a Node.js diagnostic report

Node.js diagnostic reports are intended for development, test and production problem determination. A report can include JavaScript and native stacks, heap information, libuv handles, platform details and resource data. Those fields let you distinguish a JavaScript hot path from a process waiting on handles or spending time in native code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Programmatic capture

If your service already exposes an authenticated, restricted diagnostic path, write a report only under an incident-approved control. The following minimal example uses the built-in report API; verify the API and options against the exact Node.js version deployed:

import process from 'node:process';

export function writeIncidentReport() {
  const file = `/var/log/my-service/report-${Date.now()}.json`;
  process.report.writeReport(file);
  return file;
}

Do not add an unauthenticated HTTP endpoint that anyone can call. Restrict who can trigger capture, protect the destination, record the report filename in incident notes and transfer it through your approved channel. If the process is already failing, use the report triggers configured by your runtime or process manager rather than inventing a signal or shell command for an unknown deployment.

What to inspect in the report

  • JavaScript stack: look for repeated application frames, the current request handler and the function doing synchronous work.
  • Native stack and libuv handles: identify whether the process is blocked in native activity or simply has pending handles.
  • Heap and resource data: check whether allocation pressure or an unexpectedly large input is amplifying the incident.
  • Platform and runtime metadata: use it to match the capture to the deployed Node.js build and operating system.

A report is a snapshot, not a timeline. Take only the captures your incident procedure allows and correlate each one with CPU and latency graphs.

Profile CPU to find the hot function

Use a CPU sampler when the process is consuming CPU or when a safe reproduction is available. Clinic.js Doctor is designed to classify performance symptoms; Clinic.js Flame collects CPU samples and renders a flamegraph. Their documentation also supports collection-only workflows so data can be collected in one environment and visualized elsewhere. Visual Studio Code can open JavaScript .cpuprofile files and show CPU flame views.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Live process versus reproduction

Approach Best use Trade-off
Live-process capture Intermittent production-only input or a currently stuck worker Operational risk and data-handling concerns; tool compatibility must be checked first
Representative reproduction Repeatable input, queue message or request in a staging environment Safer and easier to inspect, but may miss production-only data or configuration
Diagnostic report Broad context: stacks, heap, handles and platform details Snapshot rather than focused CPU attribution
CPU flamegraph Focused view of where sampled CPU time is spent Sampling window can miss intermittent work and does not prove non-termination

Wide blocks in a flamegraph are candidates for investigation because they account for more sampled time. Repeated application frames can point to a loop or repeated computation. Confirm the finding in source; do not treat a wide frame as a mathematical proof that the loop cannot end.

Collect without adding more load

Avoid high-volume synchronous logging while the event loop is already under pressure. Prefer a bounded number of samples, correlation identifiers and an input fingerprint. Confirm the profiler’s supported Node.js and operating-system versions before attaching it to production. The reviewed tool documentation does not establish a universal overhead percentage, so use your service’s own change and incident controls.

Trace the hot stack back to the bug

Check loop progress and termination

function consume(items) {
  let index = 0;
  while (index < items.length) {
    handle(items[index]);
    index += 1; // verify this executes on every path
  }
}

Inspect every branch for a missing or conditional mutation of the control variable. Check integer overflow, a condition that can never become false, mutation of the collection being traversed, and exceptions that are caught and retried without a limit.

Look for unbounded retries and recursion

async function fetchUntilSuccess(load) {
  for (;;) {
    try {
      return await load();
    } catch (error) {
      await delay(100);
    }
  }
}

This code yields between attempts, so it may not peg the CPU, but it can still create an outage by waiting forever. Add an attempt or deadline budget and surface the final error. Recursive code needs the same treatment: establish a maximum depth or convert it to an iterative traversal with an explicit work limit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check input size and repeated per-request work

Parsing, regular-expression evaluation, graph traversal and synchronous compression can be finite yet exceed your latency budget for an unexpectedly large payload. Measure input size and iteration count, then enforce limits before entering the expensive path. If the work is legitimately CPU-heavy, move it away from the request event loop using an architecture appropriate to your service, such as a worker pool or separate job process.

Mitigate the incident before deploying the fix

  • Protect capacity: shed or isolate the affected route, tenant, queue partition or job class according to your incident playbook.
  • Roll back a suspect change: use the deployment system’s approved rollback, preserving the build and evidence that identified the regression.
  • Replace an unhealthy process: restart or reschedule only through the controls of your runtime, container platform or process manager. A restart removes the symptom and evidence, so capture what is safe first.
  • Bound the work: add iteration, recursion, retry, payload-size and wall-clock limits with explicit failure behavior.
  • Move CPU-heavy work: keep request handlers responsive by routing suitable computations to workers or asynchronous jobs.

These are environment-dependent operational actions; no single signal, shell command or container procedure is correct for every deployment.

Verify the code fix

  1. Build a reproduction from the same input class, configuration and dependency behavior that triggered the incident.
  2. Assert termination: maximum iterations, maximum recursion depth, retry count and deadline should all be testable values.
  3. Run a CPU profile before and after the change. The hot frame should disappear or become bounded, and unrelated paths should not become dominant.
  4. Exercise cancellation, timeout, malformed input, empty collections and the largest accepted payload.
  5. Release gradually, watching event-loop delay, CPU, latency, errors and queue age. Keep the old build available for a controlled rollback.

Common failure modes and fixes

Symptom Likely cause Next action
CPU near saturation; repeated frames in a profile Synchronous loop, recursion or repeated computation Inspect the termination variable and bound the work
CPU low; requests wait on a dependency Slow I/O, exhausted pool or unavailable service Review dependency timings, timeouts and pool metrics
Only large inputs trigger the incident Finite algorithm with unbounded input cost Enforce size limits and profile a representative large case
Report has no obvious JavaScript hot frame Wrong capture window, native work or insufficient samples Correlate timestamps, capture again if safe and profile a reproduction
Adding logs makes latency worse Synchronous or high-volume logging on a blocked loop Remove noisy logging; use bounded, asynchronous telemetry
Fix works locally but not in production Different Node.js version, OS, flags, input or configuration Match the production build and replay production-shaped input
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you need a clean screenshot of an incident dashboard, reproduction page or runbook artifact while documenting the investigation, ScreenshotNeo provides a single HTTP request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for options such as full-page capture, selector targeting, waits, custom headers, cookies, blocking, PDF settings and asynchronous jobs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free for ScreenshotNeo.

Further reading

Clinic.js Doctor and Flame documentation cover symptom classification and flamegraph collection. Node.js’s Diagnostic report documentation explains report contents and version-dependent options. Microsoft’s “Performance profiling JavaScript” documentation explains viewing .cpuprofile data in Visual Studio Code. A book titled Node.js High Performance is relevant background reading, but the available older edition does not establish current retail availability.

Frequently Asked Questions

Can an asynchronous loop still cause an outage?

Yes. A loop that awaits between attempts may leave requests or jobs waiting indefinitely without pegging CPU. Give retries a finite count or deadline and report the terminal failure.

Should I restart the process immediately?

Follow your incident procedure. Restarting can restore capacity, but capture an approved report or profile first when safe because replacement destroys the live evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What if profiling changes the symptom?

Treat that as a clue, not proof. Compare a safe reproduction, correlate timestamps with telemetry and use the least intrusive supported collection mode for the deployed runtime.

How do I prove the fix cannot loop forever?

Make progress and limits explicit, then test maximum iterations, recursion depth, retries, deadlines and malformed or oversized inputs. A bounded profile and representative production-shaped reproduction provide additional confidence.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.