Recommended Free Tools
Start by identifying the termination reason, not by guessing a memory value. In AKS, Reason: OOMKilled means the container exceeded its memory cgroup limit. Exit code 137 by itself only means the process received SIGKILL; kubelet eviction, a failed liveness probe, or manual deletion can produce the same code. Inspect the pod’s last state, events, memory trends, node pressure, and Chrome logs before changing limits.
1. Establish what actually crashed
Run these commands for the affected pod:
kubectl describe pod <pod-name> -n <namespace>
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.status.containerStatuses[*].lastState.terminated}'
kubectl get events -n <namespace> --sort-by=.lastTimestamp
kubectl top pod <pod-name> -n <namespace>
kubectl top node
kubectl top requires the metrics pipeline to be available. In describe, look at Last State, Reason, Exit Code, restart count, and the Events section.
Interpret the termination fields
- OOMKilled with exit code 137: the container crossed its memory limit.
- Error with exit code 137: the process was killed, but the cause is not proven. Check eviction events, probe failures, and operator actions.
- Page crashed while the pod remains healthy: Chromium or a renderer may have failed without the container being OOM-killed.
- Repeated restarts during node pressure: investigate node capacity and eviction signals, not only the individual pod.
Capture container-level memory from your monitoring system or cgroup files and compare the peak with the configured limit. A stable high plateau suggests demand or sizing; memory that keeps rising and never returns toward baseline suggests retained data, browser processes, or a leak.
2. Compare requests, limits, and node capacity
Inspect the deployment and the node that hosts it:
kubectl get deployment <deployment> -n <namespace> -o yaml
kubectl describe node <node-name>
kubectl get pod <pod-name> -n <namespace> -o wide
Memory requests influence scheduling; memory limits define the cgroup ceiling. A pod can be scheduled successfully yet still be killed when Chrome and Node.js exceed that ceiling. Conversely, raising a limit does not help if the node cannot supply the requested capacity or if several replicas create node-wide pressure.
#1 Best Overall
Check for unsuitable requests or limits, overcommitted nodes, insufficient node resources, and controls that do not match the workload. Increasing the limit can be a short-term restart mitigation, but treat it as a hypothesis to validate with representative PDF jobs and peak measurements.
3. Measure the PDF workload and concurrency
Reproduce the exact URLs, page lengths, fonts, images, and JavaScript used by failing jobs. Record:
- Peak container memory, including Node.js and every Chrome child process.
- Number of simultaneous browsers, contexts, pages, and PDF jobs.
- Job duration and whether memory falls after a job completes.
- Output size and the point at which failures begin.
There is no universal AKS memory number for Puppeteer PDFs. Page dimensions and PDF bytes alone cannot predict Chromium’s peak. One historical Puppeteer issue reported Chromium using more than 1.4 GB for a particular PDF/screenshot workload; that observation is not a sizing target.
Reduce peak demand while measuring
- Put a bounded queue in front of PDF generation and lower simultaneous jobs until peaks are understood.
- Reuse a browser only when you can prove contexts and pages are closed; otherwise test one browser per job and compare the cost.
- Always close pages, contexts, and browsers in
finallyblocks. - Do not retain full HTML, screenshots, or PDF buffers in an in-memory job list longer than necessary.
import puppeteer from 'puppeteer';
export async function renderPdf(url, outputPath) {
const browser = await puppeteer.launch({
// Prefer a working sandbox; do not add --no-sandbox routinely.
});
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle0', timeout: 90000 });
await page.pdf({ path: outputPath, format: 'A4', printBackground: true });
await page.close();
} finally {
await browser.close();
}
}
4. Validate Puppeteer, Chrome, and the container
Browser compatibility
Record the Puppeteer version, Node.js version, container image digest, launch options, and browser executable path. A normal Puppeteer installation downloads a compatible Chrome for Testing build. If your image points to a system Chromium executable, verify that its version is supported by the installed Puppeteer version and test that exact image in CI.
Rank #2
Sandbox
Chrome can fail to launch when Linux sandbox support is unavailable. Puppeteer documents --no-sandbox only for content the operator absolutely trusts and strongly discourages using it as a routine fix. Prefer configuring a usable sandbox in the image and runtime. A sandbox error is a browser-startup problem, not evidence that the pod needs more memory.
Writable profile and cache paths
Restricted or read-only containers must provide writable locations for Chrome’s profile, cache, and configuration. Mount an emptyDir or another appropriate writable volume and point temporary directories there when required by your image. Chrome may fail before Puppeteer connects if these paths cannot be created. Also check for zombie Chrome processes after a failed job; they can consume memory until the container is restarted.
5. Check readiness and the PDF call
Puppeteer’s PDF API is page.pdf(). Its documented behavior waits for fonts to load by default, but that does not guarantee that your application’s data, images, or client-side rendering is complete. Make page readiness explicit:
await page.goto(url, { waitUntil: 'networkidle0', timeout: 90000 });
await page.waitForSelector('#report-ready', { timeout: 30000 });
await page.evaluate(() => document.fonts.ready);
await page.pdf({ path: '/tmp/report.pdf', format: 'A4' });
Use a selector, application-side “ready” marker, or a carefully chosen delay when network-idle is not appropriate. Avoid an unbounded wait: a page that never becomes ready can hold a Chromium process until the pod is killed.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →6. Apply the fix that matches the evidence
| Evidence | Likely scope | Action |
|---|---|---|
OOMKilled and usage reaches the limit during concurrent jobs |
Concurrency-driven peak | Bound the queue, lower parallelism, then right-size the limit from measured peaks. |
| Memory rises after every job | Leak or retained browser data | Close pages and contexts, inspect heap/process counts, and isolate a minimal reproducer. |
| One unusually large document fails | Workload spike | Route large jobs separately, limit their concurrency, and measure their peak. |
| Eviction or node-pressure events | Cluster capacity | Review requests, spread replicas, and add or resize node capacity. |
No usable sandbox!, profile errors, or launch timeout |
Runtime configuration | Fix sandbox and writable paths; do not treat a memory increase as the solution. |
| Browser and Puppeteer versions differ | Compatibility | Use Puppeteer’s downloaded browser or validate the explicitly selected executable. |
7. Verify the change safely
- Deploy the candidate image and resource settings to a staging namespace.
- Run a representative mix of short, long, image-heavy, and font-heavy PDFs at production concurrency.
- Watch pod peak memory, restart count, node pressure, Chrome process count, job duration, and PDF correctness.
- Confirm that memory returns toward baseline after jobs and that no orphaned browser processes remain.
- Promote only after the termination reason disappears under sustained load, not merely after one successful PDF.
Common failures and recovery steps
“It is exit 137, so I increased memory”
Recheck Reason and Events. Exit 137 is not synonymous with OOM. If the reason is eviction or a probe action, repair node capacity or probe behavior instead.
The limit was raised, but restarts continue
Compare the new peak with the new limit and inspect node allocatable memory. If usage keeps growing between jobs, find the leak; if several pods peak together, reduce concurrency or add capacity.
Chrome says “No usable sandbox!”
Provide a supported sandbox in the container. Do not make --no-sandbox the default, especially for untrusted URLs.
Chrome cannot create files
Check the effective user, mounted volumes, TMPDIR, profile, and cache paths. Ensure the paths exist and are writable before launching.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #4
PDF generation hangs
Set navigation, selector, and overall job timeouts. Log the URL stage that is waiting, and make sure a failed job closes its page and browser in finally.
Pages crash but the container is not OOM-killed
Collect Chrome stderr and browser version details, reduce concurrency temporarily, and test the same URL in the same image. A renderer crash, incompatible executable, or page-specific browser fault can occur below the pod limit.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your requirement is a clean website image or PDF rather than operating Chromium in AKS, ScreenshotNeo provides a single HTTP request. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options, including full-page capture, lazy-image loading, CSS selectors, device presets, retina scale, PDF paper and margin controls, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous webhooks, bulk capture, and usage reporting.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Best Value
- Used Book in Good Condition
Frequently Asked Questions
Is exit code 137 proof that Puppeteer ran out of memory?
No. It proves SIGKILL was delivered. Confirm OOMKilled in the termination state and check events for eviction, probe failure, or manual deletion.
What memory limit should an AKS Puppeteer PDF pod use?
No universal safe value is established. Measure peak usage for your pages and concurrency, then leave capacity for Node.js, Chrome child processes, and other container work.
Should I add --no-sandbox to stop crashes?
Only for content you absolutely trust, and not as a routine fix. Configure a usable Linux sandbox whenever possible.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

