Dev frameworks
How to Run Puppeteer in Production Without It Crashing (and When to Stop)
October 4, 2026 · 8 min read · Grabbit Team

Puppeteer in production has a signature failure: it works perfectly on your laptop, passes every test, and then falls over the moment real traffic hits it. The browser that launched instantly in development cannot find Chrome in the container. The capture that always worked times out on a slow page. Memory that sat flat under one request climbs until the box is OOM-killed at concurrency. None of it showed up locally because local never looked like production.
This guide covers why that gap exists and the fixes teams actually ship for each failure mode: launching Chromium in a container, surviving navigation timeouts and disconnected browsers, and keeping memory bounded under load. Then the honest part: if screenshots are the whole job, the most reliable fix is to not run Chromium at all.
The short answer
- Launch Chromium correctly in a container. Use
--no-sandboxand--disable-dev-shm-usage, install the system libraries your base image is missing, and run under an init process so crashed browsers get reaped. - Make every capture fallible. Set explicit timeouts, close pages in a
finallyblock, handle the browserdisconnectedevent, and retry transient failures with a capped backoff. - Bound concurrency to RAM. Each Chromium instance wants a few hundred MB. Queue captures and cap workers to what the container can actually hold.
- Recycle the browser. A long-lived browser drifts upward across thousands of navigations. Treat Chromium as disposable and recycle it after a fixed number of jobs.
- Then ask if you need Chromium at all. If the deliverable is an image of a URL, a hosted screenshot API removes the fleet and the failure modes with it.
The rest explains each one.
Why it works locally and breaks in production
The four differences between your laptop and a production container cause nearly every "works on my machine" Puppeteer crash:
- Chrome is already installed locally. Your desktop has a browser; a slim container image does not, and it is often missing the shared libraries Chromium links against. You see
Could not find Chromeor a launch that dies immediately. - Shared memory is tiny in containers. Chromium uses
/dev/shmheavily. On a desktop that is gigabytes; in Docker it defaults to 64MB. Run a few captures and Chromium crashes with a renderer failure that never reproduces on your machine. - Production runs captures in parallel. Local testing is one page at a time. Under real traffic you run many at once, and memory that was comfortable serially becomes an out-of-memory kill.
- Production hits the messy web. Slow pages, network blips, bot walls, and redirects that your happy-path local URL never exercised. Each one is a timeout or a thrown error your code has to survive.
So the problem is rarely a bug in your logic. It is that local never stressed the three things production does: a bare environment, shared memory pressure, and concurrency against real pages.
Fix 1: launch Chromium correctly in a container
Most Docker crashes are one of three things: no shared memory, missing shared libraries, or orphaned processes. Launch with flags that address the first, install dependencies for the second, and run an init process for the third:
const browser = await puppeteer.launch({
headless: true,
args: [
'--no-sandbox', // required as non-root in most containers
'--disable-dev-shm-usage', // use /tmp instead of the tiny default /dev/shm
'--disable-gpu',
],
});
--disable-dev-shm-usage is the single most common fix: it makes Chromium write to /tmp instead of the 64MB /dev/shm that containers ship with, which removes the renderer crashes that only appear under a little load. For the missing libraries, your base image needs the fonts and shared objects Chromium links against (libnss3, libatk-bridge2.0-0, libgbm1, libasound2, and the rest); install them in the image or start from a base that already bundles them. And run the container with --init (or tini as PID 1) so a crashed Chromium does not leave an orphaned process the OS still counts against your memory.
Fix 2: survive timeouts and disconnected browsers
In production, captures fail. Pages hang, browsers die, containers get killed mid-job. The code that works is the code that assumes failure and cleans up anyway:
async function capture(browser, url) {
const page = await browser.newPage();
try {
await page.goto(url, { waitUntil: 'networkidle2', timeout: 30000 });
return await page.screenshot({ fullPage: true });
} finally {
await page.close(); // runs on success, throw, and timeout alike
}
}
The finally is the point: a page closed only on the happy path leaks every time goto throws, which is exactly when captures pile up. Beyond that, set a timeout on goto and an overall timeout on the whole job so a stuck page cannot hang a worker forever. Listen for the browser disconnected event and replace the browser instead of reusing a dead one, because a disconnected error almost always means the Chromium process died (OOM, a crash, or a killed container). Retry transient failures once or twice with a short backoff, but cap the retries so a permanently broken URL does not loop.
Fix 3: bound concurrency to the RAM you have
The crash that surprises teams most is the one that only appears under load. Each Chromium instance can use a few hundred megabytes when it is rendering a heavy page, so on a 1GB container you can run a small handful of concurrent captures, not dozens. Spawning a browser per incoming request is the fastest route to an OOM kill.
Put a fixed-size queue in front of captures and size the worker pool to the container's memory:
let active = 0;
const MAX = 3; // tune to the container's RAM, measured under load
const queue = [];
async function runCapture(job) {
if (active >= MAX) {
await new Promise((resolve) => queue.push(resolve));
}
active++;
try {
return await job();
} finally {
active--;
queue.shift()?.();
}
}
Tune MAX by watching real resident memory under representative traffic, not by guessing. A queue that holds the line at three concurrent captures is more reliable than ten that OOM the box on a bad minute. For a dedicated tool, puppeteer-cluster does the same queueing with built-in retry and per-task error handling.
Fix 4: recycle the browser
Even with clean page hygiene and bounded concurrency, one browser held open for the life of the process drifts upward as state accumulates across navigations. The pattern teams settle on is a pooled browser recycled after a fixed number of jobs:
let browser;
let jobs = 0;
const MAX_JOBS = 200;
async function getBrowser() {
if (!browser || jobs >= MAX_JOBS) {
if (browser) await browser.close();
browser = await puppeteer.launch({
args: ['--no-sandbox', '--disable-dev-shm-usage'],
});
jobs = 0;
}
jobs++;
return browser;
}
You reuse the browser enough to amortize the several-hundred-millisecond startup cost, then throw it away before accumulated state becomes a problem. Treating Chromium as disposable is the reliability idea, not any single flag. The deeper version of this, including how to find the real source of a leak, is in the guide on Puppeteer memory leaks in production.
When to stop running Chromium
Here is the honest part. Every fix above is real, and if you need a browser you should ship all of them. But look at what the work actually is: sandbox flags, shared-memory tuning, chasing missing shared libraries, concurrency caps, zombie reaping, browser recycling, and an on-call rotation to catch the crash that only happens at 3am under load.
If the only reason you reached for Puppeteer is to capture a screenshot of a URL, that is a lot of infrastructure for a job that is one HTTP request. A hosted screenshot API runs the browser for you and returns a hosted image, so there is no Chromium to launch, no /dev/shm to tune, no dependencies to install, and no memory graph to watch:
curl -X POST https://grabbit.live/api/v1/grabs \
-H "Authorization: Bearer sk_live_your_key" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"width": 1280,
"full_page": true,
"format": "webp"
}'
The parameters map to the Puppeteer options you were already setting: width (320 to 1920), height (240 to 1080), full_page for the whole scroll height, format (png, jpeg, or webp), delay_ms (0 to 10000) to wait for late content, and selector to capture a single element. Grabbit is $0.002 per grab. Annual-plan credits reset at renewal rather than monthly, while separately purchased top-up credits do not expire while your account remains active, so you pay for captures rather than browser uptime. The Node.js screenshot guide shows the same request from fetch.
This is not "never use Puppeteer." Keep the library when you click, type, scrape, or run end-to-end tests and genuinely need to drive the browser; for that decision, see Puppeteer alternatives by job. But when the deliverable is a picture of a page, handing the browser off is how the crash stops being your problem.
The takeaway
Puppeteer runs fine locally and crashes in production because production is a bare container with tiny shared memory, running many captures at once against the messy real web. Fix it by launching Chromium correctly (--no-sandbox, --disable-dev-shm-usage, the missing libraries, an init process), making every capture fallible with timeouts and finally-block cleanup, bounding concurrency to the RAM you actually have, and recycling the browser as disposable. Then ask the honest question: if screenshots are the whole job, the most durable fix is to stop running Chromium yourself.
FAQ
- Why does Puppeteer work locally but crash in production?
- Four things differ between your laptop and a production container. Your machine already has Chrome installed and plenty of RAM; a container often has neither. The default shared-memory mount (/dev/shm) is 64MB in Docker versus gigabytes on a desktop, and Chromium crashes when it runs out. Your laptop runs one capture at a time while production runs many at once, so memory that was comfortable becomes an out-of-memory kill. And production hits real network failures, bot walls, and slow pages that your happy-path local tests never see. The result is code that passes every test and falls over under load.
- How do I stop Puppeteer from crashing in Docker?
- Launch with --no-sandbox and --disable-dev-shm-usage so Chromium uses /tmp instead of the tiny default /dev/shm, install the Chromium system dependencies your base image is missing (fonts, libnss3, libatk, and friends), and run under an init process (Docker's --init, or tini) so crashed browsers get reaped instead of piling up as zombies. Most Docker crashes are one of those three: no shared memory, missing shared libraries, or orphaned processes.
- How do I handle Puppeteer navigation timeouts and disconnected browsers?
- Treat every capture as fallible. Wrap page.goto in a try/finally that always closes the page, set an explicit timeout on goto and on the whole job, and catch the browser 'disconnected' event so a dead browser is replaced rather than reused. Retry transient failures once or twice with backoff, but cap retries so a permanently broken URL does not loop forever. The disconnected error almost always means the browser process died (OOM, a crash, or a killed container), so the fix is upstream: more memory headroom or lower concurrency.
- How many Puppeteer instances can I run at once?
- Fewer than you expect. Each Chromium instance can use a few hundred megabytes under load, so on a 1GB container you can safely run a small handful of concurrent captures, not dozens. Put a fixed-size queue in front of your captures and tune the worker count to the container's RAM, measuring real resident memory under load rather than guessing. Spawning a browser per incoming request is the fastest way to OOM a box.
- Is it worth running Puppeteer in production just for screenshots?
- If capturing images of URLs is the only reason you run Chromium, the sandbox flags, shared-memory tuning, dependency chasing, concurrency limits, and zombie reaping are pure overhead for a job a hosted API does in one HTTP request. Self-hosting Puppeteer earns its keep when you also click, type, scrape, or run end-to-end tests and genuinely need to drive the browser. When the deliverable is a picture of a page, a hosted screenshot API removes the browser fleet and the on-call rotation that comes with it.
Capture any website with one API call
Get a free test key and wire your first request in two minutes.
Written by
Grabbit Team
Screenshots as a service
The team behind Grabbit, the screenshot API for developers and AI agents. We write about web capture, rendering, and automating screenshots at scale.
Keep reading

How to Fix Puppeteer Memory Leaks in Production (and When to Stop Running Chromium)
Why Puppeteer leaks memory in production, the fixes people actually ship (disposable browsers, --js-flags, per-container caps), how to find the real source, and when to stop running Chromium yourself and hand screenshots to an API instead.
Jul 21, 2026 · 9 min read

Puppeteer Alternatives in 2026: What to Use Instead (by Job)
An honest guide to Puppeteer alternatives in 2026, matched to the job: Playwright for automation, Selenium for cross-language, Cypress for app tests, and a hosted API when all you need is a screenshot of a URL.
Jul 15, 2026 · 6 min read

How to Take a Website Screenshot in Node.js (4 Ways)
Four ways to screenshot a website in Node.js: Puppeteer, Playwright, the capture-website wrapper, and a hosted API. Working code, the trade-offs, and when to skip running Chromium.
Jun 18, 2026 · 4 min read