Back to blog

Testing & monitoring

BackstopJS: Visual Regression Testing with the Open-Source Standard (and Where a Screenshot API Fits)

October 7, 2026 · 6 min read · Grabbit Team

BackstopJS: Visual Regression Testing with the Open-Source Standard (and Where a Screenshot API Fits)

If you search for BackstopJS you land on a GitHub README and a handful of five-minute tutorials. This is the version that explains the whole loop in one place: what BackstopJS does, the four commands you actually run, what goes in the config, why Docker keeps coming up, and the one part the tutorials skip, which is how to get a consistent capture when you test deployed URLs.

For the wider landscape of tools, see the best visual regression testing tools. This post is specifically about BackstopJS.

What BackstopJS is

BackstopJS is an open-source visual regression testing tool. It takes a screenshot of a page, saves it as a baseline, re-takes the screenshot after you change code, and compares the two images pixel by pixel. When the difference crosses a threshold you set, the test fails and you get an HTML report showing the before, the after, and a highlighted diff.

It renders pages with headless Chrome, driven by Puppeteer or Playwright, and you control everything from a single backstop.json file. There is no hosted service and no account. The baselines live in your repo, the diffs run on your machine or your CI runner, and the review happens in a generated HTML report. That is the whole appeal: it is free and it stays inside your own environment.

The four commands

The entire workflow is four CLI commands:

  1. backstop init scaffolds a backstop.json config and the folder structure in your project.
  2. backstop reference captures the baseline screenshots, the images every future test is compared against.
  3. backstop test captures fresh screenshots and diffs them against the baselines, then opens an HTML report of the differences.
  4. backstop approve promotes the latest test images to baselines, which is how you accept an intentional change.

The mental model: reference says "this is correct," test asks "is it still correct?", and approve says "the change I just made is the new correct." You run test on every change and approve only when a diff is expected.

The config

Everything BackstopJS does is declared in backstop.json. The three keys that matter are viewports (the screen sizes to test), scenarios (the pages and the interactions to run before capturing), and engine (Puppeteer or Playwright). A minimal scenario looks like this:

{
  "viewports": [
    { "label": "mobile", "width": 375, "height": 667 },
    { "label": "desktop", "width": 1280, "height": 800 }
  ],
  "scenarios": [
    {
      "label": "Homepage",
      "url": "https://example.com",
      "selectors": ["document"],
      "delay": 500
    }
  ],
  "engine": "puppeteer"
}

selectors controls what gets captured: "document" is the full page, or you pass a CSS selector to diff a single component. delay waits before capturing, which matters for anything that animates or loads late. For richer cases you add clickSelector, hoverSelector, or a custom onReadyScript to drive the page into a specific state first.

Why Docker keeps coming up

Search for BackstopJS and Docker is the second thing everyone asks about, for one reason: screenshots are not deterministic across machines. Fonts render differently on macOS and Linux, antialiasing differs, and a diff that is green on your laptop fails in CI because the pixels are not identical. That is a false positive, and false positives are what kill trust in a visual test suite.

Docker fixes this by pinning the exact browser and OS so the rendering is byte-for-byte the same everywhere. BackstopJS ships an official image for it. You either set a dockerCommandTemplate in the config or run:

backstop test --docker

The cost is real, though. You are now building or pulling a container image, keeping it current, and running it in CI alongside everything else. For a team already deep in Docker that is nothing. For a small project it is a second system to maintain just so the screenshots match.

Where this gets expensive: deployed URLs

BackstopJS scenarios accept any URL, so testing a staging or production page looks trivial. In practice it is where the maintenance lives. To diff a deployed page reliably you need the same viewport, the same fonts, the same network conditions, and the same wait for content to settle on every single run. Miss any of those and you get diffs that are real pixel changes but not real regressions.

That is the self-hosted headless browser tax, and it is the part the tutorials never mention. One developer put the ops version of it plainly: running your own Chromium at concurrency is "where we lose weekends, memory creep, sticky workers, restart storms." A visual test suite that flakes on rendering noise gets muted within a month, and a muted suite catches nothing.

This is the seam where a screenshot API fits. Instead of running and babysitting a headless browser in CI, you call an API that captures the live URL at a fixed viewport on hosted infrastructure and hands back a stable image. You keep BackstopJS, or any diffing library, for the comparison and the report. The API just makes the capture deterministic without a browser in your pipeline:

curl -X POST https://www.grabbit.live/api/v1/grabs \
  -H "Authorization: Bearer sk_live_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://staging.example.com/pricing",
    "width": 1280,
    "height": 800,
    "full_page": true,
    "format": "png",
    "delay_ms": 500
  }'

You get back a hosted PNG of the rendered page. Pin width, height, and delay_ms and the capture is the same every run, which is exactly the consistency a pixel diff needs. Grabbit bills a flat $0.002 per successful capture, so a CI suite that captures a few hundred pages a day costs cents, and prepaid credits do not expire while your account is active. Test keys (sk_test_) return a deterministic placeholder for free so you can wire up the integration before spending anything.

When to use which

Reach for BackstopJS on its own when you are testing local components or a local dev server, you already run Docker, and you want a zero-cost, self-hosted suite you fully control. It is the open-source standard for a reason.

Reach for a screenshot API as the capture step when you are diffing deployed URLs, you do not want a headless browser in CI, or your existing suite flakes because captures are not consistent across runs. The diffing stays yours; only the rendering moves off your infrastructure.

For the broader comparison of self-hosted and hosted options, see the best visual regression testing tools, and for the workflow itself, visual regression testing: a practical guide.

FAQ

What is BackstopJS?
BackstopJS is an open-source visual regression testing tool. It captures baseline screenshots of your pages, re-captures them after a change, and diffs the two images pixel by pixel, producing an interactive HTML report of any differences. It runs headless Chrome through Puppeteer or Playwright and is driven entirely from a backstop.json config and a four-command CLI: init, reference, test, and approve.
Is BackstopJS still maintained?
BackstopJS is a mature, widely used project with thousands of GitHub stars and a stable command set that has not changed in years. Development is slower than it was at its peak, and the open issues skew toward edge cases such as false positives on dynamic content. It is reliable for the page-level visual diffing it was built for; teams that want a hosted review dashboard and managed baselines often pair it with, or move to, a SaaS tool.
What is the difference between BackstopJS and Percy?
BackstopJS runs entirely in your own environment: it captures screenshots with a local headless browser, stores baselines in your repo, and reviews diffs in a generated HTML report. Percy is a hosted service that renders screenshots on its own infrastructure, stores baselines in the cloud, and gives you a review dashboard with approve and reject buttons. BackstopJS is free but you run and maintain the browser; Percy is paid but removes the rendering and review infrastructure.
How do I run BackstopJS in Docker?
BackstopJS ships an official Docker image so the headless Chrome rendering is identical on every machine, which removes the font and antialiasing differences that cause false diffs across operating systems. Set engine to puppeteer and add dockerCommandTemplate to your backstop.json, or run backstop test --docker. The tradeoff is that you now build, pull, and keep a container image current in CI alongside the test run.
Can BackstopJS test deployed URLs instead of local components?
Yes. BackstopJS scenarios point at any URL, so you can diff a staging or production page directly. The catch is that the capture has to be deterministic: the same viewport, the same fonts, and the same wait for content to settle every run, or you get false diffs. Running your own headless browser in CI makes that consistency your problem. A screenshot API captures the live URL for you at a fixed viewport and hands back a stable image you can diff.

Capture any website with one API call

Get a free test key and wire your first request in two minutes.

Written by

Grabbit Team

Screenshots as a service

The team behind Grabbit, the screenshot API for developers and AI agents. We write about web capture, rendering, and automating screenshots at scale.

Keep reading