A runnable fixture reproducing four visual regressions on an Optimizely CMS 12 / Alloy-style campaign page, checked by three tools side by side.
Three things, all measured on the same fixture, on the same afternoon:
- A competent Playwright suite ships a visibly broken page. 12 of 12 tests pass on a build where the checkout button is dark-on-dark and unreadable — including the tests at the exact viewport where the bug lives.
- Playwright's own screenshot comparison fails when nothing is wrong. 4 of 4 red on a byte-identical page, because a headed browser is 15px narrower than a headless one. It is a dimension mismatch, so no threshold tuning can absorb it.
- Applitools is green when the page is right and red when it isn't. 28/28 on the good build, 8 of 28 on the regression, both times across Chrome at three widths, Safari, Edge, iPhone 14 Pro and Galaxy S22 — from a single capture per test, on a Linux runner where Safari and Edge are not otherwise available at all.
It holds up on live, non-deterministic content too: with a price rotating through four locales and a per-request account email, scoped regions stay 28/28 green and still catch the regression at exactly the same 8 renders. A pixel threshold big enough to absorb that same content is ~22,400 pixels wide and blind to anything smaller, anywhere on the page.
And the part that closes the argument for a DXP estate: Optimizely ships no automated visual-regression check. Its own compare-versions view diffs content properties, and none of these four defects touch a content property. The reviewer gets a clean compare view and a broken page.
Nothing here is asserted. Every number came out of a run, and the Applitools counts were read back from the Eyes server rather than inferred from an exit code.
| Suite | What it represents |
|---|---|
tests/functional.spec.ts |
A typical CI Playwright suite — 3 widths × 4 Visitor Groups, 12 tests |
tests/native-visual.spec.ts |
Playwright's built-in toHaveScreenshot() — what a team already has |
tests/visual.spec.ts |
Applitools Eyes + Ultrafast Grid — 4 captures → 28 renders |
tests/dynamic-visual.spec.ts |
The same 28 renders against the page unfrozen — price, email, stock and timestamp all changing per request |
tests/fixture-integrity.spec.ts |
Proves each defect is genuinely broken while functional signals stay clean — 11 tests, no API key |
Read the write-up: True Red, False Green (15pp) —
or the condensed edition, Two Ways To Be Wrong (6pp),
same argument and numbers in about a third of the length. Both are committed PDFs, built
from docs/true-red-false-green.src.html and
docs/true-red-false-green-short.src.html via
npm run doc:pdf / npm run doc:short:pdf. An interactive HTML version of each is also
published: full ·
condensed.
Nothing here is a strawman. The functional suite runs the same viewport × Visitor Group matrix as the visual suites. The native suite masks both dynamic elements, disables animations, hides the caret, and covers all four Visitor Groups — i.e. it is what a competent team writes after a couple of rounds of flake triage.
Node 20+. Ports 4300 (app) and 4301 (stand-in CDN origin) must be free.
npm ci
npm run setup # fetches the brand font + installs Chromiumnpm run setup:font downloads Caveat (SIL Open Font License 1.1) to cdn/brand.ttf.
The font is not committed and any distinctive display face works — see
Brand font.
Then run the suites that need no credentials:
npm run test:fixture # 11 pass — proves the defects are real
DEFECTS="" npm run test:functional # 12 passFor the Applitools suites, add a key. Copy .env.example to .env, or
export it:
export APPLITOOLS_API_KEY=... # Applitools dashboard -> Account -> My API keyCapture your own baselines before comparing anything. They are not portable and are
not committed: Playwright keys its snapshots by platform (…-darwin.png) and by whichever
brand font was present at capture time, and Applitools baselines live in your own Eyes
account.
DEFECTS="" npm run baseline:native
DEFECTS="" NATIVE_WIDTH=1280 npm run baseline:native
DEFECTS="" npm run baseline:visual # needs APPLITOOLS_API_KEY
DEFECTS="" npm run baseline:dynamic # needs APPLITOOLS_API_KEYNow the head-to-head below reproduces. .github/workflows/visual.yml
runs the no-key suites plus the Applitools suite on Linux CI, and deliberately omits the
native screenshot suite — its baselines cannot cross platforms, which is the whole point.
Every number below came from running the suites in this repo. The ✅/❌ marks whether the tool reached the correct verdict, not whether tests passed — a suite that goes red on a broken build is right, and a suite that stays green on one is wrong.
| # | Scenario | Correct verdict | Functional (12 tests) | Native @1920 (4) | Native @1280 (4) | Applitools (4 tests / 28 renders) |
|---|---|---|---|---|---|---|
| 1 | Good build | green | ✅ 12 pass true green |
✅ 4 pass true green |
✅ 4 pass true green |
✅ 4 pass true green |
| 2 | Good build, headed browser, page byte-identical |
green | ✅ 12 pass true green |
❌ 4 fail FALSE RED |
— | ✅ 4 pass true green |
| 3 | overlap regression only |
red | ❌ 12 pass FALSE GREEN |
❌ 4 pass FALSE GREEN |
✅ 4 fail true red |
✅ 4 fail TRUE RED |
| 4 | All four defects | red | ❌ 12 pass FALSE GREEN |
✅ 4 fail true red |
— | ✅ 4 fail true red |
Collapsed to the only two questions that matter:
| Suite | Cries wolf (red, nothing wrong) | Sleeps through it (green, page broken) |
|---|---|---|
| Functional Playwright | never — it has no visual opinion at all | rows 3 and 4 |
Native toHaveScreenshot |
row 2 — a 15px scrollbar | row 3 |
| Applitools | not observed | not observed |
Same build, same page, same masks. The only change is a headed browser instead of
headless. It fails 4/4 with Expected an image 1920px by 1168px, received 1905px by 1168px — a 15px scrollbar. Because it is a dimension mismatch, maxDiffPixelRatio
is never consulted, so no threshold tuning can absorb it. Applitools passes on the same
build, because the Grid renders a DOM snapshot at a declared viewport instead of
screenshotting whatever the local browser happened to do.
The overlap defect lives in a 1024px–1366px media query. The functional suite passes
all 12 tests including the ones at 1280 and 1024, because pointer-events: none on the
scrim lets Playwright's hit-target check resolve straight through to the buried button.
The native suite at the usual 1920 CI viewport passes 4/4. Applitools goes red on all four
Visitor Groups — the 1024 and 1280 Chrome renders, 8 of 28 — with no extra enumeration.
Point the native tool at 1280 and it catches this (0.29 diff ratio). The axis is not capability, it is coverage economics: every configuration you want is one you must both think of and launch a real browser for, and Safari and Edge are not available to you on a Linux CI runner at all. Applitools got 7 configurations from 4 page loads because the Grid re-renders one capture server-side.
| Flag | What breaks | Real-world analogue | Why functional tests miss it |
|---|---|---|---|
overlap |
Hero gradient scrim regresses to 118% height, drapes over the CTA block | A shared block partial or CSS framework bump | Only in the 1024–1366px band. pointer-events: none lets click() succeed on a visually buried button. |
font |
CDN host dropped from the CSP allowlist; brand font blocked, site falls back to Georgia | A Cloudflare / CDN rule change by ops — no commit, no build, no pipeline run | No text, role, or count assertion observes font family. getComputedStyle returns the declared stack whether or not the font loaded. |
vgcollapse |
Personalized block clips to 0px — returning-customer only |
A gutter variable change interacting with a Visitor Group variant | CI runs anonymous. And the wrapper clips with max-height:0; overflow:hidden, so the <h2> keeps its own non-empty layout box — clipping is paint-time, not layout-time — and toBeVisible() passes on an element no human can see. |
heroimg |
Hero renders at native 3000px; document overflows | Image-processor / Cloudflare Polish config differing between DXP slots | The <img> has a valid bounding box and alt text. Overflow is not an assertion anyone writes. |
Assumes you have run the Quickstart already. Ports 4300 (app) and 4301
(stand-in CDN origin). APPLITOOLS_API_KEY is needed only for test:visual,
test:dynamic and the baseline:visual / baseline:dynamic scripts.
npm ci && npm run setup
# Prove the fixture is honest -- each defect measurably broken, functional signals clean
npm run test:fixture
# Capture baselines from the good build (once)
DEFECTS="" npm run baseline:native
DEFECTS="" NATIVE_WIDTH=1280 npm run baseline:native
DEFECTS="" npm run baseline:visual
# TRUE GREEN -- everything agrees the good build is fine
DEFECTS="" npm run test:visual # 4 pass
# FALSE RED -- native cries wolf on a byte-identical page
DEFECTS="" npm run test:native -- --headed # 4 fail (scrollbar)
DEFECTS="" npm run test:visual # 4 pass
# FALSE GREEN vs TRUE RED -- a real regression only Applitools catches
DEFECTS="overlap" npm run test:functional # 12 pass -- missed
DEFECTS="overlap" npm run test:native # 4 pass -- missed
DEFECTS="overlap" npm run test:visual # 4 fail -- caught
# DYNAMIC CONTENT -- the page changes correctly on every request
DEFECTS="" npm run baseline:dynamic # once
DEFECTS="" NATIVE_MAX_DIFF=0 npm run test:native # 4 fail -- 581-2,284px, nothing broken
DEFECTS="" npm run test:dynamic # 4 pass = 28/28 renders green
DEFECTS="overlap" npm run test:dynamic # 4 fail = 20 pass / 8 unresolved
# One report, functional and visual results together
npx playwright show-report
# Build the shareable doc, and a local PDF of it
npm run doc:build # inlines evidence PNGs -> docs/true-red-false-green.html
npm run doc:pdf # -> docs/true-red-false-green.pdf (A4, 15pp)
# Condensed edition -- same argument, ~1/3 the length, 5 exhibits
npm run doc:short # -> docs/true-red-false-green-short.html
npm run doc:short:pdf # -> docs/true-red-false-green-short.pdf (A4, 6pp)
# Regenerate the page screenshots in docs/evidence/ (dashboard shots need an
# authenticated Chrome over CDP -- see docs/capture-dashboard.mjs)
npm run doc:shotsDefects are also switchable per request for manual demoing:
http://localhost:4300/campaigns/spring-launch?defects=overlap,font
playwright.config.ts registers the Applitools reporter
alongside the standard HTML one:
reporter: [
['list'],
['html', { open: 'never' }],
['@applitools/eyes-playwright/reporter'],
],Eyes status and a deep link to every checkpoint land in the same HTML report as the
functional results — a verified 140 dashboard links after a full run. That removes the
"the build went red, now go find out where" step: one artifact, both kinds of result. And
failTestsOnDiff: 'afterEach' attributes each diff to the Visitor Group that produced it
rather than emitting one anonymous worker-level error.
The suite uses the fixture API — import { test } from '@applitools/eyes-playwright/fixture'
with configuration in use.eyesConfig — which is what the reporter hooks into.
Worth being exact, because the imprecise version gets pushed back on:
toBeVisible()resolves to non-empty bounding box and notvisibility: hidden. There is deliberately no occlusion or z-order test — that would make the API slow and flaky. Overlap, contrast, and clipping are structural blind spots, not coverage gaps you can close by writing more assertions.click()does hit-test, at the element's centre. It catches occlusion when three conditions hold: a test clicks that element, at that viewport, and the occluder accepts pointer events. Decorative overlays and gradient scrims routinely carrypointer-events: none, which breaks the third — seeoverlap.toHaveText()/toHaveCount()read the DOM. Font substitution, colour, clipping, reflow, and overflow are outside what they can express.
All of them cost real debugging time; they are in the code as comments.
1. Never use setBaselineEnvName() to separate DXP slots. An Applitools environment
is OS + browser + viewport, and baselines are keyed on (app, test, environment) — so
Chrome@1920 and iPhone 14 Pro@393×852 get separate baselines automatically, with nothing
to configure. setBaselineEnvName() overrides that key with one fixed string, collapsing
the whole matrix into a single shared baseline. The dashboard then shows a 393×852 iPhone
baseline diffed against a 1920×1080 desktop checkpoint, and exactly one config per test
passes while the rest report permanent diffs — including on a run against the very build
the baseline came from. Put the slot in the app name instead.
2. DeviceName has no iPhone_14 key (iPhone_14_Pro and iPhone_14_Plus exist).
A wrong key yields undefined and the Grid reports UFG environment(s): "undefined",
which reads like a service outage rather than a typo. Verify device keys against the
installed SDK.
3. The manual API has a failure surface the fixture API does not have.
If you hand-roll new Eyes(runner, config): runner.getAllTestResults() throws on diffs
by default, which silently makes any reporting code after it unreachable — pass false,
report, then fail deliberately. And eyes.abort() must go in a finally, or a test that
throws before close() leaves an open session and the batch never completes. Both were
real bugs here before the suite moved to the fixture API, where neither can happen.
4. Set the lowest layoutBreakpoints value at or below your narrowest device.
Galaxy S22 is 360px wide. With a floor of 1024 — or even 390 — the SDK warns that those
renders fall below the smallest breakpoint and captures their resources at 1023px instead,
which can mask a mobile-only regression.
Run it against the Production slot URL before Complete-EpiDeployment. That is the
only window where real Production config — CSP, CDN rules, blob paths, image processing —
is observable on a URL you can still abandon. A red batch means you skip the slot swap
instead of rolling back a live site.
# after Start-EpiDeployment, before Complete-EpiDeployment
- script: npm run test:visual
env:
BASE_URL: $(ProdSlotUrl)
SLOT: production
APPLITOOLS_BATCH_ID: $(Build.BuildId)BASE_URL also suppresses the local fixture servers, so the same suite runs against a
real environment unchanged.
The bigger gap this fixture cannot show: content publishing in Optimizely is not a deploy. A Content Approval workflow puts restructured content live via a runtime content event — no commit, no build, no pipeline trigger — so a PR-triggered suite does not run at all. That needs a visual run on a content-published webhook or a schedule against Production. Functional testing has no story for it.
5. A region is a rectangle derived from the element's box — so pin the box. Both
ignoreRegions and layoutRegions are computed from the element's bounding box. If the
element resizes between baseline and checkpoint, the rectangle resizes with it and the
sliver at the edge gets compared anyway. Left unpinned, the stock-counter span measured
114–127px across renders — not only from the digit count, but because the brand face has
variable digit advance widths — and 2 of 28 renders came back unresolved on a page with
nothing wrong. Pinning all four dynamic elements to a fixed width took it to 28/28.
This bites ignore regions just as hard as layout regions.
The fixture runs both frozen (?freeze=1, deterministic) and unfrozen (live). Unfrozen,
four things change on every request: promo-price rotates through
$1,249.00 / 1.249,00 € / £1,079.00 / CHF 1'119.00, account-email gets a new
local part, stock-counter moves between 3 and 19, and session-timestamp is never
repeatable.
| Approach | Good build, live page | overlap build |
|---|---|---|
| Pixel diff, zero tolerance | ❌ 4 fail — 581–2,284px differ, nothing is wrong | 4 fail |
Pixel diff, maxDiffPixelRatio: 0.01 + masks |
✅ 4 pass | ❌ 4 pass @1920 — FALSE GREEN |
| Applitools, scoped regions | ✅ 28/28 renders green | ✅ 20 pass / 8 unresolved — server-verified |
That 0.01 budget is ~22,400 pixels of the full-page 1920×1168 capture — ten to forty times the dynamic
content it was bought to absorb, applied to the whole page, unable to tell a rotating
currency from a regression of the same size. The Applitools row buys quiet from four
scoped selectors instead, and the overlap column proves the sensitivity survived: it is
the identical 8-render signature the frozen suite produces.
Region choice matters and the two kinds are not interchangeable:
layoutRegions(price, stock counter) — ignore the glyphs, keep asserting the structure. A currency switch stays green; a price row that collapses or shoves the CTA down does not. A pixel mask cannot express this: it is opaque, so it discards the layout along with the content.ignoreRegions(render stamp, account email) — ignore the area outright, for content with no invariant worth asserting. Still scoped to a selector, unlike a page-wide pixel budget.
- The Visitor Group hook is a
qa-visitor-groupcookie. Against a real estate this assumes a customIVisitorGroupCriterionreading a cookie or header — small, but it is work, and it gates the whole personalization matrix. - Experiment variations are not covered. Adding Optimizely Web Experimentation needs the force-variation query parameter for your snippet generation; the shape has changed across versions, so confirm it against the account rather than trusting a documented string.
tests/visual.spec.tsfreezes the dynamic content with?freeze=1;tests/dynamic-visual.spec.tsruns the same page live. Both are kept deliberately. But the regions are tuned against this page — four selectors and four pinned boxes is a small, honest surface, and a real estate has more of both. The technique transfers; the specific region list does not.- Optimizely product names were verified against Optimizely's own docs on 2026-08-21: Visual Builder is the default editing experience in CMS 13 (it replaced on-page edit, which is disabled in 13); content approvals is current for CMS 12 and 13, and the same mechanism is approval sequences in CMS (SaaS). No first-party automated visual-regression product was found. CMS 13 went GA 31 March 2026, after this fixture was built — the fixture models CMS 12 / Alloy, which still receives security and bug fixes, and the four failure modes are architectural rather than version-specific.
- Brand font. The fixture needs a distinctive display face so
that the
fontdefect — which drops the CDN host from the CSP allowlist — makes the page visibly fall back to a system serif. This repo ships no font:npm run setup:fontfetches Caveat (SIL Open Font License 1.1) tocdn/brand.ttf, which is gitignored. The face originally used here was Chalkduster, which is Copyright 2008 Apple Computer, Inc. and not redistributable; it was removed for that reason. Any distinctive.ttfdropped atcdn/brand.ttfworks — nothing refers to the font by name. Changing the font changes rendering, so recapture baselines afterwards. - Baselines are not portable and are not committed. Native snapshots are keyed by OS
and by the font present at capture time; Applitools baselines live in your Eyes account.
Run the
baseline:*scripts once against a known-good build before trusting any comparison.
MIT. The fetched brand font is licensed separately under the SIL Open Font License 1.1 and is not redistributed by this repository.
Not affiliated with, endorsed by, or sponsored by Optimizely. "Optimizely", "Episerver" and related marks belong to their respective owners, and are used here only to describe the platform this fixture models. The fixture itself renders a fictional brand ("Contoso") and contains no Optimizely code or assets. Product names and behaviour were verified against Optimizely's public documentation on the date noted above; re-verify before relying on them, because Optimizely renames things.