Visual regression
Skill Syo-M/codex-frontend-skills/plugins/codex-frontend-skills/skills/visual-regression
Codex-native frontend skills, custom agents, profiles, and an evidence-backed evaluation harness for React, Next.js, Vite, and Astro.
npx -y skills add Syo-M/codex-frontend-skills --skill visual-regressionAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 24 days oldThe repository was created 24 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Visual regression testing with Storybook or Playwright screenshots, deterministic baselines, and diff review. 日本語の依頼例:「スクショテスト」「ビジュアルリグレッション」「見た目の差分」「Chromatic」。
SKILL.md
4.3 KB, 890 tokens by cl100k_base, as published. Nobody here has run it
Visual Regression Testing
VRT catches what DOM assertions can't: layout breaks, token regressions, theme bugs, overflow. It complements — never replaces — behavioral tests. If a DOM assertion can express the check (text content, element presence), use that; VRT is for how things look.
Stories are the VRT surface
- Component-level VRT runs over Storybook stories (Chromatic, or a Playwright/Vitest screenshot run against the built Storybook). Every visual state already has a story per the
storybookskill — that catalog IS the snapshot suite. Don't build a parallel page-screenshot harness for component states. - Runner precedence: if a vendor service (Chromatic) is already configured, add the story and stop — it's covered. Otherwise default to Playwright
toHaveScreenshotagainst the built Storybook, baselines committed alongside the spec and (re)generated only in the CI container. - Page-level VRT:
expect(page).toHaveScreenshot()in Playwright for a handful of critical pages/themes only, as its own Playwright project running on fixture data — never bolttoHaveScreenshotonto a real-backend functional E2E spec (those followtesting-playwright). Full-page screenshots of every route is a flake farm, not coverage.
Determinism — every flake source must be pinned
A VRT suite that cries wolf gets ignored; eliminate nondeterminism at the source:
- Time: mock the clock (
page.clock/ fake timers); seed any randomized data. Relative dates ("3 minutes ago") in fixtures, never realDate.now(). - Network: all data via MSW/fixtures (same handlers as the test suites). No real API in VRT.
- Fonts & images: wait for
document.fonts.readyAND for in-view images to finish decoding before capture — a half-loaded image is a diff. Self-hosted fonts only (a CDN hiccup is a diff). - Animation: disable it —
toHaveScreenshot({ animations: 'disabled' })is the primary mechanism. Emulatingprefers-reduced-motionalone is insufficient: permotion, state-conveying animations reduce to a quick fade, which can still be captured mid-fade. Never screenshot mid-animation. - Environment: baselines are generated in the SAME environment that compares them — the CI container (one OS, one browser, pinned versions, pinned
TZ=UTC+ locale, sinceIntloutput differs per environment). Local-macOS baselines vs Linux CI = permanent font-rendering diffs. Use--update-snapshotsin the CI container (or the vendor cloud) to (re)baseline (the bare flag updates only changed snapshots; passallto force everything). - Scrollbars & caret: content-dependent scrollbars and a blinking text caret cause intermittent diffs — size stories to their content and keep the caret hidden (Playwright's
toHaveScreenshotdoes this by default). - Leftovers: fixed viewport sizes;
maskgenuinely dynamic regions (avatars, maps) rather than widening thresholds.
Thresholds & scope
- Keep
maxDiffPixelRatiostrict (≤ 0.01); a loose threshold silently approves real regressions. If a region forces looseness, mask that region instead. - Screenshot the component/element under test, not the whole page, when the subject is a component — smaller surface, fewer false positives.
- Cover both themes and the key breakpoints for layout-critical pieces; don't multiply every story × every viewport × every browser reflexively (cost grows multiplicatively, signal doesn't).
Baseline review — diffs are code review artifacts
- A visual diff is either a regression (fix the code) or an intended change (approve the new baseline in the PR). Never auto-accept baselines, never regenerate-and-commit without looking at the diff — that converts the suite into a rubber stamp.
- Baseline updates ship in the same PR as the change that caused them, so the reviewer sees code and pixels together.
- Recurring flake in one snapshot = fix the determinism cause or delete the snapshot; regenerating or re-running until the diff happens to pass is banned — same retried-pass-is-a-flake-report policy as
testing-playwright.
Gives 0 of the 12 instructions most e2e browser skills give in 890 tokens
Counted across 407 of the 410 authors here whose files we hold, read 2026-08-06
- use page object model patternin 35 of 407, across 25 files
- Snapshot to get element refsin 24 of 407, across 14 files
- keep tests independentin 23 of 407, across 18 files
- Interact using refs from the latest snapshotin 23 of 407, across 11 files
- clean up test data after each testin 21 of 407, across 15 files
- test user behavior not implementationin 20 of 407, across 14 files
- quarantine flaky tests explicitlyin 19 of 407, across 10 files
- wait for specific network conditionsin 18 of 407, across 8 files
- re-snapshot after navigation or dom changesin 17 of 407, across 10 files
- Detect running dev servers before writing test codein 17 of 407, across 7 files
- use web-first assertionsin 17 of 407, across 14 files
- capture screenshots or videos on test failurein 17 of 407, across 14 files
Said here and by no other author read
- use DOM assertions instead of visual regression where possible
- add stories to existing vendor services before adding tests
- use Playwright toHaveScreenshot against built Storybook by default
- mock the clock and seed randomized data for tests
- use fixtures or mocks for all network requests
- wait for fonts and images to load before capture
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.