agentsclimarketplace

Browser test

Skill TheophilusChinomona/idev/skills/browser-test

Code-as-action browser testing: verify UI changes and user flows by writing re-runnable Playwright scripts, capturing screenshots/console/network evidence, and producing test reports. Use when verifying a web feature in a real browser, reproducing a UI bug, smoke-testing a flow after changes, or asked to test the site / run browser tests.From its SKILL.md

Install
npx -y skills add TheophilusChinomona/idev --skill browser-test

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

5.1 KB, ~1.2k tokens by cl100k_base, as published. Nobody here has run it

Browser Test — Code-as-Action Browser Verification

Verify web features by writing Playwright scripts, not by imagining what a browser would show. The script is the durable artifact: it lands in the project's test library and re-runs on every future change to that flow. (Concept adapted from Microsoft's Webwright: treat the browser as a disposable environment your code spawns; the workspace — scripts, logs, screenshots — is the state.)

State layout

.claude/idev/browser-tests/
├── scripts/        # one script per flow — the accumulated E2E library
├── artifacts/      # screenshots, console/network logs per run (gitignorable)
└── reports/        # structured test reports

Setup detection (in order)

  1. Project already has @playwright/test in package.json → use it (npx playwright test). Respect the project's playwright.config.*.
  2. Python playwright importable → use the sync API (python3 script.py).
  3. Neither → tell the user the one-time setup (npm i -D @playwright/test && npx playwright install chromium) and ask before installing anything.

Script conventions

  • Reuse first: grep scripts/ for the flow before writing a new script. Adapt and parameterize an existing one rather than duplicating it.
  • One flow per script, named for the flow: login-flow.spec.ts, checkout-smoke.spec.ts.
  • Base URL from env, never hardcoded: const BASE = process.env.BASE_URL ?? 'http://localhost:3000' — scripts must run against dev, staging, or CI unchanged.
  • Assert real outcomes, not just "page loaded": visible text, URL after navigation, row counts, toast messages. Prefer role/label selectors (getByRole, getByLabel) over CSS chains — they survive refactors.
  • Capture evidence on every run: screenshot at each key step and on failure; collect console errors and failed network requests:
import { test, expect } from '@playwright/test';
const BASE = process.env.BASE_URL ?? 'http://localhost:3000';

test('user can log in', async ({ page }) => {
  const consoleErrors: string[] = [];
  page.on('console', m => m.type() === 'error' && consoleErrors.push(m.text()));
  page.on('requestfailed', r => consoleErrors.push(`NET ${r.url()}`));

  await page.goto(`${BASE}/login`);
  await page.getByLabel('Email').fill(process.env.TEST_USER ?? '[email protected]');
  await page.getByLabel('Password').fill(process.env.TEST_PASS ?? 'test');
  await page.getByRole('button', { name: 'Sign in' }).click();

  await expect(page).toHaveURL(/dashboard/);
  await page.screenshot({ path: '.claude/idev/browser-tests/artifacts/login-ok.png' });
  expect(consoleErrors, `console/network errors: ${consoleErrors}`).toHaveLength(0);
});
  • Credentials via env vars only (TEST_USER, TEST_PASS) — never commit secrets into scripts; note required vars at the top of the script.
  • Headless by default; headed only while debugging a script.

Failure discipline (the important part)

When a run fails, classify before acting:

  • Script bug (wrong selector, missing wait, bad assumption) → fix the script and re-run. Max 3 repair iterations; then report what's blocking.
  • App bug (console error, failed request, wrong behavior) → do NOT silently change the app to make the test pass. Capture the evidence (screenshot, error text, request) and report it as a finding.
  • Environment issue (server not running, missing test data) → report what's needed; see the project's run conventions.

Report format

Write to reports/<YYYY-MM-DD>-<flow>.md:

# Browser Test Report: <flow> — <date>
Target: <BASE_URL>   Script: scripts/<name>   Result: PASS | FAIL

| # | Check | Result | Evidence |
|---|-------|--------|----------|
| 1 | login redirects to /dashboard | PASS | artifacts/login-ok.png |
| 2 | zero console errors | FAIL | "TypeError: x is undefined" |

## App bugs found
- <error + repro + evidence path>  (also log per the lessons-learned skill if recurring)

## Not covered
- <flows/states this run did not exercise>

Report only what actually ran — no inferred results. If a check didn't execute, it's "not covered", not "pass".

Anti-patterns

  • Verifying UI changes by reading the JSX/HTML and declaring it works — that's what this skill exists to replace.
  • Writing a throwaway script and deleting it — save it to scripts/; the library is the point.
  • waitForTimeout(3000) sprinkled to "fix" flakiness — wait for conditions (expect(...).toBeVisible(), waitForURL) instead.
  • Testing through 10 UI steps what one API call could set up — use the API for arrange, the browser for act/assert.

Concept adapted from Webwright (MIT, © Microsoft Corporation).

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most e2e browser skills give in ~1.2k tokens

Counted across 499 of the 513 authors here whose files we hold, read 2026-09-06

  • Capture screenshots, videos, and traces on failurein 32 of 499, across 23 files
  • Close the browser when donein 22 of 499
  • Interact with elements using snapshot refsin 21 of 499, across 20 files
  • Wait for specific network responses instead of fixed timeoutsin 20 of 499, across 10 files
  • Keep tests independent with no shared statein 19 of 499, across 17 files
  • Use Page Object Model classes to encapsulate page interactionsin 19 of 499, across 9 files
  • Locate elements with data-testid attributesin 19 of 499, across 10 files
  • Quarantine flaky tests with fixme or skipin 17 of 499, across 7 files
  • Upload test artifacts after every CI runin 17 of 499, across 8 files
  • Wait on conditions instead of using fixed sleepsin 17 of 499, across 13 files
  • Clean up test data after each testin 17 of 499, across 16 files
  • Test user-visible behavior, not implementation detailsin 16 of 499, across 10 files

Said here and by no other author read

  • Write Playwright scripts instead of imagining browser behavior
  • Write one script per flow
  • Take the base URL from environment variables
  • Assert real outcomes using role or label selectors
  • Capture screenshots and console/network errors on every run
  • Limit script repair to three iterations, then report blockers

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.