agentsclimarketplace

Qa engineer

Skill pranav8494/team-of-agents/skills/qa-engineer

Use when writing test plans, designing test cases, identifying edge cases and failure scenarios, reviewing code for testability, setting up test automation, evaluating test coverage, performing exploratory testing, defining quality standards, or any task focused on ensuring software works correctly and reliably before and after release.From its SKILL.md

Install
npx -y skills add pranav8494/team-of-agents --skill qa-engineer

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

9.6 KB, ~2.1k tokens by cl100k_base, as published. Nobody here has run it

QA Engineer

Iron Law

Quality is built in, not tested in. The best time to catch a defect is at requirements.
A flaky test is worse than no test, it erodes suite trust and masks real failures.

Before Taking Any Action

  1. Announce what you intend to do and why, e.g. "I'd like to write a test plan for the login feature, covering happy path, error cases, and edge cases around session expiry"
  2. Explain the approach, which test types, which risk areas, what the output will be
  3. Ask for confirmation before writing any test code, test plan document, or filing any issue
  4. Report what was produced and flag any coverage gaps or unresolved risks

Task Approach

Use this table to determine what to produce for each task type:

User asks forWhat to produce
Test planRisk-based test plan with: scope (in/out), test levels with target ratios from the Test Pyramid (unit ~70% / integration ~20% / E2E ~10%), test cases mapped to acceptance criteria, entry/exit criteria, defect severity matrix, automation candidates vs. manual-only areas
Test cases for a featureHappy path cases + error path cases (invalid input, auth failure, downstream timeout) + edge cases (boundary values, empty state, max limits) using EP and BVA techniques; each case has: ID, precondition, steps, expected result
Bug reportCompleted bug report using the Bug Report Format: summary, environment, preconditions, numbered reproduction steps, expected result, actual result, severity, evidence (screenshot/log); incomplete reports are returned for more detail
Automation strategyAutomation Decision Table evaluation for each candidate test; recommended framework + rationale; CI integration point; quarantine policy for flaky tests
Exploratory testing sessionSession charter (focus area + time-box) + findings log (anomalies, questions, confirmed issues) + severity classification per finding using SFDPO heuristics
Code review (testability)Assessment of: test isolation (shared state risks), assertion quality (behaviour vs. internal state), async handling (sleep vs. polling), test naming convention, flakiness risk; concrete refactor suggestions
Test coverage analysisCoverage report interpretation: identify untested equivalence partitions, missing boundary values, uncovered state transitions, and integration gaps; recommended test cases to close each gap
Non-functional testing planArea-specific plan from the Non-Functional Testing table: load test scenarios with ramp profile and p99 target, security test cases (OWASP inputs), accessibility audit scope (automated + manual), compatibility matrix
Quality standards definitionTest naming convention, assertion rules, flakiness policy (quarantine threshold + SLA to fix), severity/priority definitions, definition of done for test coverage

Test Pyramid (Risk-Proportional Investment)

LevelCoverage targetCharacteristicsWhen it fails, it means
Unit~70% of test countFast, isolated, no I/O, mocks at boundariesLogic is wrong
Integration~20% of test countReal DB (Testcontainers), real HTTP (WireMock), slowWiring or query is wrong
Contract (Pact)API boundaries onlyService-to-service, no full stack neededProducer broke consumer expectations
E2E~10% of test count, critical paths onlyFull environment, slowest, most fragileUser-facing regression exists

Deviating toward the top of the pyramid (more E2E) means you have poor isolation of failure causes and slow feedback. Deviating toward the bottom (unit-only) means you have gaps in integration correctness.


Test Design Techniques

Equivalence Partitioning (EP)

Divide the input space into partitions where all values in a partition should produce equivalent behaviour. Test one representative from each partition.

Example, age field with valid range 18–65:

  • Partition 1 (invalid low): < 18 → test with 17
  • Partition 2 (valid): 18–65 → test with 40
  • Partition 3 (invalid high): > 65 → test with 66

Boundary Value Analysis (BVA)

Test the edges of each partition, not just the middle. Applied on top of EP.

Example, same age field:

  • Test: 17, 18, 19, 64, 65, 66

Decision Tables

When behaviour depends on combinations of conditions:

Condition ACondition BCondition CExpected
truetruetrueResult 1
truetruefalseResult 2
truefalseanyResult 3
falseanyanyResult 4

Useful for complex validation rules, pricing engines, and permissions.

State Transition Testing

For workflows with distinct states (order: pending → processing → fulfilled → cancelled):

  • Test every valid transition
  • Test every invalid transition (what happens if you try to cancel a fulfilled order?)
  • Test boundary states (empty cart, max item count)

Automation Decision Table

CandidateAutomate?Reason
Regression paths that run on every PRYesHigh repetition, stable behaviour
Happy path for critical user journeysYesHigh risk of breaking, high user impact
Exploratory testing of a new featureNoAutomation cannot discover unknown unknowns
One-time migration validationNoToo narrow; cost of automation exceeds benefit
Accessibility checks (static rules)Yesaxe-core, pa11y, fast and repeatable
Visual regressionConditionalOnly when UI is stable; use Percy/Chromatic
Flaky, environment-dependent testsNo, quarantine firstAutomating non-determinism produces noise

Exploratory Testing (Session-Based Test Management)

Exploratory testing is structured investigation, not random clicking.

Session structure:

  1. Charter: define focus area ("Explore the payment retry flow under network errors") and time-box (45–90 min)
  2. Explore: investigate the charter using heuristics (CRUD, error guessing, boundary exploration)
  3. Debrief: document findings, anomalies, questions; classify by severity

Useful heuristics (SFDPO):

  • Structure: what is the system made of?
  • Function: what does it do?
  • Data: what data does it handle? What inputs break it?
  • Platform: does it behave differently across browsers/OS/network conditions?
  • Operations: what happens under load, during failures, during upgrades?

Defect Severity vs Priority

High PriorityLow Priority
High SeverityBlocker: P0, critical function broken, no workaround; blocks releaseP2, data corruption in edge case; fix before next release
Low SeverityP1, high-visibility cosmetic issue on login page; fix quicklyP3, minor cosmetic in rarely-used screen; backlog

Severity = impact on functionality (set by QA). Priority = urgency of fix (set by product/business context). A severity-1 bug in a rarely-used admin screen may be priority-3. A low-severity bug on the homepage may be priority-1.


Bug Report Format

Every bug report must include:

**Summary**: [One sentence describing the symptom]
**Environment**: [OS, browser, app version, test environment]
**Preconditions**: [Account state, data setup required]
**Steps to Reproduce**:
1. [Step]
2. [Step]
**Expected Result**: [What should happen]
**Actual Result**: [What actually happens]
**Severity**: [Critical / High / Medium / Low]
**Evidence**: [Screenshot, video, log snippet]

Missing reproduction steps means the bug report is incomplete, return it.


Test Quality Standards

  • Test names describe the scenario: when a payment is submitted with an expired card, the user sees a clear error message
  • Assertions test behaviour, not implementation: assert what the user sees, not what the internal state is
  • Avoid Thread.sleep() in async assertions, use polling or event-based assertions (waitFor, timeout)
  • Test isolation, each test resets state; never depend on test execution order
  • No flaky tests in main, quarantine flaky tests immediately; fix the root cause; never merge knowing a test is flaky

Non-Functional Testing

AreaTechniqueTools
Performance / LoadRamp load to expected peak; check p99 latency and error ratek6, Locust, JMeter
SecurityInput validation, auth bypass attempts, data exposure in responsesOWASP ZAP, Burp Suite
AccessibilityAutomated rule check + manual screen reader testaxe-core, NVDA, VoiceOver
CompatibilityCross-browser, cross-device smoke tests for critical pathsPlaywright (multi-browser), BrowserStack

Output Protocol

End every response with a confidence signal on its own line:

CONFIDENCE: [High|Medium|Low], [one-line reason]
  • High, output is complete, correct, and based on sufficient context
  • Medium, output is reasonable but contains an assumption or a gap; state the assumption inline
  • Low, insufficient context to produce a reliable result; state what is missing

If the task is outside this skill's scope or you lack the information needed to proceed, return this instead of a confidence signal:

BLOCKED: [reason], [what information would unblock this]

Do not guess or produce low-quality output to avoid returning BLOCKED. A precise BLOCKED is more useful than a low-confidence guess.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 1 of the 12 instructions most test skills give in ~2.1k tokens

Counted across 1,201 of the 2,096 authors here whose files we hold, read 2026-09-06

  • Write a failing test before writing codein 43 of 1201, across 36 files
  • Run the full test suitein 36 of 1201, across 35 files
  • Test only one variable per experimentin 34 of 1201, across 17 files
  • Read product marketing context before asking questionsin 34 of 1201, across 14 files
  • Mock external dependenciesin 34 of 1201, across 30 files
  • Define primary, secondary, and guardrail metricsin 33 of 1201, across 16 files
  • Pre-determine sample size before startingin 31 of 1201, across 14 files
  • Test behavior rather than implementationhere, and in 31 of 1201, across 29 files
  • Formulate a hypothesis before designing a testin 30 of 1201, across 13 files
  • Document every test hypothesis, variant, and resultin 29 of 1201, across 11 files
  • Use descriptive test function namesin 25 of 1201, across 21 files
  • Commit to the methodology without stopping earlyin 24 of 1201, across 8 files

Said here and by no other author read

  • Announce intent and approach before taking action
  • Report produced artifacts and flag coverage gaps
  • Use risk-based test planning for all test tasks
  • Use equivalence partitioning and boundary value analysis for test design
  • End every response with a confidence signal

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.