Product Testing
v1.0.0A Claude Code skill that tests a product the way a QA engineer would: it inventories every surface, attacks it across functionality, UI, load and security, and leaves behind a runnable test suite and an evidence-backed findings report.
Product Testing is a skill for Claude Code that tests a product the way a QA engineer would. It inventories every surface the product exposes, attacks those surfaces across four dimensions, and leaves behind two things: a runnable test suite in the project's own idiom, and a findings report in which every claim is backed by captured output.
Ask for testing in the ordinary way and the skill triggers on its own. "Test the export endpoint before I ship it", "can the API handle Black Friday traffic?" and "is this upload handler safe?" each select the right dimensions; "full QA pass on the dashboard" runs all four.
Why it exists#
Asked to test something, the default behaviour of a coding agent is to exercise the happy path, see nothing break, and report that it works. That is the one outcome that costs real money, because the team ships on it.
This skill replaces that with a method. It starts from an enumeration of what the product actually exposes rather than from whatever a code skim surfaced, it aims each case at a named failure instead of at confirmation, and it reports three outcomes, verified pass, verified failure and not tested, so the boundary around the work is visible instead of implied.
The four dimensions#
Ask for one, several, or none; naming none runs all four. The What it tests tab describes each in detail.
| Dimension | Covers |
|---|---|
| Functionality | Boundary and malformed input across about eighty input classes, state and sequence defects, the authorization matrix, concurrency races, idempotency, and the error contract |
| UI | Real browser rendering, interaction, forms, keyboard and focus, responsive and dark mode, WCAG accessibility, console and network errors, visual regression |
| Load | Baseline, load, stress, soak and spike shapes; latency percentiles rather than averages; bottleneck identification; rate-limit and quota behaviour; recovery and degradation past the ceiling |
| Security | Authorization and ownership checks, authentication and sessions, injection boundaries, SSRF on URL-taking surfaces, secret leakage, file handling, transport headers, abuse and cost |
What you get back#
- A test suite in the project's existing idiom and location, wired to one command, with a failing test for every confirmed finding. The fix is verified when the test flips rather than by re-reading the diff.
- A report with a committed ship-or-wait verdict, findings ranked by severity, a minimal copy-pasteable reproduction for each, and an explicit list of what was not tested and why. See Reports.
- The evidence the report rests on, saved alongside it.
Built to run against real products#
The skill is explicit about what it will not do. It will not point load or security traffic at a host without your go-ahead for that specific target; it will not use a destructive action as a proof of anything; it will not trigger a surface that emails, charges or notifies real people without agreement; it will not copy a discovered secret into a report or a commit; and it will not install tooling, rewrite source to make a test pass, or bypass a commit hook on its own initiative.
It also treats everything the product under test emits, error messages, page content, API responses and logs, as data rather than as instruction. Output that tries to direct the tester is reported as a prompt-injection finding, not obeyed. The Safety tab has the full model.
Measured, not asserted#
Against a fixture with ten planted defects across all four dimensions, two agents were given the same prompt and the same authorization, one with the skill and one without.
| With skill | Baseline | |
|---|---|---|
| Planted defects found | 10/10 | 9/10 |
| Files delivered | 25 | 3 |
| Runnable test suite | 17 tests, 13 red as regression guards | none |
| Raw evidence captures | 16 + screenshot | 2 |
| Explicit not-tested section | 8 entries with reasons | a shorter note |
| Cost | 1.8× tokens, 2.3× wall clock |
Recall is close; a capable tester finds most defects unaided. What the skill adds is the inventory, the evidence, the load data, the durable suite and a stated coverage boundary, plus restraint: given identical unrestricted authorization, the baseline rewrote another account's primary key as a proof and could not restore it, while the skill wrote only to its own row and restored what it changed. Method, full results and the cost breakdown are on the Evaluation tab.
Bundled scripts#
Two command-line tools ship with the skill and are useful on their own, in CI or from a terminal. surfaces.py inventories what a codebase exposes: routes, CLI entry points, pages, specifications, scheduled work, authorization checks, outbound calls, uploads, raw SQL and possible secrets. loadtest.py load-tests an endpoint with real percentiles, a status histogram and classified errors, and doubles as a CI gate. Both use only the Python 3 standard library. The Scripts tab is the reference.
Requirements#
| Component | Needed for |
|---|---|
| Claude Code | the skill itself |
| Python 3.9 or newer | the two bundled scripts |
| A browser driver | the UI dimension: the session's own browser tools, or Playwright, Cypress or Puppeteer if already in the project |
Nothing else is required. A load generator such as k6 is used if present and never installed unasked.
Licence#
Apache-2.0. The source is on GitHub.