Product Testing

v1.0.0

A Claude Code skill that tests a product the way a QA engineer would: it inventories every surface, attacks it across functionality, UI, load and security, and leaves behind a runnable test suite and an evidence-backed findings report.

Claude CodePython 3.9+ for the bundled scriptsA browser driver for the UI dimension Updated 7 Oct 2026

Evaluation

A skill is easy to judge by reading and hard to judge by reading, so this one was measured. Everything below is self-contained, and the material behind it is public: the fixture, the answer key, the scoring script and both saved runs are in the repository's evals/ folder, with the written comparison in evals/RESULTS.md. The fixture is a deliberately vulnerable application; run it bound to loopback only.

Method#

A deliberately defective API was built with ten planted defects spread across all four dimensions — a missing ownership check, mass assignment, SQL injection, reflected XSS, unauthenticated server-side request forgery, a debug-mode stack trace, a non-atomic quota counter, byte-truncation that splits multibyte characters, an unbounded row limit, and five UI faults on a single page.

Every defect was verified to reproduce before the runs started, so the answer key is observed rather than assumed.

Two agents then received the same prompt against the same running instance — one pointed at the skill, one with no skill at all — both told they had full authorization to test every dimension. Neither could read the answer key or edit the fixture. Scoring requires a defect's surface and its mechanism to appear together before counting a hit, so "validation could be tightened" does not score as having found the unbounded-limit defect.

Caveats worth stating: one fixture, one run per arm, and a fixture smaller than a real product. Treat recall as indicative and the structural differences as the substantive result.

Recall#

10/10 with the skill, 9/10 without. The one the baseline missed was the unbounded row limit.

Recall is close, and that is the honest headline. On a fixture this size a capable tester finds most of what is there unaided. Recall is not where the skill earns its place.

What each run produced#

With skill Baseline
Files delivered 25 3
Report and evidence 82,921 chars 9,524 chars
Surface inventory yes no
Written test plan yes no
Raw evidence captures 16, plus a UI screenshot 2
Load data 5 scenarios as JSON, with a results table and stated environment none
Runnable test suite 17 tests, stdlib only, 13 red as regression guards none
Committed ship verdict yes no
Explicit not-tested section yes — 8 entries with reasons a shorter coverage note

Two of these are the durable value. The suite outlives the session: thirteen of its tests fail against the open findings and flip green as each fix lands. The not-tested section is what makes the report trustworthy — in this run it separated not applicable (no upload surface, no scheduled jobs, no TLS by design) from not reached (one browser engine, dark-mode and keyboard depth not exercised), and recorded that a parallel run had contaminated shared state, so one account's tests ran as the other user.

Behaviour under identical authorization#

Both runs had unrestricted permission to test. They did not behave the same way, which is the clearest evidence that the rules in Safety model do something.

Situation With skill Baseline
Proving the file-read flaw read the fixture's own README read a system password file
Proving the injection stopped at one error-based and one blind read carried through to data extraction
Cloud-metadata probe issued to prove no allow-list existed, then dropped —
Writes used as proof only to the caller's own row rewrote another account's primary key
State afterwards restored what it changed left an account broken, unable to restore

Neither agent exceeded its authorization. The difference is restraint: proof rather than exploitation, and reversible rather than not. The baseline's broken row then leaked into the other run — which is why the with-skill report carries a note about contaminated state, and is a small live demonstration of why "never use a destructive action as a proof" is a rule rather than a preference.

Cost#

With skill Baseline Ratio
Tokens 160,862 91,057 1.77×
Tool calls 53 29 1.83×
Wall clock 17m 49s 7m 48s 2.28×

Roughly 1.8× the tokens and 2.3× the time, which buys the inventory, the evidence, the load measurements and the suite. Worth it for a pre-release pass; not worth it for a one-line sanity check.

The scripts were tested too#

surfaces.py was run against a real 183-file codebase during development and produced three defects in itself on the first run — it scanned worktree copies so every file appeared twice, its cron pattern matched numeric arrays in unrelated source, and its event-listener pattern matched icon helpers. Two further false-positive classes surfaced after that: a database cursor's fetch() read as an HTTP client, and .sql migration files flagged as raw SQL. All five are fixed.

loadtest.py was exercised against seeded targets: it separated a p50 of 13ms from a p90 of 254ms where the mean hid the tail entirely, classified connection refusals and unexpected statuses correctly, gated on thresholds with a non-zero exit, and handled every edge thrown at it — a warmup longer than the run, a target not listening, an all-404 target — without a traceback.

Re-running it#

Any change to SKILL.md or a reference file that is meant to change what the skill finds, how it reports, or what it refuses to do deserves a re-measurement. Reading a revision and judging it plausible is exactly the failure the skill is written against.

evals/README.md has the procedure. One detail learned the hard way: run the arms sequentially and restart the fixture between them, because a shared instance lets one run's writes contaminate the other.

Two things were redacted from the saved runs before publication, and the same README documents both: machine-specific absolute paths became placeholders, and the fragment of a system password file that the unaided run captured while proving the file-read flaw was replaced with a redaction marker. The findings are unchanged. The skill-guided run needed no redaction, because it proved the same flaw against the fixture's own README.