# Product Testing: Evaluation

> A Claude Code skill that tests a product the way a QA engineer would: it inventories every surface, attacks it across functionality, UI, load and security, and leaves behind a runnable test suite and an evidence-backed findings report.

Page: https://developers.zactonz.com/skills/product-testing/evaluation/

Complete documentation for this product: https://developers.zactonz.com/skills/product-testing.md

A skill is easy to judge by reading and hard to judge by reading, so this one was
measured. Everything below is self-contained, and the material behind it is
public: the fixture, the answer key, the scoring script and both saved runs are in
the repository's [`evals/`](https://github.com/zactonz/zactonz-product-testing/tree/main/evals)
folder, with the written comparison in
[`evals/RESULTS.md`](https://github.com/zactonz/zactonz-product-testing/blob/main/evals/RESULTS.md).
The fixture is a deliberately vulnerable application; run it bound to loopback
only.

## Method

A deliberately defective API was built with ten planted defects spread across all
four dimensions — a missing ownership check, mass assignment, SQL injection,
reflected XSS, unauthenticated server-side request forgery, a debug-mode stack
trace, a non-atomic quota counter, byte-truncation that splits multibyte
characters, an unbounded row limit, and five UI faults on a single page.

Every defect was verified to reproduce before the runs started, so the answer key
is observed rather than assumed.

Two agents then received the same prompt against the same running instance — one
pointed at the skill, one with no skill at all — both told they had full
authorization to test every dimension. Neither could read the answer key or edit
the fixture. Scoring requires a defect's **surface and its mechanism** to appear
together before counting a hit, so "validation could be tightened" does not score
as having found the unbounded-limit defect.

Caveats worth stating: one fixture, one run per arm, and a fixture smaller than a
real product. Treat recall as indicative and the structural differences as the
substantive result.

## Recall

**10/10 with the skill, 9/10 without.** The one the baseline missed was the
unbounded row limit.

Recall is close, and that is the honest headline. On a fixture this size a capable
tester finds most of what is there unaided. Recall is not where the skill earns its
place.

## What each run produced

| | With skill | Baseline |
|---|---|---|
| Files delivered | 25 | 3 |
| Report and evidence | 82,921 chars | 9,524 chars |
| Surface inventory | yes | no |
| Written test plan | yes | no |
| Raw evidence captures | 16, plus a UI screenshot | 2 |
| Load data | 5 scenarios as JSON, with a results table and stated environment | none |
| Runnable test suite | 17 tests, stdlib only, 13 red as regression guards | none |
| Committed ship verdict | yes | no |
| Explicit not-tested section | yes — 8 entries with reasons | a shorter coverage note |

Two of these are the durable value. The **suite** outlives the session: thirteen of
its tests fail against the open findings and flip green as each fix lands. The
**not-tested section** is what makes the report trustworthy — in this run it
separated *not applicable* (no upload surface, no scheduled jobs, no TLS by design)
from *not reached* (one browser engine, dark-mode and keyboard depth not
exercised), and recorded that a parallel run had contaminated shared state, so one
account's tests ran as the other user.

## Behaviour under identical authorization

Both runs had unrestricted permission to test. They did not behave the same way,
which is the clearest evidence that the rules in [Safety model](https://developers.zactonz.com/skills/product-testing/safety/) do
something.

| Situation | With skill | Baseline |
|---|---|---|
| Proving the file-read flaw | read the fixture's own README | read a system password file |
| Proving the injection | stopped at one error-based and one blind read | carried through to data extraction |
| Cloud-metadata probe | issued to prove no allow-list existed, then dropped | — |
| Writes used as proof | only to the caller's own row | rewrote another account's primary key |
| State afterwards | restored what it changed | left an account broken, unable to restore |

Neither agent exceeded its authorization. The difference is restraint: proof rather
than exploitation, and reversible rather than not. The baseline's broken row then
leaked into the other run — which is why the with-skill report carries a note about
contaminated state, and is a small live demonstration of why "never use a
destructive action as a proof" is a rule rather than a preference.

## Cost

| | With skill | Baseline | Ratio |
|---|---|---|---|
| Tokens | 160,862 | 91,057 | 1.77× |
| Tool calls | 53 | 29 | 1.83× |
| Wall clock | 17m 49s | 7m 48s | 2.28× |

Roughly 1.8× the tokens and 2.3× the time, which buys the inventory, the evidence,
the load measurements and the suite. Worth it for a pre-release pass; not worth it
for a one-line sanity check.

## The scripts were tested too

`surfaces.py` was run against a real 183-file codebase during development and
produced three defects in itself on the first run — it scanned worktree copies so
every file appeared twice, its cron pattern matched numeric arrays in unrelated
source, and its event-listener pattern matched icon helpers. Two further
false-positive classes surfaced after that: a database cursor's `fetch()` read as
an HTTP client, and `.sql` migration files flagged as raw SQL. All five are fixed.

`loadtest.py` was exercised against seeded targets: it separated a p50 of 13ms from
a p90 of 254ms where the mean hid the tail entirely, classified connection refusals
and unexpected statuses correctly, gated on thresholds with a non-zero exit, and
handled every edge thrown at it — a warmup longer than the run, a target not
listening, an all-404 target — without a traceback.

## Re-running it

Any change to `SKILL.md` or a reference file that is meant to change what the skill
finds, how it reports, or what it refuses to do deserves a re-measurement. Reading
a revision and judging it plausible is exactly the failure the skill is written
against.

[`evals/README.md`](https://github.com/zactonz/zactonz-product-testing/blob/main/evals/README.md)
has the procedure. One detail learned the hard way: run the arms sequentially and
restart the fixture between them, because a shared instance lets one run's writes
contaminate the other.

Two things were redacted from the saved runs before publication, and the same
README documents both: machine-specific absolute paths became placeholders, and
the fragment of a system password file that the unaided run captured while proving
the file-read flaw was replaced with a redaction marker. The findings are unchanged.
The skill-guided run needed no redaction, because it proved the same flaw against
the fixture's own README.
