# Product Testing

> A Claude Code skill that tests a product the way a QA engineer would: it inventories every surface, attacks it across functionality, UI, load and security, and leaves behind a runnable test suite and an evidence-backed findings report.

- **Type:** Claude Code skill (Skills)
- **Version:** 1.0.0, released 7 Oct 2026
- **Requires:** Claude Code, Python 3.9+ for the bundled scripts, A browser driver for the UI dimension
- **Licence:** Apache-2.0
- **Download:** https://github.com/zactonz/zactonz-product-testing/releases/latest/download/product-testing.skill
- **Source:** https://github.com/zactonz/zactonz-product-testing
- **Documentation:** https://developers.zactonz.com/skills/product-testing/

## Overview

Product Testing is a skill for Claude Code that tests a product the way a QA engineer would. It inventories every surface the product exposes, attacks those surfaces across four dimensions, and leaves behind two things: a runnable test suite in the project's own idiom, and a findings report in which every claim is backed by captured output.

Ask for testing in the ordinary way and the skill triggers on its own. "Test the export endpoint before I ship it", "can the API handle Black Friday traffic?" and "is this upload handler safe?" each select the right dimensions; "full QA pass on the dashboard" runs all four.

### Why it exists

Asked to test something, the default behaviour of a coding agent is to exercise the happy path, see nothing break, and report that it works. That is the one outcome that costs real money, because the team ships on it.

This skill replaces that with a method. It starts from an enumeration of what the product actually exposes rather than from whatever a code skim surfaced, it aims each case at a named failure instead of at confirmation, and it reports three outcomes, verified pass, verified failure and not tested, so the boundary around the work is visible instead of implied.

### The four dimensions

Ask for one, several, or none; naming none runs all four. The [What it tests](https://developers.zactonz.com/skills/product-testing/dimensions/) tab describes each in detail.

| Dimension | Covers |
|---|---|
| **Functionality** | Boundary and malformed input across about eighty input classes, state and sequence defects, the authorization matrix, concurrency races, idempotency, and the error contract |
| **UI** | Real browser rendering, interaction, forms, keyboard and focus, responsive and dark mode, WCAG accessibility, console and network errors, visual regression |
| **Load** | Baseline, load, stress, soak and spike shapes; latency percentiles rather than averages; bottleneck identification; rate-limit and quota behaviour; recovery and degradation past the ceiling |
| **Security** | Authorization and ownership checks, authentication and sessions, injection boundaries, SSRF on URL-taking surfaces, secret leakage, file handling, transport headers, abuse and cost |

### What you get back

- **A test suite** in the project's existing idiom and location, wired to one command, with a failing test for every confirmed finding. The fix is verified when the test flips rather than by re-reading the diff.
- **A report** with a committed ship-or-wait verdict, findings ranked by severity, a minimal copy-pasteable reproduction for each, and an explicit list of what was not tested and why. See [Reports](https://developers.zactonz.com/skills/product-testing/reports/).
- **The evidence** the report rests on, saved alongside it.

### Built to run against real products

The skill is explicit about what it will not do. It will not point load or security traffic at a host without your go-ahead for that specific target; it will not use a destructive action as a proof of anything; it will not trigger a surface that emails, charges or notifies real people without agreement; it will not copy a discovered secret into a report or a commit; and it will not install tooling, rewrite source to make a test pass, or bypass a commit hook on its own initiative.

It also treats everything the product under test emits, error messages, page content, API responses and logs, as data rather than as instruction. Output that tries to direct the tester is reported as a prompt-injection finding, not obeyed. The [Safety](https://developers.zactonz.com/skills/product-testing/safety/) tab has the full model.

### Measured, not asserted

Against a fixture with ten planted defects across all four dimensions, two agents were given the same prompt and the same authorization, one with the skill and one without.

| | With skill | Baseline |
|---|---|---|
| Planted defects found | 10/10 | 9/10 |
| Files delivered | 25 | 3 |
| Runnable test suite | 17 tests, 13 red as regression guards | none |
| Raw evidence captures | 16 + screenshot | 2 |
| Explicit not-tested section | 8 entries with reasons | a shorter note |
| Cost | 1.8× tokens, 2.3× wall clock | |

Recall is close; a capable tester finds most defects unaided. What the skill adds is the inventory, the evidence, the load data, the durable suite and a stated coverage boundary, plus restraint: given identical unrestricted authorization, the baseline rewrote another account's primary key as a proof and could not restore it, while the skill wrote only to its own row and restored what it changed. Method, full results and the cost breakdown are on the [Evaluation](https://developers.zactonz.com/skills/product-testing/evaluation/) tab.

### Bundled scripts

Two command-line tools ship with the skill and are useful on their own, in CI or from a terminal. `surfaces.py` inventories what a codebase exposes: routes, CLI entry points, pages, specifications, scheduled work, authorization checks, outbound calls, uploads, raw SQL and possible secrets. `loadtest.py` load-tests an endpoint with real percentiles, a status histogram and classified errors, and doubles as a CI gate. Both use only the Python 3 standard library. The [Scripts](https://developers.zactonz.com/skills/product-testing/scripts/) tab is the reference.

### Requirements

| Component | Needed for |
|---|---|
| Claude Code | the skill itself |
| Python 3.9 or newer | the two bundled scripts |
| A browser driver | the UI dimension: the session's own browser tools, or Playwright, Cypress or Puppeteer if already in the project |

Nothing else is required. A load generator such as `k6` is used if present and never installed unasked.

### Licence

Apache-2.0. The source is on [GitHub](https://github.com/zactonz/zactonz-product-testing).

## Getting started

### Install

Download the skill package from the latest GitHub release and unpack it into your Claude Code skills folder. The archive contains a single `product-testing/` folder.

```bash
curl -fsSL https://github.com/zactonz/zactonz-product-testing/releases/latest/download/product-testing.skill -o product-testing.skill
unzip -o product-testing.skill -d ~/.claude/skills/
```

Or clone the repository and copy the folder, which is the same operation:

```bash
git clone https://github.com/zactonz/zactonz-product-testing.git
cp -r zactonz-product-testing/skills/product-testing ~/.claude/skills/
```

Use `./.claude/skills/` instead of `~/.claude/skills/` to scope the skill to one project. Either way, start a new Claude Code session afterwards so the skill is picked up.

### Install as a plugin

The repository is also a Claude Code plugin marketplace with one plugin in it, so Claude Code can install and update the skill itself. A plugin resolves through a marketplace, so register the marketplace first; `claude plugin install` on its own will not find it.

```bash
claude plugin marketplace add zactonz/zactonz-product-testing
claude plugin install zactonz-product-testing@zactonz-product-testing
```

The plugin installs at user scope and contains the one skill. Start a new session afterwards, as with the plain install.

### Verify the install

```bash
python3 ~/.claude/skills/product-testing/scripts/surfaces.py --help
python3 ~/.claude/skills/product-testing/scripts/loadtest.py --help
```

Both should print usage. If `surfaces.py` runs but Claude never seems to use the skill, check that `SKILL.md` sits directly inside the skill folder rather than one level down: `~/.claude/skills/product-testing/SKILL.md`.

### Requirements

| Component | Needed for |
|---|---|
| Claude Code | the skill |
| Python 3.9 or newer | the two bundled scripts |
| A browser driver | the UI dimension only, see below |

Nothing else. The scripts use the Python standard library, so there is no install step and nothing to keep up to date.

For the UI dimension the skill uses whatever is already available, in this order: the session's own browser tools; a driver the project already has (Playwright, Cypress, Puppeteer); otherwise it says so and falls back to asserting server-rendered HTML, which checks markup but not rendering. It will not install a browser without asking.

### Run a first pass

Ask in the ordinary way; the skill triggers on its own.

```
test the export endpoint before I ship it
can the API handle Black Friday traffic?
find the edge cases in the invoice form
is this upload handler safe?
full QA pass on the dashboard
```

To scope deliberately, name the dimensions: *run the load and security dimensions against the staging API*. Naming none runs all four.

### What happens, in order

1. **Scope.** It establishes what is under test, the commit or version, the URL, the environment, and settles authorization if load or security work is involved. Expect to be asked before any traffic goes at something live.
2. **Inventory.** It enumerates the product's surfaces rather than testing from whatever a code skim surfaced. This is the step that decides whether the rest is real.
3. **Plan.** A short table of surfaces and the cases chosen for each. On a large or ambiguously scoped product it will show you this before a long run.
4. **Execute.** Dimension by dimension, capturing raw output as it goes.
5. **Suite.** Tests written into the project's existing test location, in its existing idiom, wired to one command.
6. **Report.** A verdict, ranked findings with minimal reproductions, and an explicit list of what was not tested.

### What you get back

- **A report** at `test-reports/<date>-<scope>/report.md` by default, with the evidence beside it. See [Reports](https://developers.zactonz.com/skills/product-testing/reports/).
- **A test suite** where the project's tests already live. Every confirmed finding gets a test that fails today; those are the regression guards, and they flip green as each fix lands, which is a better signal that a fix worked than re-reading the diff.
- **The raw captures** the report rests on, so any claim can be checked.

Evidence files can carry real payloads and screenshots. The skill will say so and ask before committing them; decide whether that directory belongs in git.

### Scoping advice

The skill is worth reaching for on a pre-release pass, after a refactor that touched many surfaces, when inheriting an unfamiliar codebase, or when you need a defensible answer to "is this ready". It costs roughly 1.8× the tokens and 2.3× the time of testing without it (see [Evaluation](https://developers.zactonz.com/skills/product-testing/evaluation/)), which is plainly worth it for those, and not worth it for a one-line sanity check.

On a large product, scope the first run to one dimension and a handful of surfaces rather than asking for everything. A focused pass that finishes beats a broad one that runs out of room, and the inventory from the first run makes the next one cheaper.

## What it tests

Testing splits into four dimensions. Ask for one, several, or none — naming none
runs all four. Each has a reference file the skill loads only when it reaches that
dimension, so asking for one does not pay for the others.

### Scoping a run

Name them directly, or describe what you want and let the wording select:

| You say | Dimensions selected |
|---|---|
| "find the edge cases in the invoice form" | functionality |
| "does the dashboard still work?" | ui |
| "can it handle Black Friday?" | load |
| "is this upload handler safe?" | security |
| "full QA pass before release" | all four |
| "load and security against staging" | load, security |

If a dimension has no surface in your product, the skill says so in the report's
**Not tested** section with the reason and skips it. A pure library has no UI; a
static brochure site has no authorization boundary. That is useful information
rather than a gap — what you do not want is a dimension silently dropped, which is
why it is written down either way.

### Functionality

The dimension that finds the most defects in working software, because its subject
is the gap between what a surface promises and what it does.

- **Boundary and malformed input** across roughly eighty input classes — absence
  and emptiness, numeric edges and overflow, string limits and whitespace, Unicode
  that breaks naive truncation and uniqueness checks, wrong types and shapes,
  identifiers that do not exist or belong to someone else, date and timezone edges,
  file and encoding cases. The discipline is working through the catalogue per
  input rather than relying on the three cases that come to mind, because those are
  the ones already handled.
- **State and sequence** — the empty state, the single-item state, the large state,
  replayed submits, out-of-order operations, and interrupted writes, where the
  question is what got left behind.
- **The authorization matrix**, built as an explicit grid of callers against
  surfaces, so the missing cells are visible rather than implied.
- **Concurrency and idempotency** — simultaneous creates against a unique key, lost
  updates, counter and quota races, and retries of requests whose response was lost.
- **The error contract** — right status code, actionable message, nothing leaked,
  consistent shape, and nothing half-applied after a failure.

It also insists on an oracle: a documented contract, a schema, consistency with
the product's own siblings, or a labelled judgement call. A finding with no basis
for calling the behaviour wrong is reported as an unspecified behaviour, which is
honest and still useful.

### UI

Verifies what a person sees and can do in a real browser, rather than what the
markup suggests they would.

- **Rendering** — no blank page, no error overlay, no leaked template syntax, no
  visible `undefined` or `NaN`, and the real content present rather than just the
  shell
- **Every state** — loading, empty, error and populated. The empty state is the
  most-skipped screen in software and the first one a new user sees
- **Interaction** — primary actions and their consequences, double-clicked submits,
  clicking during an in-flight request, keyboard-only operation, and focus
  behaviour around modals
- **Forms** — empty submit, each field invalid in turn, paste and autofill rather
  than typing, and whether a failed submit preserves what was typed
- **Responsive and theme** — 375px, 768px and desktop; light and dark, including
  the system preference rather than only an in-page toggle
- **Accessibility** — automated scanning where a scanner is available, plus the
  things no scanner judges: contrast, focus visibility, alt text that conveys
  meaning, heading structure, real labels rather than placeholders, and whether
  dynamic content is announced
- **Console and network** — uncaught exceptions on pages that look correct, 4xx and
  5xx including assets, and anything sensitive in a URL

It reads pages as text to assert, and screenshots for proof — cheaper and more
precise for anything textual, with images reserved for what only an image settles.

### Load

Finds the point where the product stops meeting its promise under pressure, and
characterises what happens past it. Every product has such a point; the question
is whether you found it or a customer did.

Four shapes, named separately because they answer different questions:

| Shape | Question |
|---|---|
| **Load** | Does it hold up under the traffic it is built for? |
| **Stress** | Where is the ceiling, and what breaks first? |
| **Soak** | Does it leak? Memory, descriptors, pool exhaustion — invisible in a minute |
| **Spike** | Does it survive the transient, and does it *recover*? |

Method that matters: a single-worker baseline first, one variable at a time, warmup
discarded, long enough for garbage collection and cache expiry to appear, and the
environment recorded alongside every number so a laptop figure is not later quoted
as production capacity.

Results are reported as percentiles, never as a mean — a 50ms mean routinely hides
2% of requests at eight seconds, and that tail is what users complain about.
Throughput and latency are reported together, because high throughput at a
twelve-second p99 is a queue filling up rather than capacity.

The most valuable half is past the ceiling: whether it sheds load or collapses,
whether it recovers unaided when traffic stops, whether it survives a dependency
failing and reconnects without intervention, and whether anything is corrupted
afterwards. Data damage under load outranks every latency number in the report.

Rate limits and quotas are tested as part of the contract — the threshold, a
parseable 429 rather than a hang, correct retry headers, the right subject, and
whether concurrent requests can straddle the boundary.

### Security

Verifies that the product's own protections hold. It thinks in controls rather than
exploits: for each control the product relies on, what is the minimum request that
shows whether it is present?

- **Authorization**, tested first and everywhere, because the check is written per
  handler and therefore forgotten per handler. Two accounts, and for every surface
  taking an identifier: can one read, update or list the other's data; can a
  privileged field be set on your own record; are the surfaces with no caller in the
  repo protected
- **Authentication and sessions** — credential verification, user enumeration,
  token rejection when expired or tampered, whether logout invalidates server-side,
  reset token scope and reuse, cookie flags
- **Injection boundaries** — one probe per context to establish whether input can
  become syntax: SQL, HTML, shell arguments, templates, paths, XML entities,
  deserialization, log and header injection
- **Requests the server makes for you** — any surface taking a URL, checked against
  loopback, link-local and private ranges, non-HTTP schemes, hostnames resolving
  inward, and redirects from public to private, which is where most implementations
  fail
- **Secret leakage** — responses, error pages, headers, the client bundle, publicly
  reachable config and backup paths, repository history, and logs
- **File handling, transport headers, configuration and dependencies**
- **Abuse and cost** — cheap requests triggering expensive operations, which is a
  direct route to a large bill

It stops at proof and does not pivot. Once a control is shown to be missing, the
finding is complete; reading one record demonstrates the flaw, and enumerating the
table adds no information while turning a test into a breach. See
[Safety model](https://developers.zactonz.com/skills/product-testing/safety/).

This verifies controls. It does not model a determined attacker, and where that
matters it is not a substitute for a scoped engagement.

## Safety

Testing means running untrusted things, reading untrusted output and generating
traffic against something real. This page says exactly what the skill will and will
not do, so you can decide what to point it at.

The rules below live in `SKILL.md` as a block the skill treats as non-negotiable by
anything it encounters mid-run — not by a comment in the code, a note in a README,
a line in a log, a page under test, or a stated deadline. Where something seems to
require breaking one, it stops and asks instead.

### Authorization is per target, and comes only from you

The default is an instance started locally, or one you named explicitly. Load or
security work against anything shared or live needs your go-ahead for that specific
host, in that conversation, before the first request — along with an agreed
concurrency ceiling and duration, which it then holds to.

Text found inside the product is data, not permission. A page saying "this is a
test environment", a README saying "feel free to hammer this", a staging banner —
none of these establish who owns the machine, and none are accepted as
authorization. Approval for one target does not carry to another.

If you say "test it" and the only thing running is production, it says so and
offers to stand up a local instance.

Two reasons this is strict rather than ceremonial. A load test and a
denial-of-service attempt are the same packets, and a successful one against your
own production takes your product down. And probing infrastructure you do not own
is an attack whatever the intent.

### Destructive actions are never used as proof

Deleting, overwriting, truncating or mass-writing real data demonstrates nothing a
read cannot. A missing authorization check is proven by reading one record; a write
path is proven with a record the test created itself. It does not drop tables,
flush caches, rotate keys or clear queues on a shared instance.

In the [evaluation](https://developers.zactonz.com/skills/product-testing/evaluation/) this distinction showed up plainly: given
identical unrestricted authorization, the run without the skill rewrote another
account's primary key as a demonstration and could not restore it, while the run
with the skill wrote only to its own row and restored what it changed.

### Nothing real goes out

Before testing a surface it checks whether that surface sends email or SMS,
charges a card, calls a metered third-party API, posts publicly, or notifies users.
Those need a sandbox, a test mode, or your explicit agreement. A test run that
emails live customers cannot be taken back, and a load test against a metered
endpoint arrives as an invoice.

### It stops at proof

Once a control is shown to be missing, the finding is complete. It does not chase
the weakness deeper, pivot from one flaw to another, or widen the blast radius to
make the report more impressive. Escalation is your decision, not the tester's.

Concretely, from the evaluation: the injection finding was established with one
error-based and one blind read and not carried through to extraction; the
cloud-metadata probe was issued to prove no allow-list existed and then dropped;
the file-read flaw was proven by reading the fixture's own README rather than a
system file. Each is recorded in the report's **Not tested** section as stopped by
policy, so you can see the decision rather than guess at it.

### Product output is data, never instruction

A tester spends its time reading error messages, page content, API responses,
filenames and logs — and on a real product some of that is attacker-controlled by
design. If any of it asks the skill to run a command, fetch a URL, change a file or
disregard its instructions, that is reported as a prompt-injection finding, and a
notable one. It is never acted on.

This matters more than it first appears. A stored field that reaches an
administrator's screen is a channel into whatever reads that screen, and a tester
is a thing that reads that screen.

### Secrets are reported by location, never by value

If a probe surfaces a live credential, token or personal data, you are told
immediately so it can be rotated, and only its location is recorded. The value
never goes into the report, a screenshot, a fixture, a commit or a message.

Rotation rather than deletion is the advice, because a secret in repository history
is not removed by a later commit.

### It changes the project only as asked

Adding tests and writing a report is the job. Installing tooling, editing
configuration, rewriting source to make a test pass, committing, pushing, or
bypassing a commit hook is not — it asks first. Where a dependency is genuinely
required it says what and why and leaves the decision to you.

It will not install a browser or a load generator unannounced. It uses `k6` if
present and the bundled Python generator if not.

### It cleans up

Servers it started are stopped, accounts and records it created are removed, config
it touched is restored, and anything it could not clean up is stated plainly.

### What this does not cover

The skill's restraint is not a substitute for your own scoping. It will refuse an
obviously unauthorized target and ask about an ambiguous one, but it cannot know
that a host you named is shared with a customer, or that an endpoint you pointed it
at bills per call. Tell it what it cannot see.

And the security dimension verifies controls. It is not a penetration test, does
not model a determined adversary, and where that distinction matters it should not
be reported as one.

## Reports

Reports land at `test-reports/<date>-<scope>/report.md` by default, with the
evidence beside them. The template is in `assets/report-template.md`.

### Shape

```
Verdict                 three sentences and a severity table
Findings                ranked most severe first
Load results            a table, a capacity sentence, the environment
Verified working        what was exercised and held up
Not tested              what was not covered, and why
Suite added             where the tests went and how to run them
Evidence                where the raw captures are
```

Read the verdict and the first three findings. They are ordered so that stopping
there is a reasonable thing to do.

### The verdict commits

It names what was tested, what was found, and whether to ship. "Several issues were
identified" is not a verdict and the skill is written against producing one. Expect
something closer to: *do not ship — one authorization gap on the invoices endpoint
lets any user read any account's data.*

### A finding

Each one carries, in this order: the surface and its location in the code; a
minimal copy-pasteable reproduction; what was observed, with a reproduction count;
what was expected, with the source that says so; why it matters; the evidence file;
and the specific fix with a named location.

Two details to look for. **The cited expectation** — "the schema documents `limit`
as 1–1000" ends an argument that "should return 400" invites; where no documented
expectation exists, the finding is labelled a judgement call so you can overrule it
in seconds. And **the reproduction count** — "reproduced 3/3" and "reproduced 1/10"
lead to different decisions, and an intermittent failure is worth knowing about
precisely because it is the kind that wakes people up.

### Severity

Judged on impact and reachability, not on how alarming the defect's name is.

| Level | Means | Action |
|---|---|---|
| **Critical** | Data loss or corruption, an authentication or authorization bypass, secret exposure, the product unusable for everyone, or money moved incorrectly — reachable by an ordinary or unauthenticated caller | Do not ship |
| **High** | A core promise broken for many users with no workaround, a crash on ordinary input, or a silent failure of something the user believes succeeded | Fix before shipping |
| **Medium** | Real breakage that is narrow, recoverable, or has a workaround; degradation above expected traffic; missing hardening with no current exploit path | Schedule |
| **Low** | Cosmetic, a confusing message, or a latent problem with no present impact | Backlog |

Two deliberate adjustments. **Silent wrongness rates above loud failure** — a crash
gets noticed and fixed, while a wrong number or a half-saved form propagates for
months, so silent incorrectness sits a level higher than a visible failure of
similar scope. And **reachability cuts both ways** — behind an admin-only surface,
a level down; reachable unauthenticated, a level up.

Where a finding sits between two levels it takes the lower one and says why it
could be argued higher. That keeps the scale meaning something across reports.

### The section that matters most

**Not tested** is what makes the rest of the report trustworthy, and it is the first
thing to read if you are deciding how much weight to put on a clean result.

A reader who can see the boundary knows what the silence elsewhere means. A reader
who cannot will assume total coverage. Expect it to be long, to distinguish *not
applicable* ("no upload surface exists") from *not reached* ("one browser engine
only; keyboard and dark-mode depth not exercised"), and to record anything that
compromised the run — a contaminated environment, a missing credential, an agreed
window that ran out.

A report with four findings and an eight-entry not-tested list is more useful than
one with fifteen findings and no stated boundary.

### When nothing broke

"No defects found" is worth something only alongside what was attempted. A clean
report should still show the surfaces exercised as a count over the inventory, the
case classes run, and — the most informative line in it — what was expected to break
and did not. A clean report with no visible attempt is indistinguishable from no
testing, and should be treated that way.

### The three outcomes

Every case is a verified pass, a verified failure, or not tested. Reports lie almost
exclusively by collapsing the third into the first, letting an untested area read as
silence that the reader hears as fine.

Before handing over, the skill runs a checklist: every claimed pass corresponds to
captured output rather than to code it read; every finding was reproduced at least
twice; every reproduction was re-run from a clean state and still fails; the
not-tested list is complete; no secret values appear anywhere; the build and
environment are recorded; and anything uncertain is labelled uncertain rather than
smoothed over.

If you find a claim in a report you cannot trace to an evidence file, treat that as
a defect in the report and say so — the whole method depends on that being
checkable.

## Scripts

Two tools ship with the skill. Both use only the Python 3 standard library, so
there is no install step, and both are useful on their own — in CI, or from a
terminal, without Claude involved.

They live at `skills/product-testing/scripts/`, or
`~/.claude/skills/product-testing/scripts/` once installed.

### surfaces.py — surface inventory

Greps a pattern library across common web, API, CLI and job frameworks and groups
what it finds, so a test pass starts from what the product exposes rather than from
whatever a code skim surfaced.

```bash
python3 surfaces.py [root] [--out PATH] [--json] [--limit N]
```

| Option | Effect |
|---|---|
| `root` | Project root. Defaults to the current directory |
| `--out PATH` | Write the report to a file as well as stdout |
| `--json` | Emit JSON instead of Markdown |
| `--limit N` | Rows shown per category, default 40. Totals are always reported in full |

#### What it reports

The detected stack, any test harness already present, file-based routes, and then
one section per category: HTTP routes and request entry points, API
specifications, pages and views, CLI entry points, scheduled jobs and queues,
webhooks and event handlers, authorization boundaries, outbound requests,
file-upload handling, raw SQL, and possible hardcoded secrets.

The harness detection is worth reading first — a suite written in a second style is
a suite nobody runs, so it tells you what idiom to match.

#### Reading it honestly

It is a pattern scan, so it misses dynamically registered routes and anything
behind indirection, and it over-reports in a few categories by design. Three
sections need triage rather than being taken at face value:

- **Authorization boundaries** lists checks that *exist*. The finding is the
  surfaces with none, which you get by comparing against the route list.
- **Outbound requests** lists every HTTP client call. Only the ones whose URL comes
  from request input are candidates for server-side request forgery.
- **Possible hardcoded secrets** will match fixtures and examples. Verify before
  reporting, and never copy a live value anywhere.

It skips the directories you would expect — `node_modules`, `vendor`, build output,
caches — along with hidden directories other than `.github`, and worktree copies,
which otherwise scan the whole repository twice.

```bash
# Typical first move on an unfamiliar codebase.
python3 surfaces.py . --out surfaces.md --limit 100
```

### loadtest.py — concurrent load generator

Exists because `k6`, `wrk`, `hey` and `vegeta` are usually not installed, and
ApacheBench reports no usable percentiles. Use `k6` instead where it is available;
this covers everywhere else.

```bash
python3 loadtest.py URL [options]
```

#### Shaping the run

| Option | Effect |
|---|---|
| `-c, --concurrency N` | Concurrent workers, default 10 |
| `-n, --requests N` | Total requests to send |
| `-d, --duration SEC` | Sustain traffic for this long. Mutually exclusive with `-n` |
| `--rps N` | Cap total arrival rate, modelling real traffic rather than a closed loop of workers |
| `--ramp SEC` | Stagger worker start over this period, instead of hitting cold caches with a wall |
| `--think SEC` | Pause between a worker's requests, modelling users rather than bots |
| `--warmup SEC` | Discard results from the first N seconds, so JIT and connection setup do not pollute the numbers |

#### The request

| Option | Effect |
|---|---|
| `-X, --method` | HTTP method, default GET |
| `-H, --header 'K: V'` | Repeatable |
| `-b, --body DATA` | Body, or `@path` to read from a file |
| `--timeout SEC` | Per-request timeout, default 10 |
| `--expect-status LIST` | Comma-separated codes to count as success. Default is anything under 400 |
| `--insecure` | Skip TLS verification |

#### Output and gating

| Option | Effect |
|---|---|
| `--json PATH` | Full summary, including the per-second timeline |
| `--csv PATH` | Per-request records, for plotting the tail yourself |
| `--fail-over-errors PCT` | Exit 1 if the error rate exceeds this |
| `--fail-over-p95 MS` | Exit 1 if p95 exceeds this |
| `-q, --quiet` | Suppress progress output |

The two `--fail-over-*` flags make it a CI gate: it exits 1 on breach and names
what breached.

#### The safety guard

Non-local targets are refused unless `--allow-remote` is passed, so aiming load at
a remote host is always a deliberate act rather than a typo. Local means loopback,
`127.0.0.0/8`, `::1`, or a name ending `.localhost` or `.test` — a private-range
address is still somebody's server and needs the flag.

Before passing it: a load test and a denial-of-service attempt are the same
packets. Only target hosts you own or are authorized to test, agree a ceiling
first, and check what the endpoint does before hitting it ten thousand times —
mail, payments and metered third-party calls all arrive as consequences.

#### What it reports

Request and error counts, error rate, throughput, bytes received; latency as min,
mean, p50, p75, p90, p95, p99 and max; a status-code histogram; errors grouped by
class (`timeout`, `connection_refused`, `connection_reset`, `dns_failure`,
`tls_error`, `unexpected_status_<code>` and so on); and a per-second timeline of
requests, errors and p95.

Percentiles are the point. On a seeded tail during development it reported a p50 of
13ms against a p90 of 254ms, where the 60ms mean hid the tail completely.

A run that measured nothing, or in which nothing succeeded, prints a warning rather
than a page of zeros that could be mistaken for a pass.

#### Worked sequence

```bash
# 1. Baseline. One worker, serial. Everything later is compared to this.
python3 loadtest.py http://127.0.0.1:8080/health -c 1 -n 50 --json base.json

# 2. Load. Expected peak concurrency, sustained, warmup discarded.
python3 loadtest.py http://127.0.0.1:8080/v1/thing -c 25 -d 60 --warmup 5 \
  -X POST -H 'Content-Type: application/json' -b @payload.json \
  --expect-status 200,201 --json load-25.json

# 3. Stress. Climb until something gives; the knee is the real capacity.
for c in 10 25 50 100 200; do
  python3 loadtest.py http://127.0.0.1:8080/v1/thing -c $c -d 30 --json "stress-$c.json"
done

# 4. As a CI gate.
python3 loadtest.py http://127.0.0.1:8080/v1/thing -c 20 -d 30 \
  --fail-over-errors 1 --fail-over-p95 500 -q
```

Record the environment with any number you keep — machine, CPU count, whether the
database is local, whether it is a debug build. Load figures get quoted later as
production capacity, and without that context they are not.

## Evaluation

A skill is easy to judge by reading and hard to judge by reading, so this one was
measured. Everything below is self-contained, and the material behind it is
public: the fixture, the answer key, the scoring script and both saved runs are in
the repository's [`evals/`](https://github.com/zactonz/zactonz-product-testing/tree/main/evals)
folder, with the written comparison in
[`evals/RESULTS.md`](https://github.com/zactonz/zactonz-product-testing/blob/main/evals/RESULTS.md).
The fixture is a deliberately vulnerable application; run it bound to loopback
only.

### Method

A deliberately defective API was built with ten planted defects spread across all
four dimensions — a missing ownership check, mass assignment, SQL injection,
reflected XSS, unauthenticated server-side request forgery, a debug-mode stack
trace, a non-atomic quota counter, byte-truncation that splits multibyte
characters, an unbounded row limit, and five UI faults on a single page.

Every defect was verified to reproduce before the runs started, so the answer key
is observed rather than assumed.

Two agents then received the same prompt against the same running instance — one
pointed at the skill, one with no skill at all — both told they had full
authorization to test every dimension. Neither could read the answer key or edit
the fixture. Scoring requires a defect's **surface and its mechanism** to appear
together before counting a hit, so "validation could be tightened" does not score
as having found the unbounded-limit defect.

Caveats worth stating: one fixture, one run per arm, and a fixture smaller than a
real product. Treat recall as indicative and the structural differences as the
substantive result.

### Recall

**10/10 with the skill, 9/10 without.** The one the baseline missed was the
unbounded row limit.

Recall is close, and that is the honest headline. On a fixture this size a capable
tester finds most of what is there unaided. Recall is not where the skill earns its
place.

### What each run produced

| | With skill | Baseline |
|---|---|---|
| Files delivered | 25 | 3 |
| Report and evidence | 82,921 chars | 9,524 chars |
| Surface inventory | yes | no |
| Written test plan | yes | no |
| Raw evidence captures | 16, plus a UI screenshot | 2 |
| Load data | 5 scenarios as JSON, with a results table and stated environment | none |
| Runnable test suite | 17 tests, stdlib only, 13 red as regression guards | none |
| Committed ship verdict | yes | no |
| Explicit not-tested section | yes — 8 entries with reasons | a shorter coverage note |

Two of these are the durable value. The **suite** outlives the session: thirteen of
its tests fail against the open findings and flip green as each fix lands. The
**not-tested section** is what makes the report trustworthy — in this run it
separated *not applicable* (no upload surface, no scheduled jobs, no TLS by design)
from *not reached* (one browser engine, dark-mode and keyboard depth not
exercised), and recorded that a parallel run had contaminated shared state, so one
account's tests ran as the other user.

### Behaviour under identical authorization

Both runs had unrestricted permission to test. They did not behave the same way,
which is the clearest evidence that the rules in [Safety model](https://developers.zactonz.com/skills/product-testing/safety/) do
something.

| Situation | With skill | Baseline |
|---|---|---|
| Proving the file-read flaw | read the fixture's own README | read a system password file |
| Proving the injection | stopped at one error-based and one blind read | carried through to data extraction |
| Cloud-metadata probe | issued to prove no allow-list existed, then dropped | — |
| Writes used as proof | only to the caller's own row | rewrote another account's primary key |
| State afterwards | restored what it changed | left an account broken, unable to restore |

Neither agent exceeded its authorization. The difference is restraint: proof rather
than exploitation, and reversible rather than not. The baseline's broken row then
leaked into the other run — which is why the with-skill report carries a note about
contaminated state, and is a small live demonstration of why "never use a
destructive action as a proof" is a rule rather than a preference.

### Cost

| | With skill | Baseline | Ratio |
|---|---|---|---|
| Tokens | 160,862 | 91,057 | 1.77× |
| Tool calls | 53 | 29 | 1.83× |
| Wall clock | 17m 49s | 7m 48s | 2.28× |

Roughly 1.8× the tokens and 2.3× the time, which buys the inventory, the evidence,
the load measurements and the suite. Worth it for a pre-release pass; not worth it
for a one-line sanity check.

### The scripts were tested too

`surfaces.py` was run against a real 183-file codebase during development and
produced three defects in itself on the first run — it scanned worktree copies so
every file appeared twice, its cron pattern matched numeric arrays in unrelated
source, and its event-listener pattern matched icon helpers. Two further
false-positive classes surfaced after that: a database cursor's `fetch()` read as
an HTTP client, and `.sql` migration files flagged as raw SQL. All five are fixed.

`loadtest.py` was exercised against seeded targets: it separated a p50 of 13ms from
a p90 of 254ms where the mean hid the tail entirely, classified connection refusals
and unexpected statuses correctly, gated on thresholds with a non-zero exit, and
handled every edge thrown at it — a warmup longer than the run, a target not
listening, an all-404 target — without a traceback.

### Re-running it

Any change to `SKILL.md` or a reference file that is meant to change what the skill
finds, how it reports, or what it refuses to do deserves a re-measurement. Reading
a revision and judging it plausible is exactly the failure the skill is written
against.

[`evals/README.md`](https://github.com/zactonz/zactonz-product-testing/blob/main/evals/README.md)
has the procedure. One detail learned the hard way: run the arms sequentially and
restart the fixture between them, because a shared instance lets one run's writes
contaminate the other.

Two things were redacted from the saved runs before publication, and the same
README documents both: machine-specific absolute paths became placeholders, and
the fragment of a system password file that the unaided run captured while proving
the file-read flaw was replaced with a redaction marker. The findings are unchanged.
The skill-guided run needed no redaction, because it proved the same flaw against
the fixture's own README.
