# Claude QA Plan — Queue North Website ``` Status: Current Owner: _null Last reviewed: 2026-08-18 Governs: what a QA round consists of Review trigger: Any new user-facing surface, or a defect class that got through ``` > The playbook. What a round *is*, so two rounds are comparable and a gap is > visible rather than assumed covered. ## Before a round - Build from a clean checkout at a known SHA, and record that SHA. - Run from a detached worktree if other work is in flight, so uncommitted changes cannot contaminate what is under test. - Note the environment: device, OS version, browser, screen size — whatever the product's behaviour actually depends on. ## The passes Each pass gets a letter, so `ClaudeQACoverage.md` can report per pass and a skipped one is visible. | Pass | What it covers | | --- | --- | | A | First run: install or load, cold start, permissions, empty states | | B | The core flow, end to end, as a real user would do it | | C | The core flow with things going wrong: no network, denied permission, invalid input | | D | Persistence: quit and return, background and resume, restart | | E | The end of the loop — the state that is hardest to reach on purpose | | F | Accessibility: keyboard only, screen reader labels, contrast, text scaling | | G | Performance under the load this product will actually see | | H | Abuse, and what a stranger can reach — **rewritten for this project, see below** | ### Why H is separate from B, here The template's pass H is about authorization: authenticated is not owning, and owning is not permitted. **None of that applies to this project** — there is no login, no session, no role and no per-user row. Every one of those checks would be a box ticked against nothing. What does apply is the other half, and it is the most exposed thing here: two unauthenticated `POST` endpoints that write to a database, on a public origin, behind nothing but a rate limiter and a reCAPTCHA score. Pass B submits those forms as a real visitor would. Pass H is somebody who is not being polite: - both endpoints called directly, off the form, with hand-made payloads - the rate limiter driven past `RATE_LIMIT_PER_MINUTE`, and the 429 body read - the honeypot field and the reCAPTCHA path bypassed deliberately - the 1 MB body limit and the 30 s request timeout tripped on purpose - a duplicate email driven into the 409, to confirm the Zoho forward still fires - **`dist/` inspected** — `VITE_RECAPTCHA_SITE_KEY` is inlined into the bundle at build time, which is correct for a *site* key and catastrophic for a secret one. `scripts/secrets.sh --built dist/` is the mechanical half of this `scripts/preflight.sh` covers headers and TLS and nothing else. Its `--auth` checks stay off: there is no login to rate-limit and no account to enumerate. **Pass I — money flowing backwards — is deleted from this plan rather than carried as skipped.** No money moves through this site: no checkout, no subscription, no refund path. A pass that never applies is noise, and one that is always skipped is a lie. ## Measure it before you file it Batches 10 and 11 were twenty UI and accessibility issues written from reading source. **Seven of ten misstated their own evidence** — two contrast figures were wrong, one proposed a colour measuring 1.96:1 against the background it would sit on, one asked for a change that would have spread a Level A failure, one described a clipping ancestor that does not exist, one an overlap that is a 12px gap, one an attribute that is not in the file. None of that was carelessness about *whether* something was wrong. It was confidence about *how much*, without measuring. So, before filing a UI defect and before acting on one: ```bash node scripts/qa-browser.mjs # production, five pages, five widths node scripts/qa-browser.mjs --url http://localhost:3099 --viewports 320,768 ``` Contrast is arithmetic — composite the colour over its background and compute the ratio, do not judge it by eye. Overlap is two rectangles. "Does it clip" is a computed style you can read off the ancestors. A filed defect's numbers are a claim to check, not a measurement. ## What counts as a finding A finding needs: what was done, what happened, what should have happened, and the build SHA. Without the SHA it cannot be re-tested, and a finding that cannot be re-tested cannot be closed. ## Severity Findings are filed as issues, labelled: - **P0** — ships broken, or loses data - **P1** — materially wrong, but shippable - **P2** — cosmetic or low impact - **release-blocker** — a release built today would be wrong rather than merely incomplete Exactly these label names: the Command Center queries them by name, and a repository that spells them differently has its defects reported as *not adopted* rather than counted wrongly. This one did, until 2026-08-18 — the labels read `P0 Critical`, `P1 High` and `P2 Medium`, and 205 issues reported as *not adopted* rather than as 87% complete. **Do not use `P3`.** It exists on 21 closed issues from before the convention and is frozen. It is not one of the four names anything queries, so a defect filed `P3` today is counted by nothing. Severity is what it costs, not how annoying it is to fix. ## After a round File each finding as a labelled issue. Update `ClaudeReport.md`'s run-state block and its overall sentence, and `ClaudeQACoverage.md` with what each pass actually reached. A pass that could not be run is recorded as blocked, with what blocks it — never quietly left out, which reads identically to "passed". Then **push, and reconcile**. The verdict on the project screen at privacyllc.dev is read out of `ClaudeReport.md` in the pushed repository, so a round whose report is committed but not pushed — or pushed but not reconciled — leaves a stakeholder reading the previous round's judgment with no indication that a newer one exists. The rest of the cycle is in `docs/WORK_CYCLE.md`.