Queue-North-Website/docs/qa/ClaudeQAPlan.md

7.4 KiB

Claude QA Plan — Queue North Website

Status: Current
Owner: _null
Last reviewed: 2026-08-18
Governs: what a QA round consists of
Review trigger: Any new user-facing surface, or a defect class that got through

The playbook. What a round is, so two rounds are comparable and a gap is visible rather than assumed covered.

Before a round

  • Build from a clean checkout at a known SHA, and record that SHA.
  • Run from a detached worktree if other work is in flight, so uncommitted changes cannot contaminate what is under test.
  • Note the environment: device, OS version, browser, screen size — whatever the product's behaviour actually depends on.

The passes

Each pass gets a letter, so ClaudeQACoverage.md can report per pass and a skipped one is visible.

Pass What it covers
A First run: install or load, cold start, permissions, empty states
B The core flow, end to end, as a real user would do it
C The core flow with things going wrong: no network, denied permission, invalid input
D Persistence: quit and return, background and resume, restart
E The end of the loop — the state that is hardest to reach on purpose
F Accessibility: keyboard only, screen reader labels, contrast, text scaling
G Performance under the load this product will actually see
H Abuse, and what a stranger can reach — rewritten for this project, see below

<Add, remove and rename to fit. A pass that never applies is noise; a pass that is always skipped is a lie.>

Why H is separate from B, here

The template's pass H is about authorization: authenticated is not owning, and owning is not permitted. None of that applies to this project — there is no login, no session, no role and no per-user row. Every one of those checks would be a box ticked against nothing.

What does apply is the other half, and it is the most exposed thing here: two unauthenticated POST endpoints that write to a database, on a public origin, behind nothing but a rate limiter and a reCAPTCHA score. Pass B submits those forms as a real visitor would. Pass H is somebody who is not being polite:

  • both endpoints called directly, off the form, with hand-made payloads
  • the rate limiter driven past RATE_LIMIT_PER_MINUTE, and the 429 body read
  • the honeypot field and the reCAPTCHA path bypassed deliberately
  • the 1 MB body limit and the 30 s request timeout tripped on purpose
  • a duplicate email driven into the 409, to confirm the Zoho forward still fires
  • dist/ inspectedVITE_RECAPTCHA_SITE_KEY is inlined into the bundle at build time, which is correct for a site key and catastrophic for a secret one. scripts/secrets.sh --built dist/ is the mechanical half of this

scripts/preflight.sh covers headers and TLS and nothing else. Its --auth checks stay off: there is no login to rate-limit and no account to enumerate.

Pass I — money flowing backwards — is deleted from this plan rather than carried as skipped. No money moves through this site: no checkout, no subscription, no refund path. A pass that never applies is noise, and one that is always skipped is a lie.

Measure it before you file it

Batches 10 and 11 were twenty UI and accessibility issues written from reading source. Seven of ten misstated their own evidence — two contrast figures were wrong, one proposed a colour measuring 1.96:1 against the background it would sit on, one asked for a change that would have spread a Level A failure, one described a clipping ancestor that does not exist, one an overlap that is a 12px gap, one an attribute that is not in the file.

None of that was carelessness about whether something was wrong. It was confidence about how much, without measuring.

So, before filing a UI defect and before acting on one:

node scripts/qa-browser.mjs                       # production, every page in its sitemap, four widths
node scripts/qa-browser.mjs --url http://localhost:3099 --viewports 320,768
node scripts/device-sweep.mjs                     # every page, ten emulated phones and tablets

Contrast is arithmetic — composite the colour over its background and compute the ratio, do not judge it by eye. Overlap is two rectangles. "Does it clip" is a computed style you can read off the ancestors. A filed defect's numbers are a claim to check, not a measurement.

A resized desktop window is not a device (added 2026-09-10)

Two defect classes reached production through a QA round that looked thorough, and both were invisible to the instrument being used.

A window sized to 768 is not 768. It has a scrollbar, so the layout viewport is about 753, so the md breakpoint never engages and the desktop layout the check is about is never on screen. That is how #214, the header CTA clipped at iPad portrait, was fixed, checked at "768", released, and was still 25px past the right edge on every page. A device profile has no scrollbar inset. Check a breakpoint on a device, or on an emulated one; never on a window you dragged.

document.scrollWidth is not evidence of fitting. This site's body carries overflow-x: hidden, so content sliced off the right edge leaves scrollWidth === innerWidth and no scrollbar to hint at it. Compare a box against its nearest clipping ancestor, which is what scripts/lib/css-audit.js does. The first sweep found 21 blocking and 530 high findings on pages that had passed every earlier round.

Touch targets are the other class that got through, for a related reason: a 17px-tall footer link is fine to click and fiddly to tap, and nothing in a desktop pass distinguishes them. The rule that came out of it is in design/OVERHAUL_PLAN.md under Tap targets.

What counts as a finding

A finding needs: what was done, what happened, what should have happened, and the build SHA. Without the SHA it cannot be re-tested, and a finding that cannot be re-tested cannot be closed.

Severity

Findings are filed as issues, labelled:

  • P0 — ships broken, or loses data
  • P1 — materially wrong, but shippable
  • P2 — cosmetic or low impact
  • release-blocker — a release built today would be wrong rather than merely incomplete

Exactly these label names: the Command Center queries them by name, and a repository that spells them differently has its defects reported as not adopted rather than counted wrongly. This one did, until 2026-08-18 — the labels read P0 Critical, P1 High and P2 Medium, and 205 issues reported as not adopted rather than as 87% complete.

Do not use P3. It exists on 21 closed issues from before the convention and is frozen. It is not one of the four names anything queries, so a defect filed P3 today is counted by nothing.

Severity is what it costs, not how annoying it is to fix.

After a round

File each finding as a labelled issue. Update ClaudeReport.md's run-state block and its overall sentence, and ClaudeQACoverage.md with what each pass actually reached. A pass that could not be run is recorded as blocked, with what blocks it — never quietly left out, which reads identically to "passed".

Then push, and reconcile. The verdict on the project screen at privacyllc.dev is read out of ClaudeReport.md in the pushed repository, so a round whose report is committed but not pushed — or pushed but not reconciled — leaves a stakeholder reading the previous round's judgment with no indication that a newer one exists. The rest of the cycle is in docs/WORK_CYCLE.md.