15 KiB
Claude QA Coverage — Privacy: Period Tracker
Status: Current
Owner: _null
Last reviewed: 2026-08-20
Governs: what each QA pass actually reached
Review trigger: Any QA round run
Pass by pass, what was reached and what was not. The point of this file is the Blocked and Not run rows: a pass left out of a report reads exactly like a pass that succeeded, and that is how untested code ships believing it was tested.
Targeted Check — 2026-08-20, launcher alias
Not a full QA round. This was the device proof for issue #31 on emulator-5554,
API 34, with the debug package dev.privacyllc.period.debug.
Fresh install resolved MAIN/LAUNCHER to
dev.privacyllc.period.PeriodLauncherAlias. Onboarding was completed through
the UI, Settings showed Discreet launcher off, and the confirmation dialog
said the home screen would show Daybook while warning that the launcher entry
is removed and re-added.
After confirming, cmd package query-activities returned one launchable
activity: dev.privacyllc.period.IncognitoLauncherAlias. Relaunching through
that alias opened the app, and Settings showed the switch on with the copy
"Home screen shows Daybook with a neutral icon." Restoring through the UI
returned the resolver to PeriodLauncherAlias. A second adb install -r
succeeded over the existing install, and pm list packages dev.privacyllc.period
still returned only dev.privacyllc.period.debug.
Round 3 — 2026-08-18 at 0451fbe, partial
Batches 04 and 05 landed: fertility estimates and the reminder system. Pass F became runnable for the first time.
Environment: emulator PeriodQA, API 36, Pixel 6 profile, debug build, plus
6 instrumented tests on the same device — 4 in NotificationPrivacyTest, 2 in
PeriodCrudTest.
| Pass | Result | Notes |
|---|---|---|
| A — First run | Pass | Re-run after the brand change; onboarding reaches a forecast, relaunch skips it. |
| B — Core loop | Pass | Unchanged and re-driven. Fertility appears on Today and the calendar once the forecast is tight enough. |
| C — Failure paths | Partial | As Round 2. Airplane mode, denied notification permission and a killed process still untried. |
| D — Persistence | Partial | As Round 2. Reboot and update-over-install untried — the second matters more now that WorkManager holds scheduled work. |
| E — Forecast under hard histories | Partial | Unit-tested against both engines; the §51 histories still have not been entered by hand. |
| F — Notification privacy | Partial, and this is the important row | Four instrumented tests on a device assert what a lock screen would render: every kind × both private modes attaches a public version and is marked VISIBILITY_PRIVATE, no public title or body carries a health word, Direct is the only mode marked public, no channel name in any of the three modes carries one either, and Discreet cannot pop over the screen. Action labels are checked by the JVM unit tests in NotificationCopyTest, not on the device. Nobody has yet looked at an actual locked screen. That is a real gap: the tests check the notification object, and the last mile is what the system chooses to draw. |
| G — Accessibility | Partial | Unchanged. Calendar verified in real greyscale again after the palette change. TalkBack and font scaling still never run. |
| H — Data ownership | Partial | Unchanged. |
What this round found
Two defects, both in how Android behaves rather than in the app's logic, and both found by running on a device:
- A notification channel is immutable after creation. Importance and lock-screen visibility cannot be changed once set, so a single shared channel would have kept whatever the user's first privacy mode chose — switching from Direct to Maximum privacy would have appeared to work and changed nothing. One channel per mode now.
checkPermissionswas green over its own target, reading a stale merged manifest because it did not depend on the task that writes one. The third guard in this project to fail its first proof.
Also confirmed: a seventeen-day "fertile window" was on screen for a user one cycle in. Arithmetically correct, useless, and only visible by looking.
Round 2 — 2026-08-18 at 19edf4c, partial
Batch 03 built every screen the app has, and each was driven by hand on the
emulator as it landed rather than in one pass at the end. Same environment:
PeriodQA, API 36, Pixel 6 profile, debug build.
| Pass | Result | Notes |
|---|---|---|
| A — First run | Pass | Clean install reaches onboarding; all seven screens walked; the flow produces a forecast and a relaunch goes straight to Today. |
| B — Core loop | Pass | Two-tap logging confirmed on the device: "Started period" → "Yes — today" → "Logged ✓". Editing, ending, and spotting reclassification all exercised. |
| C — Failure paths | Partial | Duplicate-date logging, an end date before a start, and a future date in the picker all handled. Airplane mode, denied permissions and a process killed mid-write still untried. |
| D — Persistence | Partial | Onboarding completion, records and settings all survive force-stop and relaunch. Reboot and update-over-install untried. |
| E — Forecast under hard histories | Partial | The §51 cases pass as unit tests against both engines, and the not-yet path was driven through the UI. The variable and outlier histories have still not been entered by hand. |
| F — Notification privacy | Not run | Nothing sends a notification yet — Batch 05. The onboarding choice was verified: Discreet is selected before the user touches anything. |
| G — Accessibility | Partial | Every calendar day, the confidence indicator and the hero countdown carry content descriptions, checked by reading the view hierarchy. The calendar was verified in actual greyscale and all five marks remain distinguishable. TalkBack itself, and font scaling, still untried. |
| H — Data ownership | Partial | Delete-all covered by an instrumented test. Export, app lock and artifact inspection do not exist yet. |
What driving it found that tests did not
Three defects, none of which any unit test would have caught:
- Dark mode was broken for the whole of Batch 01.
PeriodThemenever wrapped its content in aSurface, so text without an explicit colour inherited black and the app background never painted. Light mode looked correct by accident. - "Period ended" appeared to do nothing. The logic was right — a period ending today still includes today — but the screen was identical afterwards, so the button read as broken.
- Today's underline collided with the spotting dot, on the one day that was both. Invisible in the colour screenshot, obvious in greyscale.
The pattern is now two rounds old and worth stating as a rule: the defects in this project are found by opening it, not by reading it.
Round 1 — 2026-08-18 at adc5075, partial
Not a full round. Batch 01 produced the first build, and this covered the two passes that build could support. Recorded as partial rather than left out, because a pass omitted from a report reads exactly like a pass that succeeded.
Environment: emulator PeriodQA, API 36 (sdk_gphone64_x86_64), Pixel 6
profile, debug build.
| Pass | Result | Notes |
|---|---|---|
| A — First run | Pass | Clean install, cold start, empty state reads correctly ("No forecast yet", "No periods logged yet"), all four tabs reachable. No notification permission prompt yet — none is requested until Batch 05. |
| B — Core loop | Pass | Logged a period; cycle day, countdown, window and confidence appeared. Edited the start date twice; forecast moved from 15 Sep to 13 Sep and the record was marked edited. Deletion covered by instrumented test rather than by hand. |
| C — Failure paths | Partial | Only one path exercised, and it found a crash — see below. Airplane mode, denied permissions and a killed process mid-entry were not tried. |
| D — Persistence and migration | Partial | Covered by an instrumented test that closes and reopens a file-backed database, which is what a force-stop does. A real force-stop, a reboot and an update-over-install were not tried. |
| E — Forecast under hard histories | Not run | The §51 cases pass as unit tests; none has been driven through the UI, which is what this pass is for. |
| F — Notification privacy on a lock screen | Not run | Nothing sends a notification yet — Batch 05. |
| G — Accessibility | Not run | TalkBack, font scaling and greyscale legibility untried. The working surface is not the designed screen, so this is worth deferring to Batch 03 rather than testing a screen that is about to be replaced. |
| H — Data ownership and leakage | Partial | Delete-all covered by an instrumented test. Export, app lock and the built-artifact inspection do not exist yet. |
What pass C found
Tapping Started today twice on the same day killed the app:
FATAL EXCEPTION: main
android.database.sqlite.SQLiteConstraintException: UNIQUE constraint failed:
period_records.startDate
Not filed as an issue: it was found and fixed inside the same batch, before any
build left this repository, and a tracker issue closed by the commit that
introduced the code would be bookkeeping rather than a record. It is written
down here, in the commit that fixed it, and in
../architecture/README.md, with four regression
tests pinning the behaviour.
The interesting part is not the crash. It is that the first thing tried by hand broke immediately, on the app's primary button, with 70 unit tests green — which is the argument for pass C existing at all.
What lint found, that no pass would have
Wiring the boundary guard into ./gradlew check ran Android lint for the first
time, on f5e9fbe:
NewApi: java.time.LocalDate#ofInstant requires API 34 (minSdk is 26)
NewApi: java.time.LocalDate#EPOCH requires API 34 (minSdk is 26)
Both on the recalculation path — a crash on every device below Android 14.
Invisible to the unit tests and invisible to an API 36 emulator. Fixed, and
worth recording as a standing gap below: this project has no test on a device
at its own minSdk.
Round 4 — 2026-08-18 at f43e1c2, partial
Batch 06 began: the designed Settings screen and Delete My Data. Driven by hand
on PeriodMinSdk26, plus the font-scale pass listed under the gaps below.
What driving it found that no test would have
Two defects, both in controls that already looked finished:
- Only the radio button in a privacy option was clickable. Tapping the
label "Maximum privacy" did nothing —
PrivacyRowand onboarding'sPrivacyOptionputonClickon theRadioButtonand left the row inert, so the control that decides what a lock screen shows could only be changed by hitting a ~20dp circle. Found while trying to prove something else: the delete confirmation promises "your reminder settings are unchanged", and testing that claim meant changing a setting first, which would not work. Fixed withModifier.selectableon the row andonClick = nullon the radio, which also merges the semantics for TalkBack. - The navigation bar wrapped its labels mid-word at font scale 2.0 — recorded under the font-scaling gap below.
Neither is visible in a unit test, a preview, or a code review. Both took one person tapping the thing.
What was verified rather than assumed
- Delete My Data erases every period, spotting and prediction record, and the user's notification-privacy choice survives it — set to Maximum privacy, deleted, still Maximum privacy afterwards. That is the claim the confirmation dialog makes to the user, so it is checked on a device rather than argued from the code.
- Today renders its empty state after a deletion without crashing, which is the first-run screen reached from a direction it had never been reached from.
PrivacyViewModelTestcovers the same ground on the JVM, including a second delete while one is running.
Standing gaps
Things no round has ever covered, carried forward until they are. This list existing is not a failure; it not existing while the gaps do is.
- No lock screen has been looked at. Pass F's instrumented tests assert the notification object is built correctly, which is most of the risk and not all of it. What the system actually draws on a locked device — including heads-up behaviour and any launcher's own preview — has never been seen by anybody.
- TalkBack has never actually been run. Content descriptions are written and were checked by reading the view hierarchy, which proves they exist and not that they make sense in sequence. Pass G is not complete until somebody navigates a screen with their eyes shut.
Font scaling untried— done on 2026-08-18, and it found one defect. Driven atfont_scale1.3 and 2.0 onPeriodMinSdk26, through onboarding to Today. The hero number survives: at 2.0 the "14" and its labels still fit, and Today scrolls so nothing below is lost. Every onboarding step survives too, including the two carrying three buttons or three option cards under an illustration — their art is deliberately 104 dp where the others take 120–128. The bottom navigation bar did not. With nomaxLines, Compose wrapped the labels mid-word at 2.0: the tab bar read "Calenda / r" and "Setting / s". Fixed inPeriodApp.ktwithmaxLines = 1and an ellipsis, which degrades to "Calen…" — still recognisable, and the icon above carries the meaning. What is still unreached: font scaling has only been driven in light mode, and only on a phone-sized screen.No device at— partly closed on 2026-08-18 atminSdk0d280d5. The app has now been built, installed and driven onPeriodMinSdk26, an API 26 emulator: onboarding start to first forecast, the Material 3 date picker, Today, Calendar, Insights and Settings, plus a relaunch after the emulator was killed and cold-booted. No crash, and noNoSuchMethodError,VerifyErrororNoClassDefFoundErrorattributable to the app — the onlyNoClassDefFoundErrorin logcat belongs tocom.google.android.googlequicksearchbox. What is still unreached at API 26: notifications and their lock-screen rendering, the instrumented suites, and anything needing more than one cycle of history, so fertility never became visible. Passes C, E, F, G and H have never run at this API level.- Physical-device coverage is undecided. Passes F and G need a real device with a lock screen and TalkBack; which device that is has not been chosen, and an emulator is not a substitute for either.
- Long-horizon accuracy — whether predictions measurably improve at 3, 6 and 12 confirmed cycles — cannot be reached by a QA round at all. It needs either a simulated history harness or real elapsed time, and until one exists the product's headline claim is tested only at the unit level.