docs: record Batch 09, and close half the long-horizon QA gap
ClaudeQACoverage has carried "long-horizon accuracy is tested only at the unit level" as a standing gap for four rounds. The simulated-history half is now closed by LearningCurveTest, and the entry says so -- along with what it found, since a harness that changed nothing would be a weaker claim than one that caught two real faults. The other half stays open and is stated plainly: a simulation exercises the engine, not the app. Nothing here proves a real install accumulates those cycles correctly over a year, that snapshots survive updates and time-zone changes, or that the accuracy card shows what the engine computed. That needs elapsed time on a device. Development log entry per WORK_CYCLE step 6, including the two corrections this batch made to its own plan: the coverage dip was a 40-seed artifact, and a swept constant shipped one value away from the measured one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
866b7055f5
commit
b73bc85583
|
|
@ -32,6 +32,90 @@ written and stay true. It is exempt from review for the same reason a receipt is
|
||||||
|
|
||||||
## Entries
|
## Entries
|
||||||
|
|
||||||
|
### 2026-08-20 (later) — Batch 09: the learning loop, measured
|
||||||
|
|
||||||
|
Started from a question — *how well does this app learn a period?* — and
|
||||||
|
answered it by measuring rather than reading, which turned up more than the
|
||||||
|
question asked for.
|
||||||
|
|
||||||
|
**The engine does learn.** A stable user reaches the engine's accuracy floor,
|
||||||
|
around 0.8 days, within two or three cycles, and an intrinsically variable one
|
||||||
|
never converges and is correctly told so. That much was already true.
|
||||||
|
|
||||||
|
**The loop around it was losing evidence.** Four defects, none of which looked
|
||||||
|
broken because each produced plausible figures with something missing. A
|
||||||
|
forecast's `generatedAt` was restamped by every "Not yet", edit and delete, so a
|
||||||
|
period logged retroactively tripped the backfill guard and the forecast the user
|
||||||
|
was actually shown was deleted unscored — the app learned nothing from precisely
|
||||||
|
the cycle it got wrong. A null forecast skipped the retire step, leaving a
|
||||||
|
snapshot about a deleted history to be scored against an unrelated one. Scored
|
||||||
|
snapshots never followed the period they measured when it was edited or deleted.
|
||||||
|
And a retroactive confirm left later-dated "not yet" answers alive to depress the
|
||||||
|
next cycle's confidence. `generatedAt` is now a lineage origin carried across
|
||||||
|
revisions; scores re-score or retract with their ground truth; a resolving
|
||||||
|
confirm clears every observation. No schema change — all four are queries and
|
||||||
|
call order. Six of seven new tests were watched failing first.
|
||||||
|
|
||||||
|
**Three calibration faults, all found by a harness that did not exist.**
|
||||||
|
`LearningCurveTest` walks four simulated users through fifteen cycles, 400 seeds
|
||||||
|
each, scoring and feeding back exactly as the repository does. It closes half of
|
||||||
|
the QA doc's standing long-horizon gap and immediately earned itself:
|
||||||
|
|
||||||
|
- A textbook-regular user read High in 15.7% of forecasts. Confidence divided
|
||||||
|
raw error, so the moment real errors arrived — averaging 0.9 days against a
|
||||||
|
cycle that varies by a day — she scored worse than when she had never been
|
||||||
|
measured at all. Errors are now scored against what the window already admits
|
||||||
|
(for a Laplace, E|X| = b), and she reads High 88% of the time.
|
||||||
|
- A drifting user was tracked to within 1.2 days and told the app was losing
|
||||||
|
confidence, because spread was measured from a static centre and grew with the
|
||||||
|
history however well the trend was followed. Residuals now come off the trend
|
||||||
|
line. That also revealed the damping had been *paying for* the measurement
|
||||||
|
error: raising it 0.50 → 0.70 halved the fixture error, took fleet MAE from
|
||||||
|
0.67 to 0.44, restored 9/9 coverage and narrowed the mean window.
|
||||||
|
- The window promised 80% and delivered less early on, because the spread
|
||||||
|
estimate never accounted for how well the *centre* was known. A √(1 + 1/n)
|
||||||
|
term over the effective sample size fixes it.
|
||||||
|
|
||||||
|
**Two corrections worth recording.** The coverage dip that prompted #52 was a
|
||||||
|
40-seed artifact — at 400 seeds those cells sit at 73–85%, and the fix was
|
||||||
|
rewritten around what survived. And a sweep loop left `MEDIUM_THRESHOLD` at 0.44
|
||||||
|
when 0.40 was the measured choice; both passed the suite, so only the numbers
|
||||||
|
caught it. Fixed in the following commit.
|
||||||
|
|
||||||
|
`modelVersion` is now `personal-2`.
|
||||||
|
|
||||||
|
**Batch 09 also wrote down what the engine assumes about bodies.**
|
||||||
|
`docs/architecture/SCIENCE.md` gives every physiology constant a verdict and a
|
||||||
|
citation, and says plainly which constants are engineering and need none. One
|
||||||
|
value moved: the population default, used only for a user with a single logged
|
||||||
|
period, goes 28 → 29, which is what the large app-data studies actually measure.
|
||||||
|
The luteal phase is filed rather than changed (#56) — 14 is above the measured
|
||||||
|
mean, but moving it shifts what the fertility screen shows every user, and that
|
||||||
|
is a product decision with its own evidence. The Insights banner also stopped
|
||||||
|
promising "predictions get noticeably better" before a single forecast had been
|
||||||
|
scored.
|
||||||
|
|
||||||
|
**And the calendar reaches a year (#57, Batch 10).** `CycleProjection` is a
|
||||||
|
separate type from `Prediction` because a forecast is scored and a projection
|
||||||
|
never is. Uncertainty grows as √n, it declines rather than stretching past ten
|
||||||
|
days either side, and every month showing projected marks carries *"If your
|
||||||
|
cycles continue as they have, this is the forecast"* with the confidence for that
|
||||||
|
distance. The opposite of the widget library that prompted it, which stamps a
|
||||||
|
fixed 28 days twelve months forward and draws it exactly like the next period.
|
||||||
|
|
||||||
|
Batch 10 also holds two filed-not-built items: restore-from-export (#58 — the
|
||||||
|
export is currently a one-way door, so a new phone starts the model from zero)
|
||||||
|
and a cycle-length chart for Insights (#59).
|
||||||
|
|
||||||
|
**Next action:** decide #56 — measure a 13-day luteal phase and either move
|
||||||
|
`DEFAULT_LUTEAL_PHASE_DAYS` with its evidence or record why 14 stays. Then #58,
|
||||||
|
restore-from-export, which is the one item with a data-ownership obligation
|
||||||
|
behind it.
|
||||||
|
|
||||||
|
**Blockers:** none in code. The unreachable half of the long-horizon QA gap
|
||||||
|
stands — a simulation exercises the engine, not a year-old install — and API 26
|
||||||
|
still has no prediction coverage, both recorded in `docs/qa/ClaudeQACoverage.md`.
|
||||||
|
|
||||||
### 2026-08-20 (final) — Discreet launcher alias landed
|
### 2026-08-20 (final) — Discreet launcher alias landed
|
||||||
|
|
||||||
Issue #31 is ready. The normal launcher entry is now
|
Issue #31 is ready. The normal launcher entry is now
|
||||||
|
|
|
||||||
|
|
@ -234,7 +234,23 @@ existing is not a failure; it not existing while the gaps do is.
|
||||||
- **Physical-device coverage is undecided.** Passes F and G need a real device
|
- **Physical-device coverage is undecided.** Passes F and G need a real device
|
||||||
with a lock screen and TalkBack; which device that is has not been chosen, and
|
with a lock screen and TalkBack; which device that is has not been chosen, and
|
||||||
an emulator is not a substitute for either.
|
an emulator is not a substitute for either.
|
||||||
- **Long-horizon accuracy** — whether predictions measurably improve at 3, 6 and
|
- **Long-horizon accuracy — half closed.** Whether predictions measurably improve
|
||||||
12 confirmed cycles — cannot be reached by a QA round at all. It needs either
|
at 3, 6 and 12 confirmed cycles cannot be reached by a QA round at all: it
|
||||||
a simulated history harness or real elapsed time, and until one exists the
|
needs either a simulated history harness or real elapsed time. The harness now
|
||||||
product's headline claim is tested only at the unit level.
|
exists. `LearningCurveTest` walks four simulated users through fifteen cycles
|
||||||
|
each, 400 seeds apiece, scoring every forecast and feeding the error back
|
||||||
|
exactly as `CycleRepository` does, and asserts on the aggregates — accuracy
|
||||||
|
floors, coverage against the window's own 80% promise, and how often each user
|
||||||
|
is told High or Low.
|
||||||
|
|
||||||
|
It found what it was built to find: a textbook-regular user was reading High in
|
||||||
|
only 15.7% of her forecasts, and a drifting user was tracked to within 1.2 days
|
||||||
|
while being told the app was losing confidence. Both are fixed, and the guard
|
||||||
|
now holds them.
|
||||||
|
|
||||||
|
**What is still open is the other half.** A simulation exercises the engine,
|
||||||
|
not the app: nothing here proves that a real install accumulates those cycles
|
||||||
|
correctly over a year, that the snapshots survive updates and time-zone
|
||||||
|
changes, or that what the user sees on the accuracy card matches what the
|
||||||
|
engine computed. That needs elapsed time on a device and no harness replaces
|
||||||
|
it.
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue