Tracing the confirm -> score -> feed-back loop end to end turned up four
ways it lost or falsified its own evidence. None of them looked broken:
each produced plausible accuracy figures with something missing.
A forecast's generatedAt is now the date its LINEAGE began, not the date
its row was written. Every "Not yet", edit and delete revises the answer
to one standing question, so the replacement inherits the origin instead
of restamping today. Restamping moved the goalposts of the backfill
guard: a "Not yet" on the 29th, then a period logged on the 30th as
having started on the 28th, tripped the guard and the forecast the user
was actually shown was deleted unscored. The app learned nothing from
precisely the cycle it got wrong.
The standing snapshot is retired before the engine is consulted rather
than after, so a lineage dies with its history instead of waiting to be
scored against an unrelated one.
Scores now follow the period they are facts about. Editing a start
re-scores every snapshot recorded against it -- predictedStartDate stays
immutable, so a correction worsens the figure as readily as it improves
one -- and the same rule that refuses to score backfill retracts a score
whose start has moved behind its lineage, so an edit cannot smuggle in a
measurement the guard would have turned away. Deleting the period
retracts outright.
A confirm that becomes the newest start clears every "not yet", not just
the older ones: they all censor the same question. A date-bounded clear
left observations dated after a retroactively logged start alive to
depress the next cycle's confidence for a question already answered.
Also removes CycleData.repository's engine default. Nobody relied on it,
which is the point -- a caller who omitted the argument would compile
cleanly and ship the baseline prototype §11 calls unacceptable.
No schema change: all four fixes are queries and call order.
Six of the seven new tests were observed red against the unfixed source.
The seventh -- a deep backfill does not clear the observations censoring
the standing question -- passes both sides deliberately, pinning against
overshooting the not-yet fix.
closes#46closes#47closes#48closes#49
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PersonalPredictionEngine keeps a discrete probability distribution over
candidate start dates rather than a date with a margin bolted on. Everything the
product needs falls out of that one structure: the most likely date is its mode,
the window is the narrowest span holding 80% of the mass, and a "Not yet" is the
distribution conditioned on what the user just said — which is what §13 asks for
and what a date-plus-margin design cannot express at all.
It is better, and that is a number rather than an opinion. EngineComparisonTest
scores both engines over the §51 fixtures on every build:
engine MAE mean window within +/-2 window covered
baseline 1.00 2.67 7/9 7/9
personal 0.67 4.56 9/9 9/9
COVERAGE IS THE MEASURE, NOT WIDTH
The first version of that test asserted the new windows must not be wider, and
it failed. Measuring showed why the assertion was wrong: the fixtures where the
personal engine is wider are the ones that are genuinely less certain — a
history with a suspected missing period, and one with a 45-day outlier — and the
baseline answers both with a two-day window and misses. What a window promises
is that the period starts inside it. An engine keeping that promise 7 times in 9
has a broken promise, not a tight forecast. The test now asserts coverage, with
a ceiling so "some time this month" still fails.
THREE MODELLING BUGS THE TESTS FOUND
Each was found by a test failing, not by reading the code:
- Median absolute deviation alone reads a user alternating 25 and 37 as
perfectly consistent, because half her deviations are zero. Twenty
disagreeing cycles came back High, breaking §15's rule that volume alone
must never buy High confidence. Spread is now the larger of MAD and mean
absolute deviation; robustness comes from IntervalAnalysis down-weighting
what is questionable, which is a better place for it.
- Recency weighting assumes the recent past predicts the near future. For a
variable user that is false — her latest cycle is a draw from a wide
distribution, not a signal — and weighting it equally cost three days on the
§51 variable fixture. Recency is now trusted in proportion to how much her
cycles actually agree.
- A fixed one-day floor on trend detection fired on a 42-day-cycle history
whose medians differed by a single day, turning an exact forecast into a
wrong one. One day is a real trend at 28 and rounding error at 42, so the
floor is relative to the user's own spread.
WIRED THROUGH, NOT JUST TESTED
PredictionInput carries recentAbsoluteErrors, and CycleRepository feeds the
scored errors back in. Without that the app stores every error it makes and
never reads one back — measuring accuracy rather than learning from it, with the
widening happening only in a unit test. A repository test asserts the errors
actually reach the engine.
BaselinePredictionEngine stays as the control, and both engines run the same
§51 acceptance suite, so the next engine's improvement is measurable too.
108 tests, all passing. ./gradlew check green. Verified on a device.
closes#10closes#11closes#12closes#14
core/data is the seam between storage and everything else. Reads return domain
types, cycles are derived rather than stored, and the forecast is a function of
the data instead of a field somebody has to remember to refresh — so §11's
"recalculate after a confirmed start, after an edit, after a Not yet" is
automatic rather than three call sites.
Confirming a period is four writes in one transaction, because a partial result
is a corrupt history rather than a failed action: write the record, score the
forecast that was standing, clear the "not yet" observations it resolved, and
snapshot a fresh forecast.
THE DEFECT THIS FOUND
A test expecting one scored prediction found three. The cause was not the test:
every historical period entered during onboarding was scoring the current
forecast against a date in the past, inventing an error for a prediction nobody
had ever been shown. §16's "your predictions are getting better" would have been
populated with figures the app made up about itself — plausible ones, which is
what makes it expensive to notice.
Two rules now, both pinned by tests:
- exactly one unscored snapshot exists at a time. A forecast superseded before
its outcome was known is not a wrong forecast, and counting it lets one
cycle contribute several errors.
- a confirmed start only scores a forecast made on or before it. Anything
earlier is backfill and leaves the standing forecast alone.
Accuracy also stays quiet below three scored predictions. One lucky forecast
reading "average error: 0 days" is an overstatement, not a measurement.
THE ROOM BOUNDARY, HELD THREE WAYS
implementation rather than api on core:database; CycleRepository's constructor
internal because it names a PeriodDatabase; reads mapped to domain types in
Mappers.kt. Callers use CycleData.repository(context) and never learn Room
exists. Verified rather than asserted: grep -rn "androidx.room" app/src domain
is empty, and Room appears zero times in :app's debugCompileClasspath.
No fallbackToDestructiveMigration: it turns a forgotten migration into a silent
wipe of the user's entire cycle history on update.
Also fixed: `domain/*` inside a KDoc silently opened a nested block comment —
Kotlin block comments nest — which broke compilation in a way the error message
pointed nowhere near.
58 tests across the project, all passing.
closes#5
core/database holds the four entities from PRODUCT_PLAN.md §10 —
period_records, spotting_records, prediction_records, not_yet_observations —
with DAOs returning Flow, epoch-day/epoch-milli converters, and the schema
exported to core/database/schemas and committed.
Three constraints are structural rather than remembered:
- startDate is UNIQUE and inserts ABORT rather than REPLACE. REPLACE would
delete the original row with its createdAt and source; §14 says health
history is never modified silently.
- spotting has its own table, so no query for periods can reach it. §25: it
must never start or reset a cycle.
- a prediction snapshot can be scored but not rewritten — score() sets only
actualStartDate and absoluteErrorDays. A snapshot editable after the fact
can only ever report that the app was right, which would make §16's whole
accuracy feature a lie.
deleteEverything() is one transaction and the only bulk delete in the module: a
partial wipe leaves the cycle reconstructible from the tables the user asked to
be rid of.
14 tests, on the JVM under Robolectric — no emulator.
THE SCHEMA GUARD, AND WHY IT IS A SCRIPT
SchemaTest was written as a drift guard and proved not to be one. Room
regenerates the schema export during compilation, so adding a column to
PeriodRecordEntity without bumping VERSION leaves the suite green while the
committed schema quietly changes underneath it. That was not reasoned about, it
was run: the column was added, 1.json gained it, and every test passed. On a
device that is "Room cannot verify the data integrity" — a crash on update,
after shipping.
scripts/schema-guard.sh asks git instead, which Room cannot overwrite. Proved
both ways before being trusted: green on a clean tree, exit 1 on the injected
drift. It runs in pre-commit when an entity or the schema directory is staged,
and the hook treats its exit 2 as a refusal.
SchemaTest keeps its four tests and now documents what it does not catch.
Room's own MigrationTestHelper is not used: every constructor needs an
Instrumentation and schema assets, and AGP 9's library source-set DSL throws
DefaultAndroidLibrarySourceSet_Decorated cannot be cast to
AndroidLibrarySourceSet when you add an asset directory. Recorded so the next
person does not spend the afternoon on it.
Docs updated in this commit, as their triggers required: the migration table
now has its version 1 row and the trap that makes such tables go stale, TOOLS
explains the seventh script, and the hooks README lists the new guard.
closes#3