Commit Graph

7 Commits

Author SHA1 Message Date
null 0ae92bd134 feat: let an archive come back, and refuse everything that is not one
The export was a one-way door. A user changing phone started the prediction
model from zero, which made "your history is yours" a smaller promise than it
sounded — and the archive says, in its own prose, "keep this file, a future
version of the app will be able to read it back". Nothing checked that was true.

`ExportReader` mirrors `ExportDocument.render` in the same module, over a
hand-written strict JSON reader. Strict is the point: it refuses duplicate keys,
trailing content, non-integer numbers and unsupported escapes, because a parser
that quietly accepts a trailing comma is a parser that will one day accept
somebody else's file and import it as a cycle history. A missing magic string is
`NotOurFile`; a version above this build's is `NewerFormat` rather than a
half-read; unknown *keys* are ignored, which is what lets the format grow; an
unknown `source` becomes `IMPORTED` rather than a guess about the user.

The write is `CycleRepository.importHistory`, not a loop over
`confirmPeriodStart`, because that would score the standing forecast against
history the app never predicted — and the backfill guard does not catch all of
it, since a file exported this morning carries this morning's period. So the
scoring is not dodged by accident of date; it is not reached. The engine runs
exactly once, at the end, over the whole history. Merge keeps what is on this
phone untouched (§14 — a file does not get to rewrite a record she made here),
and replace goes through the new `PeriodDatabase.replaceEverything`, which
empties the tables inside the same transaction as the writes: a failure between
the wipe and the inserts would leave her with neither her own history nor the
file's. Delete My Data still calls `deleteEverything` and still gets its VACUUM —
replacing a history is not erasing one.

The screen asks the one question the app must not answer for her, before the
picker rather than after, because the screen that asked is the screen a re-lock
destroys. Add is the filled button and needs no confirmation; replace is
confirmed by a dialog that names the deletion and points back at the safe
option. The result lands in the banner above the tabs, beside the export's, as a
live region — it arrives with no focus change.

The archive cannot touch the app lock. It carries `biometricUnlock` and the
importer does not apply it: there is no PIN in the file and there cannot be.

Also raises the lock ViewModel test's hang budget to three minutes. 60s passed
alone and failed when four modules ran in one invocation — three real PBKDF2
derivations at 210k iterations. The budget is for catching a stuck coroutine,
not a slow one.

Verified: `:core:export` round trip against the committed golden file (17);
`CycleRepositoryImportTest` — one recalculation for a four-record archive,
nothing scored, one standing forecast, and a failed replace leaving the history
intact (prove-guard: remove `withTransaction` from `replaceEverything`, exactly
one red); `DataImporterTest`, `ImportScreenTest`, `ImportControllerTest`,
`ImportCopyTest`. On PeriodQA: Settings → Restore → Add → Downloads →
records-2026-08-19.json gives "Restored. 4 periods and 2 spotting days added.",
Insights then shows three cycles and accuracy still unscored, and the same file
a second time gives "Everything in that file was already recorded here, so
nothing changed."

closes #58

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 02:27:42 -05:00
null 861738c4e9 fix: make the learning loop keep the samples it earns
Tracing the confirm -> score -> feed-back loop end to end turned up four
ways it lost or falsified its own evidence. None of them looked broken:
each produced plausible accuracy figures with something missing.

A forecast's generatedAt is now the date its LINEAGE began, not the date
its row was written. Every "Not yet", edit and delete revises the answer
to one standing question, so the replacement inherits the origin instead
of restamping today. Restamping moved the goalposts of the backfill
guard: a "Not yet" on the 29th, then a period logged on the 30th as
having started on the 28th, tripped the guard and the forecast the user
was actually shown was deleted unscored. The app learned nothing from
precisely the cycle it got wrong.

The standing snapshot is retired before the engine is consulted rather
than after, so a lineage dies with its history instead of waiting to be
scored against an unrelated one.

Scores now follow the period they are facts about. Editing a start
re-scores every snapshot recorded against it -- predictedStartDate stays
immutable, so a correction worsens the figure as readily as it improves
one -- and the same rule that refuses to score backfill retracts a score
whose start has moved behind its lineage, so an edit cannot smuggle in a
measurement the guard would have turned away. Deleting the period
retracts outright.

A confirm that becomes the newest start clears every "not yet", not just
the older ones: they all censor the same question. A date-bounded clear
left observations dated after a retroactively logged start alive to
depress the next cycle's confidence for a question already answered.

Also removes CycleData.repository's engine default. Nobody relied on it,
which is the point -- a caller who omitted the argument would compile
cleanly and ship the baseline prototype §11 calls unacceptable.

No schema change: all four fixes are queries and call order.

Six of the seven new tests were observed red against the unfixed source.
The seventh -- a deep backfill does not clear the observations censoring
the standing question -- passes both sides deliberately, pinning against
overshooting the not-yet fix.

closes #46
closes #47
closes #48
closes #49

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 16:13:53 -05:00
null 5da7c18364 feat: fertility that declines rather than stretches
§17 and §18. Ovulation is estimated a luteal phase before the PREDICTED next
period rather than counted forwards from the last one — the luteal phase is the
stable half of the cycle, which is why §17 asks for it that way — and the
fertile window opens five days before ovulation and closes one day after,
because sperm survive and the egg does not.

The uncertainty is inherited, not invented. Ovulation is derived from a
predicted date, so it can never be more certain than that prediction.

THE PART THE DEVICE TAUGHT

The first version showed a user one cycle in a fertile window of 8 Aug – 24 Aug.
Seventeen days. Arithmetically honest, and completely useless — over half a
cycle, dressed up as a feature.

So the estimate now returns null past a usable uncertainty, and Today says "Not
enough history to estimate. Log a few more cycles and the app will be able to
estimate ovulation." Three stable cycles later the same user gets 12 Aug – 20
Aug, which is worth reading. Verified in both states on a device.

That is the same shape as PredictionAccuracy refusing figures below three scored
forecasts and CycleInsights withholding an average below two intervals, and it
is now written down in the architecture doc as a rule rather than three
coincidences: the app declines rather than stretches.

§18'S PROHIBITION IS A TYPE, NOT A CONVENTION

FertilityLikelihood has LOWER, HIGHER and UNKNOWN and no fourth value. Somebody
reading "safe" would take a decision on it; the estimate comes from a predicted
date carrying days of uncertainty; and §18 has already promised this is not
contraception. A test asserts no label contains a permission word, so adding one
is a deliberate act with a failing test.

The disclaimer travels with the feature — same screen, same time. A disclaimer
one tab away is a disclaimer nobody read.

ANOTHER GREYSCALE COLLISION

The ovulation star was centred, which put it directly behind the numeral: in
greyscale "18" and the mark merged into one smudge. Ovulation is now the fertile
ring plus a small star low in the cell, which is also semantically right — that
day IS inside the window, and the pair reads as "that window, and this day".

Nine Compose ModifierParameter warnings fixed properly rather than suppressed.

164 tests, all passing. ./gradlew check green, 0 lint errors.

closes #21
closes #22
closes #23
2026-08-18 14:56:05 -05:00
null 2fe423cf47 feat: the prediction engine section 12 specifies, and it beats the baseline
PersonalPredictionEngine keeps a discrete probability distribution over
candidate start dates rather than a date with a margin bolted on. Everything the
product needs falls out of that one structure: the most likely date is its mode,
the window is the narrowest span holding 80% of the mass, and a "Not yet" is the
distribution conditioned on what the user just said — which is what §13 asks for
and what a date-plus-margin design cannot express at all.

It is better, and that is a number rather than an opinion. EngineComparisonTest
scores both engines over the §51 fixtures on every build:

  engine      MAE    mean window   within +/-2   window covered
  baseline    1.00    2.67          7/9           7/9
  personal    0.67    4.56          9/9           9/9

COVERAGE IS THE MEASURE, NOT WIDTH

The first version of that test asserted the new windows must not be wider, and
it failed. Measuring showed why the assertion was wrong: the fixtures where the
personal engine is wider are the ones that are genuinely less certain — a
history with a suspected missing period, and one with a 45-day outlier — and the
baseline answers both with a two-day window and misses. What a window promises
is that the period starts inside it. An engine keeping that promise 7 times in 9
has a broken promise, not a tight forecast. The test now asserts coverage, with
a ceiling so "some time this month" still fails.

THREE MODELLING BUGS THE TESTS FOUND

Each was found by a test failing, not by reading the code:

  - Median absolute deviation alone reads a user alternating 25 and 37 as
    perfectly consistent, because half her deviations are zero. Twenty
    disagreeing cycles came back High, breaking §15's rule that volume alone
    must never buy High confidence. Spread is now the larger of MAD and mean
    absolute deviation; robustness comes from IntervalAnalysis down-weighting
    what is questionable, which is a better place for it.

  - Recency weighting assumes the recent past predicts the near future. For a
    variable user that is false — her latest cycle is a draw from a wide
    distribution, not a signal — and weighting it equally cost three days on the
    §51 variable fixture. Recency is now trusted in proportion to how much her
    cycles actually agree.

  - A fixed one-day floor on trend detection fired on a 42-day-cycle history
    whose medians differed by a single day, turning an exact forecast into a
    wrong one. One day is a real trend at 28 and rounding error at 42, so the
    floor is relative to the user's own spread.

WIRED THROUGH, NOT JUST TESTED

PredictionInput carries recentAbsoluteErrors, and CycleRepository feeds the
scored errors back in. Without that the app stores every error it makes and
never reads one back — measuring accuracy rather than learning from it, with the
widening happening only in a unit test. A repository test asserts the errors
actually reach the engine.

BaselinePredictionEngine stays as the control, and both engines run the same
§51 acceptance suite, so the next engine's improvement is measurable too.

108 tests, all passing. ./gradlew check green. Verified on a device.

closes #10
closes #11
closes #12
closes #14
2026-08-18 03:16:12 -05:00
null f5e9fbe53c feat: module boundary guard, and two API-level bugs it uncovered
checkModuleBoundaries holds the dependency tables in docs/architecture/README.md
as a check: every module's permitted project dependencies, plus the rule that
domain:cycle and domain:prediction must never apply an Android plugin. core:ads
is already in the map with an empty permitted set, before the module exists —
PRODUCT_PLAN.md §34 is non-negotiable, and a guard written alongside the code it
constrains is one shaped around whatever exception somebody wanted at the time.

It lists every violation rather than the first, and refuses to report a pass
when it examined no modules at all.

THE GUARD FAILED ITS OWN FIRST PROOF

prove-guard.sh injected a forbidden dependency into :domain:prediction and the
guard reported "7 modules checked, no violations". The root project is
configured before its subprojects, so reading subprojects.configurations from
the root script saw every configuration empty — it had been green over an empty
map since the moment it was written, and would have been trusted for months.

Collection moved into afterEvaluate, and the task now throws rather than passing
if it ends up with no modules. Three proofs recorded in the architecture doc,
all re-run and all red: a domain module reaching upward, :app reaching past the
repository straight to Room, and a module with no rule being reported as
unmeasured rather than assumed fine.

TWO REAL BUGS FROM WIRING IT INTO `check`

Running the whole check for the first time turned up Android lint errors that
would have shipped:

  NewApi: java.time.LocalDate#ofInstant requires API 34 (minSdk is 26)
  NewApi: java.time.LocalDate#EPOCH     requires API 34 (minSdk is 26)

Both are on the recalculation path. On any device below Android 14 — most of
the install base this app targets — that is a crash. Neither the unit tests nor
the API 36 emulator could see it; lint is the only thing that could.

Replaced with atZone().toLocalDate() and ofEpochDay(0), which are API 26.

Also cleared the lint warnings that were real: a redundant activity label, and
a round launcher icon declared but never referenced. The two that remain are
deliberate and now say so where the warning is read — targetSdk 36 is Play's
floor and raising it opts into untested runtime behaviour, and the -v26 mipmap
qualifier stays because removing it makes AAPT fail to resolve the icon at all.

./gradlew check now passes with 0 lint errors across all seven modules.
70 unit tests, all passing.

closes #7
2026-08-18 03:00:11 -05:00
null adc50751d8 feat: period CRUD end to end, and stop a double tap killing the app
The Batch 01 vertical slice from PRODUCT_PLAN.md §58 now runs on a device:
launch, log a period, it is stored, the forecast recalculates, edit or delete it
and the forecast moves again. Hilt wiring, a TodayViewModel exposing one
immutable state, and a working surface that says "Batch 01 · working surface" at
the top so nobody mistakes it for the designed Today screen, which is Batch 03.

THE DEFECT THIS FOUND, ON A DEVICE

Tapping "Started today" twice on the same day killed the app:

  FATAL EXCEPTION: main
  android.database.sqlite.SQLiteConstraintException: UNIQUE constraint failed:
  period_records.startDate

Not a hypothetical — the crash was reproduced on emulator-5580, the fix
applied, and the same two taps then produced "That day is already logged." with
the process still alive and zero FATAL lines in logcat.

The constraint is right: a duplicate must not overwrite the original row and
lose its createdAt and source. The API around it was wrong. Repeating a tap
when you are not sure the first one registered is an ordinary thing for a person
to do, not a fault, and it must never be an exception. So the period writes
return PeriodWriteResult — Added, AlreadyRecorded, Updated, Conflict, NotFound —
and only genuine faults still throw.

editPeriod had the same hole: moving a record onto a date another record holds.
That is refused rather than merged, because merging would delete a period the
user entered and only they can settle it.

The ViewModel now installs a CoroutineExceptionHandler as a backstop. In a
health app a crash mid-write is adjacent to losing what was just entered, and a
message somebody can read beats a process that vanished. The message carries the
exception type and never a record's contents (§45).

Four regression tests pin all of it, plus two instrumented tests on a real
file-backed database that close and reopen it — what a force-stop actually does,
and something an in-memory database cannot fail.

70 unit tests and 2 instrumented tests, all passing. Release APK 1.2 MB.

closes #6
2026-08-18 02:52:35 -05:00
null 6d4592467f feat: repository layer, and stop backfilled history fabricating accuracy figures
core/data is the seam between storage and everything else. Reads return domain
types, cycles are derived rather than stored, and the forecast is a function of
the data instead of a field somebody has to remember to refresh — so §11's
"recalculate after a confirmed start, after an edit, after a Not yet" is
automatic rather than three call sites.

Confirming a period is four writes in one transaction, because a partial result
is a corrupt history rather than a failed action: write the record, score the
forecast that was standing, clear the "not yet" observations it resolved, and
snapshot a fresh forecast.

THE DEFECT THIS FOUND

A test expecting one scored prediction found three. The cause was not the test:
every historical period entered during onboarding was scoring the current
forecast against a date in the past, inventing an error for a prediction nobody
had ever been shown. §16's "your predictions are getting better" would have been
populated with figures the app made up about itself — plausible ones, which is
what makes it expensive to notice.

Two rules now, both pinned by tests:

  - exactly one unscored snapshot exists at a time. A forecast superseded before
    its outcome was known is not a wrong forecast, and counting it lets one
    cycle contribute several errors.
  - a confirmed start only scores a forecast made on or before it. Anything
    earlier is backfill and leaves the standing forecast alone.

Accuracy also stays quiet below three scored predictions. One lucky forecast
reading "average error: 0 days" is an overstatement, not a measurement.

THE ROOM BOUNDARY, HELD THREE WAYS

implementation rather than api on core:database; CycleRepository's constructor
internal because it names a PeriodDatabase; reads mapped to domain types in
Mappers.kt. Callers use CycleData.repository(context) and never learn Room
exists. Verified rather than asserted: grep -rn "androidx.room" app/src domain
is empty, and Room appears zero times in :app's debugCompileClasspath.

No fallbackToDestructiveMigration: it turns a forgotten migration into a silent
wipe of the user's entire cycle history on update.

Also fixed: `domain/*` inside a KDoc silently opened a nested block comment —
Kotlin block comments nest — which broke compilation in a way the error message
pointed nowhere near.

58 tests across the project, all passing.

closes #5
2026-08-18 02:41:05 -05:00