Privacy-Period-Tracker/docs/architecture
null 2fe423cf47 feat: the prediction engine section 12 specifies, and it beats the baseline
PersonalPredictionEngine keeps a discrete probability distribution over
candidate start dates rather than a date with a margin bolted on. Everything the
product needs falls out of that one structure: the most likely date is its mode,
the window is the narrowest span holding 80% of the mass, and a "Not yet" is the
distribution conditioned on what the user just said — which is what §13 asks for
and what a date-plus-margin design cannot express at all.

It is better, and that is a number rather than an opinion. EngineComparisonTest
scores both engines over the §51 fixtures on every build:

  engine      MAE    mean window   within +/-2   window covered
  baseline    1.00    2.67          7/9           7/9
  personal    0.67    4.56          9/9           9/9

COVERAGE IS THE MEASURE, NOT WIDTH

The first version of that test asserted the new windows must not be wider, and
it failed. Measuring showed why the assertion was wrong: the fixtures where the
personal engine is wider are the ones that are genuinely less certain — a
history with a suspected missing period, and one with a 45-day outlier — and the
baseline answers both with a two-day window and misses. What a window promises
is that the period starts inside it. An engine keeping that promise 7 times in 9
has a broken promise, not a tight forecast. The test now asserts coverage, with
a ceiling so "some time this month" still fails.

THREE MODELLING BUGS THE TESTS FOUND

Each was found by a test failing, not by reading the code:

  - Median absolute deviation alone reads a user alternating 25 and 37 as
    perfectly consistent, because half her deviations are zero. Twenty
    disagreeing cycles came back High, breaking §15's rule that volume alone
    must never buy High confidence. Spread is now the larger of MAD and mean
    absolute deviation; robustness comes from IntervalAnalysis down-weighting
    what is questionable, which is a better place for it.

  - Recency weighting assumes the recent past predicts the near future. For a
    variable user that is false — her latest cycle is a draw from a wide
    distribution, not a signal — and weighting it equally cost three days on the
    §51 variable fixture. Recency is now trusted in proportion to how much her
    cycles actually agree.

  - A fixed one-day floor on trend detection fired on a 42-day-cycle history
    whose medians differed by a single day, turning an exact forecast into a
    wrong one. One day is a real trend at 28 and rounding error at 42, so the
    floor is relative to the user's own spread.

WIRED THROUGH, NOT JUST TESTED

PredictionInput carries recentAbsoluteErrors, and CycleRepository feeds the
scored errors back in. Without that the app stores every error it makes and
never reads one back — measuring accuracy rather than learning from it, with the
widening happening only in a unit test. A repository test asserts the errors
actually reach the engine.

BaselinePredictionEngine stays as the control, and both engines run the same
§51 acceptance suite, so the next engine's improvement is measurable too.

108 tests, all passing. ./gradlew check green. Verified on a device.

closes #10
closes #11
closes #12
closes #14
2026-08-18 03:16:12 -05:00
..
githooks feat: Room database for cycle history, and the schema guard that actually works 2026-08-18 02:32:07 -05:00
GUARDS.md chore: adopt the project template and add the Kotlin/Compose skeleton 2026-08-18 02:16:47 -05:00
README.md feat: the prediction engine section 12 specifies, and it beats the baseline 2026-08-18 03:16:12 -05:00

README.md

Architecture

Status: Current
Owner: _null
Last reviewed: 2026-08-18
Governs: docs/architecture/**, the Gradle module graph, and the data shapes that
         outlive a function
Review trigger: Any new Gradle module, any change to a module boundary, any change
                to a Room entity or a DAO, any new Room migration, any dependency
                added to a domain/* module

The shape

Compose UI  (app, feature/*)
    ↓
ViewModel — immutable StateFlow of screen state
    ↓
Use case / prediction engine  (domain/*)
    ↓
Repository  (core/data)
    ↓
Room + DataStore  (core/database, core/datastore)

Unidirectional: state flows down as an immutable UiState, events flow up as function calls. Nothing below the ViewModel knows Compose exists.

Modules

Seven today — a module created before it has contents is a place for things to be put by accident. The wider layout sketched in ../planning/PRODUCT_PLAN.md §9 arrives the same way, with the batch that needs it.

Module Plugin Owns May depend on
app Android application MainActivity, the four-tab navigation shell, DI wiring everything below
core/designsystem Android library Material 3 theme, colour and type tokens nothing in this project
core/database Android library Room entities, DAOs, converters, the schema export domain/cycle, domain/prediction
core/datastore Android library UserPreferences and the settings that are not health history nothing in this project
core/data Android library CycleRepository, entity⇄domain mapping, accuracy — the only module that touches a DAO core/database, domain/cycle, domain/prediction
domain/cycle Kotlin JVM PeriodRecord, SpottingRecord, CycleRecord and the rules over them nothing
domain/prediction Kotlin JVM the forecast, the window, confidence, NotYetObservation domain/cycle

Planned, with the issue that brings each one. Named without backticks on purposedoc-claims.sh reads a backticked path as a claim that the file is there, and none of these are:

Module Plugin Owns May depend on Issue
core/ads Android library the AdProvider implementation neither core/database nor domain/* Batch 07

Why domain/* is kotlin("jvm") and not an Android library

PRODUCT_PLAN.md §57.10 asks for the prediction engine to be unit-testable without Android. A convention saying "do not import android.* here" is a convention somebody breaks at 11pm; a module that cannot see the Android SDK at all is a compile error instead.

It buys the thing §50 depends on: the acceptance tests in §51 — stable 35-day user, variable user, 45-day outlier, "not yet" — run on the JVM in under a second, so they run on every commit rather than on an emulator when someone remembers.

How the Room boundary is actually held

Three mechanisms, and it matters that none of them is "people remember":

  1. core/data depends on core/database with implementation, not api, so Room never reaches the compile classpath of anything above it.
  2. CycleRepository's constructor is internal — it names a PeriodDatabase, and a public constructor would force every caller to be able to name that type too. Callers use CycleData.repository(context).
  3. Repository reads return domain types. Mappers.kt is the one place an entity and a domain object meet.

Checkable, not merely intended: grep -rn "androidx.room" app/src domain is empty, and ./gradlew :app:dependencies --configuration debugCompileClasspath lists Room zero times.

No fallbackToDestructiveMigration. It turns a forgotten migration into silent data loss on update — here, a user's entire cycle history gone with no error and no way back. A missing migration must be a crash in testing rather than a wipe in production.

Ordinary outcomes are values; only faults are exceptions

period_records.startDate is UNIQUE and inserts ABORT rather than REPLACE, so a duplicate cannot destroy the original row. Both are right. The API around them was not: confirmPeriodStart let SQLiteConstraintException out, and viewModelScope.launch has no handler, so tapping "Started today" twice killed the app — found on a device, not in a test.

The mistake was treating already recorded as an error. It is a completely reasonable thing for a person to do twice when they are unsure the first tap registered. So the period writes return PeriodWriteResultAdded, AlreadyRecorded, Updated, Conflict, NotFound — and a genuine fault (full disk, corrupt database) still throws, because that is not something a caller can carry on from.

Conflict is refused rather than resolved: moving a record onto a date another record already holds is a question only the user can settle, and merging would delete a period they entered.

The ViewModel also installs a CoroutineExceptionHandler as a backstop. In a health app a crash mid-write is adjacent to losing what the user just entered, and a message they can read beats a process that vanished. The message carries the exception type and never a record's contents — §45.

One forecast stands at a time, and backfill is not a prediction

Two rules about prediction_records that are easy to get wrong and expensive to notice, because both failure modes produce plausible accuracy figures rather than obviously broken ones.

Exactly one unscored snapshot exists at any moment. Every recalculation replaces the standing forecast instead of appending. A forecast superseded before its outcome was known is not a wrong forecast — nobody was looking at it when the period arrived — and counting it as one lets a single cycle contribute several errors to §16's figures.

A confirmed start only scores a forecast made on or before it. Onboarding asks for earlier periods and a user can add one at any time; scoring today's forecast against a date in the past invents an error for a prediction nobody was ever shown. This was a real defect, caught by a test expecting one scored prediction and finding three, and it is pinned by backfilled history does not fabricate accuracy figures.

The boundary that is not negotiable

The advertising subsystem must never receive menstrual dates, cycle length, period duration, fertility status, ovulation estimates, prediction confidence, prediction history, spotting records, or any other health-derived attribute. — PRODUCT_PLAN.md §34

Expressed structurally rather than as a rule people remember: when core/ads exists it will declare no dependency on core/database or domain/*, and a Gradle check enforces the whole table above by enumerating each module's allowed dependencies. Ads reach the UI through an AdProvider interface owned by app.

The guard, and the three proofs it survived

checkModuleBoundaries in the root build.gradle.kts holds the tables above as a check. It runs as part of ./gradlew check, lists every violation rather than the first, and refuses to report a pass when it examined no modules at all.

Per GUARDS.md §1 it is not evidence until it has been watched failing. These three are repeatable, each restores the file from a trap, and each was run:

# 1. A domain module reaching upward — the leak that would end JVM-only tests
bash scripts/prove-guard.sh domain/prediction/build.gradle.kts \
  'implementation(project(":domain:cycle"))' \
  'implementation(project(":domain:cycle"))
    implementation(project(":core:database"))' \
  ./gradlew checkModuleBoundaries

# 2. :app reaching past the repository straight to Room
bash scripts/prove-guard.sh app/build.gradle.kts \
  'implementation(project(":core:data"))' \
  'implementation(project(":core:data"))
    implementation(project(":core:database"))' \
  ./gradlew checkModuleBoundaries

# 3. A module with no rule is reported as unmeasured, not assumed fine
bash scripts/prove-guard.sh build.gradle.kts \
  '":core:data" to setOf(":core:database", ":domain:cycle", ":domain:prediction"),' \
  '' \
  ./gradlew checkModuleBoundaries

Proof 1 failed the first time, and that is the point. The guard reported "7 modules checked, no violations" with a forbidden dependency sitting in the build file. The root project is configured before its subprojects, so reading subprojects.configurations from the root script saw every configuration empty — the check had been green over an empty map since the moment it was written. Collection now happens in afterEvaluate, and the task throws rather than passing if it ends up with no modules to examine.

Thirty seconds of proving caught a guard that would otherwise have been trusted for months.

Data shapes

Defined in domain/cycle as plain Kotlin, and mirrored by Room entities in core/database once issue #3 creates it. The full field lists are PRODUCT_PLAN.md §10; what matters here is why each exists and what must not happen to it.

Type Why it exists The rule that goes with it
PeriodRecord a confirmed period, with its source and whether it is confirmed a record's source is kept; edits are recorded, never silent
SpottingRecord spotting, tracked separately must not start a cycle or reset one
CycleRecord derived interval between two confirmed starts derived, never stored as truth — toCycles() recomputes from the period records on every read, so an edit cannot leave a stale interval behind it
PredictionRecord a snapshot taken before the outcome is known this is what makes accuracy measurable at all; never overwritten in place
NotYetObservation the user said the period had not started by a date a censoring observation — the forecast is re-conditioned on it, not shifted by +1 day
Interval one gap between confirmed starts, with its recency weight and whether it looks questionable derived per calculation, never stored. A questionable interval is down-weighted, never dropped — §12 step 2
UserPreferences notification privacy, reminder time, lock, theme, ads entitlement lives in DataStore, never in the cycle database — see below

Why settings are not in the database

core/datastore could have been two more Room tables. It is not, and the reason is a deletion semantic rather than a taste in storage.

Delete My Data removes the health history and must leave the settings alone. A user exercising that control has not asked to have notification privacy returned to a default they did not choose — handing back a weaker setting at the exact moment somebody is reaching for a privacy control is the worst possible time to do it. Separate stores make that the easy implementation rather than the one you have to remember.

UserPreferencesRepository takes a DataStore rather than a Context, which is what lets its tests run on the JVM against a temporary file. The Android instance is supplied by DI at the app layer — the only place that should know where a file lives.

The prediction engine

PersonalPredictionEngine (modelVersion personal-1) is what the app ships. It keeps a discrete probability distribution over candidate start dates rather than a date with a margin bolted on, and everything the product needs falls out of that one structure: the most likely date is its mode, the window is the narrowest span holding 80% of its mass, and a "Not yet" is the distribution being conditioned on what the user just said. A date-plus-margin design cannot express that last one, which is why §13 is the reason for the shape.

Five decisions, each measured rather than assumed:

Decision Why
Weighted median, not mean §51's outlier history moves a mean to 31.7 and leaves a median at 29
Laplace, not normal cycles have heavy tails; under a bell curve a period four days late is nearly impossible, so the model stays confidently wrong
Spread is the larger of MAD and mean-AD a user alternating 25 and 37 has half her deviations at zero, so MAD alone reads a wildly variable cycle as perfectly consistent — a test caught exactly that
Recency trusted in proportion to consistency recency weighting assumes the recent past predicts the near future, which is false for a variable user; weighting it equally cost three days on the §51 variable fixture
Trend damped, with a floor relative to the user's own spread a fixed one-day floor fired on a 42-day-cycle history whose medians differed by one day and turned an exact forecast into a wrong one

BaselinePredictionEngine stays in the tree as the control. EngineComparisonTest scores both over the §51 fixtures on every build, so "better" is a number. At the swap: mean absolute error 0.67 against 1.00, and the window contained the actual start 9 times out of 9 against 7.

Coverage is the measure, not width. The personal engine's windows are wider, and where they are wider they are right to be — a history with a suspected missing period is genuinely less certain, and the baseline answers it with a two-day window and misses. What a window promises is that the period starts inside it; an engine keeping that promise 7 times in 9 has a broken promise, not a tight forecast. The comparison test asserts coverage, with a ceiling so the trivial cheat of answering "some time this month" still fails.

Unusual is relative to the user, never to a constant

IntervalAnalysis decides whether a gap is odd by comparing it to a robust centre of this user's own intervals. A global rule — "over 40 days is suspicious" — gets exactly one group wrong, and it is the group whose cycles are already unusual: the person this product exists for, and the one most tired of apps assuming she is average. 45 days is unremarkable at a usual of 43 and worth questioning at a usual of 29.

Two flags, because they earn different responses. A gap near a whole multiple of the usual is a probable missed entry and produces §14's question. A gap merely far from usual is down-weighted and left alone — asking about it would be the over-questioning §25 warns against.

Never secretly modify health history. A gap that looks like a missing entry (§14) produces a question, not a correction. That is an architectural constraint as much as a UX one: nothing in the data layer may write a PeriodRecord the user did not confirm.

Migrations

Room migrations are numbered, tested, and listed in this document — one row per migration, added in the same commit as the migration itself. The template this repository came from records why: a manual's migration table sat six behind, and every reader in between trusted it.

Version What changed Migration Guard
1 initial schema: period_records, spotting_records, prediction_records, not_yet_observations — (first version) SchemaTest + scripts/schema-guard.sh

Room's exported schemas live in core/database/schemas/ and are committed, so a migration can be tested against the real previous schema rather than a remembered one.

The trap in this table, and the guard that closes it

Room regenerates the schema export during compilation. Change an entity without bumping PeriodDatabase.VERSION and Room silently overwrites schemas/…/1.json to match — so every in-process check compares two copies of the new truth and passes. This was not reasoned about; it was proved, by adding a column and watching the whole unit suite stay green while the committed schema quietly changed underneath it.

The failure that produces on a device is Room cannot verify the data integrity — a crash on update, in front of a user, after shipping.

scripts/schema-guard.sh is the guard, and it works by asking git, which is the one party Room cannot overwrite: an already-committed schema file that now differs means an entity changed under a shipped version. It runs in .githooks/pre-commit whenever an entity or the schema directory is staged.

So: adding a row to this table is part of changing a schema, not tidying up afterwards. The version bump, the migration, the new schema file and this row belong in one commit.

Documents here

  • GUARDS.md — how to write a check that actually checks. Read it before adding a structural test or a probe.

What ships in this folder

Nothing. This project took six scripts from the template into scripts/, and ../TOOLS.md explains why the rest are absent and where the menu is.

Path What it is
scripts/secrets.sh credential shapes in a staged diff — the one that stops a keystore reaching a commit
scripts/doc-claims.sh every file a document names must exist; --covers asks the inverse
scripts/doc-triggers.py which documents a pending change fires, read from the Governs: headers
scripts/commit-mine.sh commits only the paths you name, by pathspec, after the secret scan
scripts/forgejo-issue.py files and closes issues in the tracker convention, with every rule of it as a check
scripts/prove-guard.sh breaks what a guard protects and requires the guard to go red
scripts/schema-guard.sh a Room entity may not change without the version changing with it — asks git, because Room overwrites the export during the build
.githooks/ pre-commit, commit-msg, post-commit — see githooks/README.md

What does not belong here

A note on drift

Architecture docs go stale faster than any other kind, because code changes under them silently. That is what the Review trigger above is for, and why it names a new Gradle module and a new Room migration specifically: those are the two changes here that make this document wrong without touching it.