ClaudeQACoverage has carried "long-horizon accuracy is tested only at the
unit level" as a standing gap for four rounds. The simulated-history half
is now closed by LearningCurveTest, and the entry says so -- along with
what it found, since a harness that changed nothing would be a weaker
claim than one that caught two real faults.
The other half stays open and is stated plainly: a simulation exercises
the engine, not the app. Nothing here proves a real install accumulates
those cycles correctly over a year, that snapshots survive updates and
time-zone changes, or that the accuracy card shows what the engine
computed. That needs elapsed time on a device.
Development log entry per WORK_CYCLE step 6, including the two
corrections this batch made to its own plan: the coverage dip was a
40-seed artifact, and a swept constant shipped one value away from the
measured one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>