# Operations — ``` Status: Current Owner: Last reviewed: Governs: what watches this in production, and what to do when it stops Review trigger: A new alert or a new place errors are sent; a change to the backup schedule or the restore procedure; any incident whose first ten minutes were spent working out what to run ``` > *(only for a deployed service.)* A library, a local tool or a static site with > no backend deletes this file. Do not keep it with the headings unanswered — an > empty runbook reads as one nobody wrote, which is worse than one that never > applied. ## Why this exists Every other document here describes what the project *is*. This one describes what happens when it stops being that at three in the morning, and it is the only document written for somebody who is not thinking clearly. The failure it prevents is specific: **the first ten minutes of an incident spent working out which host, which container, and which command**, while the thing everybody actually wants to know — is it down, or is it just unreachable from here — goes unanswered. ## Where errors go , alerts to " is useful at 3am and "we have error tracking" is not.> **A liveness tick is not this.** `scripts/healthcheck.sh` answers *is it up*. Error tracking answers *is it working*, and the two disagree constantly: a service returning 500 to every request is up. If this project has no error tracking, say so here in one line rather than leaving the section blank — a gap somebody chose is a different thing from a gap nobody noticed. ## What alerts, and to whom | Signal | Where it goes | Who acts on it | | --- | --- | --- | | <> | <> | <> | **An alert nobody acts on trains everybody to ignore the ones that matter.** If a signal here has fired more than twice without anyone doing anything, it is either wrong or it is not an alert — fix the threshold or move it to a dashboard, and say which happened. ## Backups | | | | --- | --- | | Schedule | | | What is captured | | | Where it lands | | | Retention | | | **Last verified restore** | **, by ** | `scripts/backup.sh` proves a dump is readable before trusting it. `scripts/restore-check.sh` proves it can be restored, which is a different claim: a file that parses is not a database you can get back. **A backup nobody has restored is a guess.** The date above is the only line in this table that says otherwise, and it is a fact with an owner — if it is empty or a year old, this project does not have backups, it has files. ## Rate limits and cost ceilings *(precautionary — no incident here has yet come from this.)* | Endpoint or job | What it costs per call | What bounds it | | --- | --- | --- | | <> | <> | <> | **Where does the spend alert go, and at what number?** <> A ceiling nobody is told about is a bill, not a limit. ## It is down — what now In this order. Each step is a command that answers one question, and none of them change anything. 1. **Is it actually down, or unreachable from here?** `bash scripts/status.sh` — it distinguishes those two, and reports the restart count. A climbing restart count with a short uptime is a crash loop, not a healthy service, and reads as "running" in every tool that shows only a state. 2. **What is it saying?** `bash scripts/status.sh --logs 200`. 3. **Is the deployed version the one you think?** `bash scripts/status.sh --deployed-version`. Non-zero means it could not be read — stop, rather than assuming. 4. **Did something change recently?** The tracker and `git log` — a deploy, a migration, a config edit. 5. **Roll forward or back:** `python3 scripts/deploy.py` to move the stack to a published image. Publishing and deploying are separate acts; this is the second one. **Write the incident down afterwards**, in `docs/history/DEVELOPMENT_LOG.md` with the date, and file what broke as an issue with a severity label. An incident nobody recorded happens again with the same surprise.