From 615043db4452b56aa4f971554d92b793a59597e8 Mon Sep 17 00:00:00 2001 From: Ripley Date: Tue, 18 Aug 2026 02:08:12 -0500 Subject: [PATCH] fix(infra): queuenorth.com is the production origin, and it is this deployment MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit _null confirmed queuenorth.com as the permanent public origin, which settles #212 and exposed a wrong claim written earlier the same day. docs/OPERATIONS.md said queuenorth.com "is not this deployment ... do not diagnose against it". That came from a DNS lookup and an assumption. It is this deployment: both hostnames serve the identical bundle and this server's own /api/health shape, 24.41.108.95 is this network's own public IP, and both reach qn-website-dev on nebula — queuenorth.com through nginx-proxy-manager on thor/exodus, qn.isnull.dev through Cloudflare. Two front doors, one container, no non-production environment. That is the worst direction for a runbook to be wrong in, so the correction quotes the wrong sentence rather than replacing it silently. The QA Round 0 table likewise gained the production observations as extra rows instead of having its originals rewritten. Dockerfile: the CORS_ORIGIN fallback was '*'. The server sets credentials:true and browsers reject '*' with credentials outright, so that fallback would have broken every form rather than over-permitting. Now the real origin. healthcheck.sh and preflight.sh now watch production by default, with the second front door reachable through their env overrides — the two ingresses terminate TLS in different places and can rot independently. Also fills a gap adoption left explicitly undone: the deploy path is a Portainer stack, id 58 on nebula, found from the container's own compose labels. OPERATIONS.md documents it, including that the stack file is a separate copy from this repository's docker-compose.yml and the two have already drifted. That drift is all that remains of #212 — one trailing slash on line 21 of the stack file. Left in place: nothing is broken today, and fixing it recreates the container and takes both front doors down together. Co-Authored-By: Claude Opus 5 (1M context) --- Dockerfile | 5 +- README.md | 4 +- docs/OPERATIONS.md | 108 +++++++++++++++++++++++----- docs/TOOLS.md | 7 +- docs/architecture/README.md | 12 ++-- docs/history/DEVELOPMENT_LOG.md | 52 ++++++++++++++ docs/history/HISTORY.md | 3 +- docs/planning/PROJECT_PLAN.md | 2 +- docs/qa/ClaudeReport.md | 16 ++++- docs/security/SECURITY.md | 12 +++- docs/security/SECURITY_CHECKLIST.md | 2 +- scripts/check-env.sh | 2 +- scripts/healthcheck.sh | 7 +- scripts/preflight.sh | 6 +- 14 files changed, 199 insertions(+), 39 deletions(-) diff --git a/Dockerfile b/Dockerfile index 19d5615..0a1db83 100644 --- a/Dockerfile +++ b/Dockerfile @@ -45,7 +45,10 @@ RUN addgroup -g 1001 -S nodejs && \ ENV NODE_ENV=production ENV SERVER_PORT=3001 ENV RATE_LIMIT_PER_MINUTE=5 -ENV CORS_ORIGIN=* +# The production origin, not '*'. A wildcard here is not merely loose: the +# server sets credentials:true, and browsers reject '*' with credentials +# outright — so the fallback would break every form rather than over-permit. +ENV CORS_ORIGIN=https://queuenorth.com ENV LOG_LEVEL=info ENV ZOHO_FORWARDING_MODE=webtolead ENV ZOHO_WEBTOLEAD_ENABLED=false diff --git a/README.md b/README.md index 13cf99a..05f5bb9 100644 --- a/README.md +++ b/README.md @@ -14,7 +14,9 @@ Review trigger: Any change to the stack, the route list, the Docker or The business website for Queue North Technologies — an 8x8 Certified Partner delivering UCaaS, contact centre, networking and managed support for SMB and enterprise organisations. Live at -[qn.isnull.dev](https://qn.isnull.dev). +**[queuenorth.com](https://queuenorth.com)**, with +[qn.isnull.dev](https://qn.isnull.dev) reaching the same container through a +second ingress. ## Objective diff --git a/docs/OPERATIONS.md b/docs/OPERATIONS.md index 9469c86..d0790aa 100644 --- a/docs/OPERATIONS.md +++ b/docs/OPERATIONS.md @@ -28,7 +28,9 @@ from here — goes unanswered. | | | | --- | --- | -| Public origin | `https://qn.isnull.dev`, fronted by Cloudflare | +| **Public origin** | **`https://queuenorth.com`** — the production front door, and the one customers use | +| Second front door | `https://qn.isnull.dev`, the same container reached through Cloudflare | +| `www` | 301s to `https://queuenorth.com/` | | Host | **nebula** (`192.168.1.11`) — the same box as the Forgejo instance | | Container | `qn-website-dev` | | Image | `dream.scheller.ltd/null/queue-north-website:dev` | @@ -37,9 +39,38 @@ from here — goes unanswered. | Logs volume | `qn-website-dev_queuenorth-logs` → `/app/logs` | | Runs as | UID 1001, non-root | -`queuenorth.com` also resolves and returns 200. **It is not this deployment** — -it is the customer's own domain and is not served by this container. Do not -diagnose against it. +### Two front doors, one container + +**Both hostnames are this application, and there is only one of it.** Verified +2026-08-18: both serve the identical bundle (`assets/index-pTFwovIx.js`) and both +answer `/api/health` with this server's exact `{"status":"ok","db":"ok"}` shape. + +```text +queuenorth.com ──► 24.41.108.95 (this network's own public IP) + └─► nginx-proxy-manager on thor/exodus (192.168.1.14) + └─┐ + ├──► qn-website-dev on nebula (192.168.1.11:3001) +qn.isnull.dev ──► Cloudflare ─┘ +``` + +`www.queuenorth.com` 301s to the apex. Both origins carry HSTS and redirect +plain HTTP. + +**This matters during an incident in two directions.** If only one hostname is +failing it is the ingress, not the app — compare the two before touching the +container. And if you restart that container you are restarting *production*, +not a dev alias, whatever the `-dev` in its name and its `:dev` image tag +suggest. + +> **An earlier version of this document said the opposite** — that +> `queuenorth.com` "is not this deployment ... do not diagnose against it". That +> was written on 2026-08-18 from a DNS lookup and an assumption, and it was +> wrong. It is corrected here rather than quietly, because a runbook that sends +> somebody away from the failing host is worse than one that says nothing. + +**There is no non-production environment.** These are two doors to the same +room. `npm run docker:test` runs the image locally, which is the closest thing +that exists. ## Where errors go @@ -166,25 +197,66 @@ change anything. the marketing pages keep serving while every form submission is failing. 5. **Did something change recently?** The tracker and `git log` — a deploy, a config edit, a Cloudflare rule. -6. **Is it Cloudflare rather than us?** Compare a request to the origin on nebula - against one to `qn.isnull.dev`. Cloudflare fronts everything here, including - the Forgejo instance on the same host. +6. **Is it the ingress rather than the app?** The two front doors take + completely different paths — `queuenorth.com` through nginx-proxy-manager on + thor/exodus, `qn.isnull.dev` through Cloudflare. If one answers and the other + does not, the container is fine and the problem is in front of it: + + ```bash + curl -s https://queuenorth.com/api/health + curl -s https://qn.isnull.dev/api/health + ``` + + Both failing while `scripts/status.sh` says healthy points at the LAN path or + the host itself. ### Rolling forward or back -**This is the gap in this runbook and it is named rather than papered over.** +**Publishing and deploying are two acts.** `npm run docker:push` builds the +image and pushes `queue-north-website:dev` to the Forgejo registry. That changes +nothing about what is running. -Publishing is `npm run docker:push`, which builds and pushes -`queue-north-website:dev` to the Forgejo registry. How nebula then *moves* to the -new image — a compose file on the host, a Portainer stack, a watchtower, a manual -`docker pull && docker compose up -d` — is not written down anywhere in this -repository, and was not recoverable from it on 2026-08-18. +**Deploying is a Portainer stack redeploy.** Established 2026-08-18 by reading +the container's own compose labels, which is why it is written down here rather +than re-derived: -Until somebody writes it here, a rollback is: find the previous image digest in -the registry, and do on nebula whatever it is that normally happens. That is not -a procedure. The template's `deploy.py` was deliberately **not** adopted for this -reason — a deploy script pointed at a path nobody has confirmed is worse than no -deploy script. +| | | +| --- | --- | +| Portainer | `https://192.168.1.11:9443` (nebula), API key in `~/.openclaw/credentials/portainer.md` | +| Stack | **id 58**, name `qn-website-dev`, endpoint 3, compose | +| Stack file on disk | `/data/compose/58/docker-compose.yml` on nebula | + +**The stack file is NOT this repository's `docker-compose.yml`.** They are two +separate files that happen to look alike, and they have already drifted: the +repository commits `CORS_ORIGIN=https://queuenorth.com` and line 21 of the stack +file says `https://queuenorth.com/`, with a trailing slash. Editing the one in +git changes nothing about production. That is the single most important sentence +in this section. + +Read the deployed file, and the environment the container actually got: + +```bash +KEY=$(grep -m1 'API Key' ~/.openclaw/credentials/portainer.md | grep -oE 'ptr_[^`]+') +curl -sk -H "X-API-Key: $KEY" https://192.168.1.11:9443/api/stacks/58/file +bash scripts/status.sh # what is running, and its image digest +``` + +To move to a newly published image, or to change an environment value, redeploy +the stack — through the Portainer UI, or its API with `PUT /api/stacks/58` +carrying the full stack file and `pullImage: true`. + +**Redeploying recreates the container, which is a restart of the customer's live +site.** Both public front doors go down together for the few seconds it takes, +because they reach the same instance. There is no non-production environment to +rehearse against; `npm run docker:test` running the image locally is the closest +thing that exists. Take the backup first — `bash scripts/backup.sh` — because +the database is in a named volume that a careless `down -v` would remove. + +The template's `scripts/deploy.py` is built for exactly this shape and reads +`DEPLOY_STACK_ID`, `DEPLOY_IMAGE`, `DEPLOY_CONTAINER` and `DEPLOY_SITE_URL` from +the environment. **It was declined on adoption** because none of the above was +known then. It is adoptable now, and worth taking the next time a deploy is +needed rather than doing it by hand twice more. **Write the incident down afterwards**, in `docs/history/DEVELOPMENT_LOG.md` with the date, and file what broke as an issue with a severity label. An incident diff --git a/docs/TOOLS.md b/docs/TOOLS.md index 90499e8..d457168 100644 --- a/docs/TOOLS.md +++ b/docs/TOOLS.md @@ -137,9 +137,14 @@ standing gap. A liveness check by hand, when you want one without the script: ```bash -curl -s https://qn.isnull.dev/api/health # {"status":"ok","db":"ok","timestamp":"…"} +curl -s https://queuenorth.com/api/health # production +curl -s https://qn.isnull.dev/api/health # same container, other ingress +# both -> {"status":"ok","db":"ok","timestamp":"…"} ``` +If those two disagree, the container is fine and the problem is in front of it. +`docs/OPERATIONS.md` has the topology. + ## Two checks that are run by hand Neither is adopted into `scripts/`, so neither runs in `verify.sh`. Both are diff --git a/docs/architecture/README.md b/docs/architecture/README.md index 46edff8..53cf22f 100644 --- a/docs/architecture/README.md +++ b/docs/architecture/README.md @@ -19,11 +19,11 @@ reverse proxy inside the container, and no second runtime. ```text browser | - v -Cloudflare ── qn.isnull.dev - | - v -Express (server/index.js, port 3001) ← the only process + +── queuenorth.com ──► nginx-proxy-manager (thor/exodus) ┐ PRODUCTION + | │ + +── qn.isnull.dev ───► Cloudflare ─────────────────────────┤ + v +Express (server/index.js, port 3001) on nebula ← the only process, one instance | +--> dist/ prerendered HTML + the React bundle, served static +--> /api/health liveness, and a real SELECT against SQLite @@ -146,7 +146,7 @@ each row says what it does *here*. | `scripts/doc-triggers.py` | which documents a pending change fires, read from the `Governs:` headers. Run it before committing, not after | | `scripts/forgejo-issue.py` | files and closes issues in the tracker convention, refusing malformed ones before they are filed | | `scripts/status.sh` | what is running on **nebula** as `qn-website-dev`, its version and its restart count. Read-only | -| `scripts/healthcheck.sh` | a liveness tick against `qn.isnull.dev`, asserting HTTP 200 **and** `"status":"ok"` — a 503 with a JSON body is a real answer, not an outage | +| `scripts/healthcheck.sh` | a liveness tick against `queuenorth.com`, asserting HTTP 200 **and** `"status":"ok"` **and** `"db":"ok"` — a 503 with a JSON body is a real answer, not an outage. `HEALTHCHECK_BASE_URL` points it at the other front door | | `scripts/preflight.sh` | headers and TLS against the live origin. No `--auth` checks: there are no accounts | | `scripts/backup.sh` | a verified SQLite dump. Its ENGINE block was rewritten for better-sqlite3's online `.backup()` — see below | | `scripts/restore-check.sh` | restores the newest dump into a scratch file, runs `PRAGMA integrity_check`, counts tables, and **times it**. A backup nobody has restored is a guess | diff --git a/docs/history/DEVELOPMENT_LOG.md b/docs/history/DEVELOPMENT_LOG.md index 0fdef49..e3fe494 100644 --- a/docs/history/DEVELOPMENT_LOG.md +++ b/docs/history/DEVELOPMENT_LOG.md @@ -46,6 +46,58 @@ sequence would have implied more. ## Entries +### 2026-08-18 — queuenorth.com confirmed as the permanent origin, and a wrong runbook claim corrected + +`_null` answered the one question left open by the adoption: **queuenorth.com is +the permanent public origin**, with `qn.isnull.dev` kept as a second ingress. + +**The answer exposed a mistake made earlier the same day.** The first draft of +`docs/OPERATIONS.md` said queuenorth.com "is not this deployment — it is the +customer's own domain and is not served by this container. Do not diagnose +against it." That was written from a DNS lookup and an assumption, and it was +wrong in the worst direction a runbook can be wrong: it sent whoever read it +away from the production host during an incident. + +What is actually true, established by request rather than inference: both +hostnames serve the identical bundle and this server's own `/api/health` shape; +`24.41.108.95` is this network's own public IP; `queuenorth.com` reaches the +container through nginx-proxy-manager on thor/exodus and `qn.isnull.dev` reaches +it through Cloudflare. **Two front doors, one container, and no non-production +environment** — restarting it restarts the customer's live site whatever the +`-dev` in its name suggests. + +The correction is written into `OPERATIONS.md` as a correction, with the wrong +sentence quoted, rather than silently replaced. Same for the QA Round 0 table, +which gained the production observations as extra rows instead of having its +originals rewritten. + +**The deploy path is no longer a gap.** Adoption left it explicitly undone — +"how nebula moves to a new image is not written down anywhere". It is a +**Portainer stack, id 58 on nebula**, file at `/data/compose/58/docker-compose.yml`, +found by reading the container's own compose labels. `OPERATIONS.md` now carries +it, including the sentence that matters most: that stack file is a separate copy +from this repository's `docker-compose.yml` and the two have already drifted. + +That drift is #212 and it is exactly one character — line 21 of the stack file +has `CORS_ORIGIN=https://queuenorth.com/` with a trailing slash. Left in place +deliberately: nothing is broken today (the forms are same-origin, and the slash +makes the header match no origin, so it fails closed), and fixing it means +recreating the container and taking both front doors down together. It belongs +in the next deploy that happens for another reason. + +`scripts/deploy.py` was declined at adoption because none of this was known. It +is built for exactly this shape and is worth taking next time. + +- **Closed:** none. #212 stays open — the repository half is done and the + production half is one line in Portainer. +- **Next action:** Unchanged — start `Batch 10 — Accessibility WCAG 2.1 AA` + with #189, walking the pages keyboard-only at 320px before fixing anything. + Fold the #212 one-character fix into whatever deploy happens next rather than + restarting production for it. +- **Blockers:** None on this work. Six issues remain blocked on the site owner + (#68, #110, #162, #213 in Batch 13; #69, #70 in Batch 14), two of which are + unsubstantiated claims live on the public site since May. + ### 2026-08-18 — Adopted the project template Merged `~/.openclaw/Projects/Template` into this repository. It was the diff --git a/docs/history/HISTORY.md b/docs/history/HISTORY.md index 6b6f457..f85310f 100644 --- a/docs/history/HISTORY.md +++ b/docs/history/HISTORY.md @@ -13,7 +13,7 @@ Review trigger: A decision reversed, a direction abandoned, or a rewrite Queue North Technologies had a working website: one large `index.html`, a hand-written stylesheet, and `main.js` doing hash-based routing. The content was good and the presentation was early-2000s. This repository is the rebuild — a -Vite + React SPA with an Express API behind it, live at `qn.isnull.dev`. +Vite + React SPA with an Express API behind it, live at `queuenorth.com`. That much is in the code. What is not in the code is which walls were already walked into, and that is what this file is for. @@ -38,6 +38,7 @@ sentence belongs somewhere else. | 2026-06-14 | Zoho switched from the REST/OAuth path to **WebToLead** as the default mode | `ZOHO_FORWARDING_MODE=webtolead`. The OAuth path works and is kept as a standby; WebToLead needed no token lifecycle for what is a single lead form | | 2026-08-01 | Every route is prerendered at build time, and the privacy policy has a server-rendered fallback | Meta's crawler does not run JavaScript. See `docs/architecture/README.md` | | 2026-08-18 | Adopted the project template; **retired phase-versioning** | See below | +| 2026-08-18 | **`queuenorth.com` is the permanent public origin**, with `qn.isnull.dev` kept as a second ingress to the same container | Asked and answered by `_null` while closing out the template adoption. It settles #212 and it corrected a mistake made earlier the same day: the first draft of `docs/OPERATIONS.md` asserted `queuenorth.com` was *not* this deployment, written from a DNS lookup rather than a request. Both hostnames serve the identical bundle and this server's own `/api/health` shape; `24.41.108.95` is this network's own public IP | | 2026-08-18 | The Forgejo repository was made **private** | Adoption committed the previously-gitignored internal documents into `docs/`. Those hold the agent narrative and internal findings about the client's own marketing claims (#108, #110), which is not material for a public tree. `null/fruit-fall` was already private and reports to the Command Center normally, so nothing was lost by it | ## What was tried and dropped diff --git a/docs/planning/PROJECT_PLAN.md b/docs/planning/PROJECT_PLAN.md index 8db6240..b195f42 100644 --- a/docs/planning/PROJECT_PLAN.md +++ b/docs/planning/PROJECT_PLAN.md @@ -82,7 +82,7 @@ re-litigated. | Storage | SQLite via better-sqlite3 | Two tables, low write volume, one process. Synchronous by design, which suits a request-scoped insert | | CRM | Zoho, WebToLead by default | Fire-and-forget behind SQLite, so a CRM outage cannot cost a lead | | Anti-abuse | Google reCAPTCHA v3 plus a honeypot field | The forms are the only writable surface | -| Hosting | Docker on **nebula**, container `qn-website-dev`, image `dream.scheller.ltd/null/queue-north-website:dev`, fronted by Cloudflare at `qn.isnull.dev` | See `docs/OPERATIONS.md` | +| Hosting | Docker on **nebula**, container `qn-website-dev`, image `dream.scheller.ltd/null/queue-north-website:dev`. Served at **queuenorth.com** through nginx-proxy-manager, and at `qn.isnull.dev` through Cloudflare — two ingresses, one container | See `docs/OPERATIONS.md` | ## Success looks like diff --git a/docs/qa/ClaudeReport.md b/docs/qa/ClaudeReport.md index c294828..6c27892 100644 --- a/docs/qa/ClaudeReport.md +++ b/docs/qa/ClaudeReport.md @@ -21,8 +21,8 @@ Review trigger: Any QA round run - **Last QA round:** Round 0 — 2026-08-18, no round has been run - **Last verified build SHA:** `ac6f5d1` on dev, deployed to nebula as `queue-north-website:dev` - **Last tested device / environment:** none under this plan. The container's own - healthcheck and a manual `curl` against `qn.isnull.dev` are the only automated - observations of the running site + healthcheck and manual `curl`s against both public origins are the only + automated observations of the running site - **Overall status:** No QA round has been run against this site under any written plan, and there is no automated test suite to fall back on — the only gate is that the build compiles. The site is up, healthy and taking real form @@ -69,14 +69,24 @@ running system on this date, and how: | Observation | How it is known | | --- | --- | +| `https://queuenorth.com/` returns 200 | `curl`, once | | `https://qn.isnull.dev/` returns 200 | `curl`, once | -| `/api/health` returns `{"status":"ok","db":"ok"}` | `curl`, once | +| `/api/health` returns `{"status":"ok","db":"ok"}` on **both** origins | `curl`, once each | +| Both origins serve the identical bundle `assets/index-pTFwovIx.js` | `curl`, once each — this is what established they are one container, not two | +| `www.queuenorth.com` 301s to the apex; both origins 301 plain HTTP and carry HSTS | `curl -I`, and `scripts/preflight.sh` clean against each | | Container `qn-website-dev` on nebula is healthy, restart count 0 | `docker inspect` | | The lead database holds 3 tables and took writes on 2026-08-18 | `sqlite_master` count and the file mtime | | Security headers, CSP and HSTS are present on the live origin | response headers | None of that is a QA pass. It is evidence that the thing is running. +**Added later the same day.** The table above originally recorded only +`qn.isnull.dev`, because this session had not yet established that +`queuenorth.com` is the production front door for the same container — an +earlier draft of `docs/OPERATIONS.md` asserted the opposite. The extra rows are +what was actually checked once that was known, and they are added rather than +substituted so the record shows the correction happening. + The first real round should start with pass B — the two form paths end to end, including whether a lead actually lands in Zoho — because that is the product's entire purpose and nothing has ever verified it end to end from outside. diff --git a/docs/security/SECURITY.md b/docs/security/SECURITY.md index 34acf76..91e8fc3 100644 --- a/docs/security/SECURITY.md +++ b/docs/security/SECURITY.md @@ -98,9 +98,15 @@ deleted rather than left as unticked boxes. ## Data in transit -TLS everywhere, terminated at Cloudflare in front of `qn.isnull.dev`. The server -additionally redirects HTTP to HTTPS in production on `x-forwarded-proto`, and -sends HSTS with a one-year max-age and `includeSubDomains`. +TLS everywhere, on both front doors and terminated in different places — +nginx-proxy-manager for `queuenorth.com`, Cloudflare for `qn.isnull.dev`. Both +carry HSTS with a one-year max-age and `includeSubDomains`, and both 301 plain +HTTP; verified with `scripts/preflight.sh` against each. The server additionally +redirects HTTP to HTTPS in production on `x-forwarded-proto`. + +**Two terminators means two places TLS can rot.** `preflight.sh` defaults to the +production origin, so the Cloudflare-fronted one is only checked when somebody +passes `PREFLIGHT_ORIGIN`. Helmet sets the CSP, and it is deliberately narrow: `default-src 'self'`, with named exceptions only for Zoho's form host, Google reCAPTCHA and Google Fonts. diff --git a/docs/security/SECURITY_CHECKLIST.md b/docs/security/SECURITY_CHECKLIST.md index 00b3d97..ea1409a 100644 --- a/docs/security/SECURITY_CHECKLIST.md +++ b/docs/security/SECURITY_CHECKLIST.md @@ -23,7 +23,7 @@ is short enough to finish rather than a document to skim. - [ ] `npm run build && bash scripts/secrets.sh --built dist/` is clean — proves the *bundle* is clean, which the tracked scan cannot tell you - [ ] `bash scripts/check-env.sh --file .env` exits 0 — proves every variable the server reads is set and shaped right, before it reads them. **Exit 2 is not a pass** - [ ] `npm audit` shows no high or critical advisory in production dependencies — proves no known-exploitable code ships -- [ ] `bash scripts/preflight.sh` against `qn.isnull.dev` is clean — proves headers, CSP and TLS survived the deploy +- [ ] `bash scripts/preflight.sh` is clean against **both** front doors — `queuenorth.com` by default and `PREFLIGHT_ORIGIN=https://qn.isnull.dev` for the other. Proves headers, CSP and TLS survived the deploy on two separate ingresses that can rot independently - [ ] The reCAPTCHA key in `dist/` is the **site** key, not the secret key — proves the one inlined value is the one that is safe to inline ## Standing checks diff --git a/scripts/check-env.sh b/scripts/check-env.sh index 95a569c..92936e3 100755 --- a/scripts/check-env.sh +++ b/scripts/check-env.sh @@ -210,7 +210,7 @@ SPEC=( # --- Always. Wrong here and the site is wrong for everybody. --- "NODE_ENV|required|one-of:development,production,test|CSP relaxes in dev and the HTTP-to-HTTPS redirect only fires in production; the wrong value ships dev CSP to the public origin, or redirects a local dev server into a loop" "SERVER_PORT|required|port|Express binds 3001 by default and the reverse proxy in front of nebula is pointed at whatever this says; a mismatch is a site that is running and unreachable" - "CORS_ORIGIN|required|url|defaults to https://queuenorth.com, which is NOT this deployment. Wrong here and every form submission from qn.isnull.dev is blocked by the browser while the server logs nothing wrong" + "CORS_ORIGIN|required|url|the production origin, https://queuenorth.com, with NO trailing slash — a slash matches no real browser origin and silently blocks every cross-origin form post while the server logs nothing wrong. Both public hostnames reach the same container, so the site's own forms are same-origin and unaffected; this is about anything that is not" "LOG_LEVEL|optional|one-of:error,warn,info,debug|defaults to info. At error you lose the request log, which is the only record of traffic this project keeps" "RATE_LIMIT_PER_MINUTE|optional|int:1..1000|defaults to 5 per IP across all of /api. Unset is fine; set to something huge and the two POST endpoints are open to a bot overnight" diff --git a/scripts/healthcheck.sh b/scripts/healthcheck.sh index 3cfe63a..b754f0d 100755 --- a/scripts/healthcheck.sh +++ b/scripts/healthcheck.sh @@ -63,7 +63,12 @@ set -euo pipefail # a config is wrong, and substituting production for it would report the health # of a site nobody asked about — quietly, and only where somebody was trying to # point this somewhere else. -BASE_URL="${HEALTHCHECK_BASE_URL:-https://qn.isnull.dev}" +# The PRODUCTION origin — the one customers use. qn.isnull.dev is the same +# container through a different ingress, and is worth checking separately when +# you are trying to tell the app apart from what is in front of it: +# +# HEALTHCHECK_BASE_URL=https://qn.isnull.dev bash scripts/healthcheck.sh +BASE_URL="${HEALTHCHECK_BASE_URL:-https://queuenorth.com}" TIMEOUT="${PRIVACY_TIMEOUT:-15}" stamp() { date -Is; } diff --git a/scripts/preflight.sh b/scripts/preflight.sh index 0db04b1..beb3d8b 100755 --- a/scripts/preflight.sh +++ b/scripts/preflight.sh @@ -77,7 +77,11 @@ die() { printf '\033[1mpreflight:\033[0m %s\n' "$*" >&2; exit 2; } # Empty on purpose: a default here is a script that attacks whatever origin the # project it was copied from happened to use. # --------------------------------------------------------------------------- -ORIGIN="${PREFLIGHT_ORIGIN:-https://qn.isnull.dev}" +# The PRODUCTION origin. The second front door goes through completely +# different infrastructure, so its headers and TLS are a separate question: +# +# PREFLIGHT_ORIGIN=https://qn.isnull.dev bash scripts/preflight.sh +ORIGIN="${PREFLIGHT_ORIGIN:-https://queuenorth.com}" # Paths used only by --auth. Unset means that check reports "not configured" # rather than guessing at /login, which on the wrong app is a POST to something