Skip to content

[316] Set tiered Cache-Control on published dashboard objects - #349

Merged
CarsonDavis merged 30 commits into
developmentfrom
316-dashboard-cache-control
Sep 9, 2026
Merged

[316] Set tiered Cache-Control on published dashboard objects#349
CarsonDavis merged 30 commits into
developmentfrom
316-dashboard-cache-control

Conversation

@CarsonDavis

@CarsonDavis CarsonDavis commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator
Category Lines added %
Tests 242 76.6%
Production code 67 21.2%
Docs 7 2.2%

Every object a publish writes now carries a Cache-Control header, so a republished dashboard reaches visitors instead of sitting stale on a cache we cannot reach.

Closes #316. PR 3 of 3 for the issue, stacked on #348 (which stacks on #330), retargeted to development as the stack merges.

A republish now reaches a customer's own cache

Previously, an operator republished a dashboard, saw the new content on our CloudFront (the publish invalidates /*), and a customer fronting that same dashboard on their own domain kept serving the old index.html.

Every object we wrote went up with no Cache-Control at all: the PUTs carried only Bucket, Key, Body, ContentLength and ContentType, and the shared-asset copies used CopyObject's default COPY directive, which carries over the source object's headers, and the upload router never set a Cache-Control on those either. With no instruction from the origin, a CloudFront cache policy falls back to its Default TTL, 86400 seconds in the managed policies and in most hand-rolled ones. The old entry page could stay pinned for a day in a distribution we have no ability to invalidate.

The fix: cacheControlForKey(key) assigns one of three tiers to every object, applied on all four write paths (uploadDirectory, uploadFile, copyPrefix, copyObjectIfExists):

  • no-cache, revalidate on every request, for index.html, build/index.html and Missions/<mission>/config.json: the entry page and the baked mission config.
  • public, max-age=31536000, immutable for build/static/{js,css,media}/… and assets/<mission>/<subdir>/uploads/…: the two classes whose filenames are content-addressed, webpack content hashes on one side and crypto.randomUUID() names that are never overwritten on the other.
  • public, max-age=300 for everything else, the keys that genuinely change in place on republish: build/static/cesium/…, public/workers/…, Missions/…/Data/mosaic_parameters.csv, and non-upload assets/….

A republish now reaches a header-honoring cache in stages: entry page and config at once, the rest within about five minutes. The two copy paths switched to MetadataDirective: REPLACE, because COPY cannot add a header the source never had; REPLACE drops the source's whole metadata set, so those copies restate ContentType as well, re-derived from the file extension.

One test guards the premise of the immutable tier: a new spec reads configuration/webpack.config.js as text and asserts a content hash in all six output-name settings, so dropping one breaks the suite instead of quietly putting a stable filename into the year-long tier.

Two un-rendered templates are no longer published

Previously, anyone hitting <dashboard>/build/index.pug or <dashboard>/public/index.html got a raw template with #{LINK_PREVIEW_TITLE}, #{AUTH} and #{user} still in it.

The publish uploads build/ and public/ wholesale, and it interpolates those placeholders into build/index.html only. The two siblings shipped exactly as they sit on disk, and uploadDirectory had no way to leave a key out.

The fix: uploadDirectory takes an optional filter(key) and publish-static.js skips those two exact keys. They are left out of the upload and out of the reported file count.

The customer doc sets a Maximum TTL and covers password rotation

Previously, docs/infrastructure/serving-a-dashboard-from-your-domain.md asked the path owner for a custom cache policy with Authorization in the cache key, all query strings, and Minimum TTL 0. It said nothing about a maximum, so a customer who set a one-day Maximum TTL would silently cap our year-long tier down to a day. Nothing told an operator what happens to a fronting cache after a dashboard's password is rotated.

The doc was written when we sent no Cache-Control at all, so there were no TTL ceilings and no revalidating tier to describe.

The fix: the setup list adds "Maximum TTL 31536000 (one year) or more", and the rationale section explains the three tiers as behavior, including what a staged republish looks like from the customer's side. It also notes that AWS's managed UseOriginCacheControlHeaders policies get the TTLs right but omit Authorization from the cache key, so a custom policy is still required. A closing paragraph tells operators to invalidate their distribution after rotating a password: the cache is keyed on Authorization, so responses fetched under the old password keep being served to anyone still presenting it, never reaching us to be rejected, for up to a year on the immutable tier.

Decisions to review

  • One-year immutable on password-gated files. public is the directive that lets a shared cache store a response to a request carrying Authorization, so cross-visitor safety rests entirely on the customer keeping Authorization in their cache key, which the doc requires and we cannot enforce. The alternative is a ceiling of hours or days, which still fixes the day-long pin and bounds the exposure after a rotation without a customer-side step.
  • The immutable tier is decided by key prefix, not by evidence of a hash in the filename. The webpack half of that claim is guarded by the new spec and the uploads half is only a code comment; the alternative is to require a hash segment in the key itself, which needs no cross-file guard. Being wrong is asymmetric, because we cannot purge a stale immutable object from a customer's edge for a year.
  • MetadataDirective: REPLACE restates only ContentType and CacheControl. Content-Encoding, Content-Disposition and x-amz-meta-* are dropped, which is exact for every type the upload router writes today; the alternatives are to HeadObject the source and carry its headers forward, or to set Cache-Control at upload time so a plain COPY suffices.
  • Our own per-dashboard distribution stays on the managed CachingOptimized policy. Its 1-second minimum TTL overrides no-cache at our edge, which does not matter because the publish already invalidates /*; the alternative is an owned cache policy in the template with minimum 0.
  • The two raw templates are kept out by an exact-key skip on a publish that only writes. A dashboard published before this keeps its copies until its bucket is emptied; the alternative is to delete those two keys on republish, which needs s3:ListBucket and s3:DeleteObject added in three IAM layers.

A republish has to show up immediately through a fronting cache we
cannot invalidate — ours or a customer's CloudFront — so every object
now carries a Cache-Control chosen by its key: the entry page and the
baked config revalidate before every use, the keys that are content-
addressed by construction are immutable, and everything else gets a
short 5-minute TTL.

Two key classes earn the immutable tier. The webpack output under
build/static/(js|css|media) is content-hashed — a premise now pinned by
tests/unit/webpackHashedOutput.spec.js, which fails if a production
naming option loses its hash. The plugin uploads under
assets/<mission>/<subdir>/uploads/ are named crypto.randomUUID() by the
upload router and never overwritten, so their keys are content-addressed
too.

The S3 copy paths (copyPrefix, copyObjectIfExists) also switch to
MetadataDirective REPLACE, since CopyObject's default COPY preserves
the source's metadata verbatim and cannot set a header the source
object never had. REPLACE means supplying ContentType too, taken from
the same extension map the uploads use.
@CarsonDavis
CarsonDavis force-pushed the 316-dashboard-cache-control branch from f0af5d7 to 10445e0 Compare September 1, 2026 14:50
# Conflicts:
#	docs/infrastructure/serving-a-dashboard-from-your-domain.md
… into 316-dashboard-cache-control

# Conflicts:
#	scripts/lib/aws-provision.js
#	tests/unit/awsProvision.spec.js
… into 316-dashboard-cache-control

# Conflicts:
#	docs/infrastructure/serving-a-dashboard-from-your-domain.md
#	src/pre/uploadKey.ts
#	tests/unit/uploadKeyClassifier.spec.js
… into 316-dashboard-cache-control

# Conflicts:
#	tests/unit/awsProvision.spec.js
#	tests/unit/uploadKeyClassifier.spec.js
…dashboard-cache-control

# Conflicts:
#	configure/src/core/upload.js
#	docs/infrastructure/serving-a-dashboard-from-your-domain.md
#	src/pre/uploadKey.ts
…dashboard-cache-control

# Conflicts:
#	tests/unit/uploadKeyClassifier.spec.js
… into 316-dashboard-cache-control

# Conflicts:
#	scripts/lib/aws-provision.js
#	scripts/publish-static.js
… into 316-dashboard-cache-control

# Conflicts:
#	configure/src/core/upload.js
#	docs/infrastructure/serving-a-dashboard-from-your-domain.md
#	scripts/lib/aws-provision.js
#	scripts/publish-static.js
#	src/pre/uploadKey.ts
#	tests/unit/awsProvision.spec.js
#	tests/unit/uploadKeyClassifier.spec.js
@CarsonDavis
CarsonDavis marked this pull request as ready for review September 9, 2026 14:29
@CarsonDavis
CarsonDavis requested a review from slesaad September 9, 2026 14:35
CarsonDavis added a commit that referenced this pull request Sep 9, 2026
## Line breakdown

| Category | Lines added | % |
|---|---|---|
| Tests | 673 | 61.7% |
| Production code | 301 | 27.6% |
| Docs | 75 | 6.9% |
| Config/build | 41 | 3.8% |
| Generated (package-lock.json) | 1 | 0.1% |
| **Total** | **1091** | |

## Overview

A published dashboard can now be served from a path on a customer's own
domain — `site.gov/tools/dashboard/` — not just from its own CloudFront
address.

Part of #316 — PR 1 of 3 (with #348 republish stack convergence and #349
cache-control tiering). Serves dashboards published after this merges;
existing dashboards pick up the new edge function via #348.

## What this does

- **Strips the customer's prefix at the edge.** The per-dashboard
CloudFront Function reads `X-Forwarded-Prefix`, validates it, and
removes it from the request path. A missing or malformed header passes
through untouched — a loud 403, never the wrong files — and the password
gate still runs first.
- **Redirects slash-less entry URLs** (`…/tools/dashboard` →
`…/tools/dashboard/`) with the query string preserved, marked
`no-store`. A character that would corrupt the rebuilt URL (a raw `&` or
`#`) is percent-escaped so the pair survives intact.
- **Single-sources the edge function.** The deployed code is read from
`infrastructure/cloudfront-function.js` at publish time instead of a
hand-copied string; tests pin it to ES5 and under CloudFront's 10 KB
cap.
- **Stores uploaded image paths slash-less** (`assets/…`, not
`/assets/…`) so they resolve inside the dashboard wherever it's mounted.
Both readers (Card renderer, Configure preview) rebase legacy
`/assets/…` values on read, so existing configs keep working.
- **Removes the other domain-root assumptions.** The PDF worker,
geologic patterns, and static-build tile/viewer paths no longer emit
`/…` or `../..` escapes; server and Docker modes are unchanged.
- **Hardens the URL readers.** Query values containing `=` survive
whole, malformed percent-encoding reads as absent instead of throwing at
startup, a valueless param reads `''` instead of the string
`"undefined"`, and bare `?_preview`/`?forcelanding` flags are
presence-based at both entry points (LandingPage now agrees with
modern.js). Pinned by a new spec.
- **Theme icons resolve through the same rule as Card images.** Both
`upload` fields share one resolver in
`src/essence/Tools/_shared/content/uploadKey.ts`; the viewer's master
image and model texture use the plain mission path instead of a
`../../../../` climb.
- **Ships the customer doc** —
`docs/infrastructure/serving-a-dashboard-from-your-domain.md`: setup
steps, a behavior table, and the why behind each setting.

## Decisions to review

- **`X-Forwarded-Prefix` is validated by deny-list** (`//`, `:`, `\`,
`/../`), not an allowlist.
- **Old `/assets/…` values are supported forever instead of migrated
once**
- **Publish reads `infrastructure/cloudfront-function.js` from the ECS
image at runtime** — nothing pins that directory into the image beyond
today's `COPY . .`; slimming the image later breaks publish at runtime.
- **`forceLanding` is accepted in both casings at every entry point**,
though nothing in the app emits the camelCase form; the alternative is
reverting the landing page to lowercase only.
CarsonDavis added a commit that referenced this pull request Sep 9, 2026
| Category | Lines added | % |
|---|---|---|
| Tests | 1538 | 52.2% |
| Production code | 1343 | 45.6% |
| Config/build | 43 | 1.5% |
| Docs | 21 | 0.7% |

Clicking Update on a published dashboard now brings its CloudFormation
stack up to the current template, so infrastructure fixes reach
dashboards that already exist. The Deployments page also stops an admin
from starting two publishes for one dashboard or deleting one while its
publish task is still running, and says why when it refuses.

Part of #316. PR 2 of 3: stacked on #330 (serve a dashboard at a
customer subpath) with #349 (cache-control) on top of this one;
retargets to `development` once #330 merges.

### Republishing a dashboard now updates its AWS infrastructure, not
just its files

Previously, an operator clicked Update on a published dashboard in
/configure → Deployments, the row went `updating` → `published`, the
bundle was replaced, the URL was unchanged, and nothing about the
dashboard's CloudFront changed. The entire update branch of
`scripts/publish-static.js` was a DescribeStacks call and a "does not
exist, publish before updating" throw.

Each dashboard's bucket, distribution and edge Function are created
once, by CreateStack at the first publish, and were never touched again.
So a dashboard already sitting behind a customer path prefix kept
serving the old edge Function and stayed broken, and a rotated
`MMGIS_DASHBOARDS_PASSWORD` never reached it. The only escape was delete
and republish, which mints a new CloudFront domain: the one thing path
owners hardcode.

The fix: Update renders the current template, which bakes in the current
password, calls UpdateStack, and polls to UPDATE_COMPLETE. Same stack,
same bucket, same URL. It logs `Converging stack 'mmgis-dashboard-N' to
the current template — this re-bakes the current dashboards password
into the auth Function.` and then `Stack '…' reached UPDATE_COMPLETE.`
When the template already matches, CloudFormation's "No updates are to
be performed" is caught and logged as `Stack '…' is already up to date.`
rather than failing the republish.

### Publishing over an existing stack either fails fast with guidance or
waits for it to settle

Previously, a publish over a stack that already existed could hang for
the full 30-minute timeout, or fail on the stack's own success status. A
stack left at ROLLBACK_COMPLETE by a failed create ended, after the
whole build, at `Stack 'mmgis-dashboard-N' reached terminal status
'ROLLBACK_COMPLETE'`: accurate, but it doesn't say what to do, and
CloudFormation cannot update such a stack at all. DELETE_IN_PROGRESS and
UPDATE_FAILED were not in the terminal list, so the poller sat on a
deleting stack until it vanished (`Stack '…' does not exist (deleted or
never created)`) and on a failed update until the 30-minute timeout. And
the wait always ran against exactly CREATE_COMPLETE, so a stack resting
at UPDATE_COMPLETE, which the converge path now produces routinely,
killed a publish re-run with `Stack '…' reached terminal status
'UPDATE_COMPLETE'`.

The cause is one list of statuses doing two jobs, "this stack has
stopped moving" and "this stack is unusable", written back when a stack
was only ever created.

The fix: a preflight, before the build, rejects the eight statuses only
a delete can clear (CREATE_FAILED, ROLLBACK_COMPLETE,
ROLLBACK_IN_PROGRESS, ROLLBACK_FAILED, UPDATE_ROLLBACK_FAILED,
UPDATE_FAILED, DELETE_FAILED, DELETE_IN_PROGRESS) with guidance matched
to the state. Five of them get `Stack '…' is in ROLLBACK_COMPLETE and
cannot be used — delete the deployment and publish it again (this mints
a new URL)`. UPDATE_ROLLBACK_FAILED is the one state where a live
dashboard, whose URL a customer may have hardcoded, can be rescued, so
its message points at the console's "Continue update rollback" first
(the stack then reads UPDATE_ROLLBACK_COMPLETE and the next republish
keeps the URL) and a delete second. The two delete states say to finish
or retry the delete; those rows already belong to the delete flow, so
that text only reaches the task log. UPDATE_FAILED also joins the
terminal list, so it throws on the first poll instead of timing out. And
publish now waits for whatever the status it finds settles at (UPDATE_*
to UPDATE_COMPLETE, UPDATE_ROLLBACK_* to UPDATE_ROLLBACK_COMPLETE,
otherwise CREATE_COMPLETE), so a stack already at rest resolves on the
first poll and one found mid-operation is waited out.
UPDATE_ROLLBACK_COMPLETE is deliberately not on the delete-only list:
that stack still has a working bucket and distribution and stays
reusable.

### A republish waits for its own update, and reports why it failed

Nothing waited on an update before, so the three problems here are
hazards the converge introduces rather than bugs an operator has seen.

Two Update clicks start two ECS tasks, and CloudFormation rejects the
loser with `Stack:arn:… is in UPDATE_IN_PROGRESS state and can not be
updated.`, which would mark that republish failed. A rollback settles at
UPDATE_ROLLBACK_COMPLETE, which nothing distinguished from success, and
the terminal-status error read StackStatusReason off the terminal
status, which CloudFormation usually leaves empty (the reason sits on
the preceding rollback status), so an operator would get a bare status
with no cause. DescribeStacks is eventually consistent, so the first
poll after an UpdateStack can still return the pre-update
UPDATE_COMPLETE, and status equality alone would resolve the wait on a
stack that had not begun converging: the task would upload files and
report `published` while CloudFormation was still working.

The fix, in three parts:

- The busy rejection is recognized, logged as `Stack '…' is busy with
another operation (UPDATE_IN_PROGRESS); waiting for it to settle, then
retrying.`, and retried, re-running this run's own UpdateStack so this
template converges rather than only the winner's. It re-reads the stack
each attempt and sleeps one poll interval first, because a read taken at
the instant of the rejection can still show the pre-operation status and
would otherwise burn the retry budget instantly. After 11 attempts it
gives up with `Stack '…' stayed busy after 11 UpdateStack attempts —
another operation may be stuck; try again shortly.` Only an
`*_IN_PROGRESS` status counts as busy, even though CloudFormation words
the wedged-status rejection identically, and DELETE_IN_PROGRESS is
excluded because waiting out a teardown can only arrive at a stack that
no longer exists.
- The converge waits specifically for UPDATE_COMPLETE, so a rollback
hits the terminal throw instead of passing as a successful publish, and
the wait now remembers the last StackStatusReason it saw that wasn't
"User Initiated" and appends it to both the terminal-status and the
timeout messages.
- The wait takes the status and LastUpdatedTime read immediately before
the UpdateStack and treats any read matching both as stale. An advanced
LastUpdatedTime is positive proof this operation landed. That read is
taken fresh on every attempt and is what a no-op converge returns, so an
Update queued behind a first publish sees the finished stack's outputs
rather than an empty pre-create snapshot.

### Failures surface before the build, with the real error

Previously, the password was only read inside the create branch, after
`npm run build:themes` and `npm run build`. A first publish with the
secret unset burned the entire bake-and-webpack build, minutes of it,
and then failed with `Missing required environment variable
'MMGIS_DASHBOARDS_PASSWORD' (lean publish flow; see sample.env)`. An
update against a stack that doesn't exist reached `Stack '…' does not
exist — publish before updating` the same way, only after the build.

Separately, the stack read treated any ValidationError as "there is no
stack here": the check was `err.name === "ValidationError" || message
includes "does not exist"`, an or rather than an and. ValidationError is
CloudFormation's name for a whole family of complaints, a rejected stack
name or a malformed request among them, so publish could go on to
CreateStack against a stack it had merely failed to read, and update
would report the stack missing while hiding the actual error.

The fix: step one renders the template and reads the stack, so a doomed
run flips the row to `failed` within seconds of the task starting. And
only the ValidationError whose message names a missing stack reads as
absence; anything else is rethrown with its own message. The delete
flow's teardown calls the same helper, so it sees the real error now
too.

### Deleting a dashboard mid-publish no longer resurrects the row

Previously, an operator who hit Delete while a publish or update task
was running saw the row go to `deleting` and the teardown start, and
then the still-running task flipped that same row back to `published`,
complete with a `cloudfront_url` and bucket for resources that were
being torn down, or to `failed`. The Deployments list showed a live
dashboard that no longer existed.

The task's two terminal writes were `deployment.update({ status:
PUBLISHED, … })` and the failed equivalent, called on the instance,
unconditionally, with no notion that anyone else might have claimed the
row.

The fix: both writes now go through `Deployments.update(…, { where: {
id, status not in [DELETING, DELETED] } })`. A row the delete flow has
claimed is left alone, and the delete flow decides its next status.

### The permissions for UpdateStack, in all three layers

Previously nothing called UpdateStack, so nothing granted it. The
task-role JSON recipe, the Terraform module's `iam.tf` and the bootstrap
`boundary.tf` each listed CreateStack, DescribeStacks and DeleteStack,
and no `cloudfront:UpdateFunction`. Had the code shipped on its own, a
republish would have died on AccessDenied, and because the permissions
boundary caps everything, granting it on the role alone would still have
failed.

The fix: all three layers gain `cloudformation:UpdateStack`,
`cloudfront:UpdateFunction`, `cloudfront:GetDistributionConfig`,
`cloudfront:GetOriginAccessControlConfig` and
`cloudfront:UpdateOriginAccessControl`. The read grants are there
because CloudFormation reads a resource's existing config before
changing it. Unused `cloudformation:DescribeStackEvents` is dropped from
all three.

### A template edit that would silently mint a new URL now fails CI

Previously nothing stopped someone adding an explicit `BucketName` to
the dashboard bucket or renaming a logical ID. Under the new UpdateStack
path, either one makes CloudFormation replace the resource: the
distribution loses its origin and the dashboard's domain changes, and
that domain is hardcoded in the customer's CloudFront forwarding rule.

The fix: a test asserts `DashboardBucket.Properties` carries no
`BucketName`, and together with the existing logical-ID test a rename
fails the suite. A second test pins
`DashboardAuthFunction.Properties.AutoPublish` to true, because without
it a converge would update the function's code while every request kept
meeting the old LIVE stage, which is to say the old password.

### The admin can no longer start two publish tasks for one dashboard

Previously, the Update button on the Deployments page was always enabled
and the update route took no lock. Two clicks, or two admins in two
tabs, started two ECS tasks against the same stack. Save & Publish in
the config editor took the same path: it picks the mission's existing
row and calls Update on it whatever state the row is in, so a second
Save & Publish during a first provision hit Update on a `provisioning`
row. The Publish form had no guard either, so a second Publish for a
mission that already had a dashboard quietly created a second stack with
a second URL.

Nothing on the server could tell a running task from a dead one. The
routes discarded the ECS task identifier that RunTask returns, and the
admin role could only start tasks, not describe them. So a task that
died without writing its row left the row at `provisioning` or
`updating` forever, and any lock keyed on that status would have locked
the row forever. That is why an earlier attempt at a server-side refusal
was reverted.

The fix: the row now remembers its task. Both routes store the task ARN
in the row's settings, and the publish task registers its own ARN at
startup from the ECS metadata endpoint, so a missing ARN means the task
never reached its first line. The admin task role gains
`ecs:DescribeTasks`, scoped to the environment's cluster, in all three
IAM layers. With that, the update route claims the row with a
compare-and-set (from a resting status to `updating`, in one statement)
and refuses the loser with a message naming the state: a publish or
update already running, a dashboard being deleted, or one already
deleted. The publish route refuses a second dashboard for a mission that
already has one; the page shows a dialog naming the existing dashboard
and lets the admin publish another on purpose. The Deployments page
disables Update on any row that is not at rest, with a tooltip per
state, and Save & Publish shows the server's reason and watches the
existing row instead of reporting a generic failure.

### A dead publish task is reported as failed, and a live one is visible

Previously, a task killed mid-run (container stopped, out of memory, a
deploy during a publish) left its row at `provisioning` or `updating`
with no error text, indefinitely. The page showed the status and nothing
else.

The fix: on every list read, a row in flight with a stored ARN is
checked against ECS. If ECS says the task stopped, or no longer lists
it, the row flips to `failed` with the stop reason and a next step
("Delete this row and publish again" for a provisioning row, "Use Update
to retry, or Delete and publish again" for an updating one). That write
is pinned to the row's status and to the ARN whose death was observed,
so a stalled ECS answer about an old task can never land on a row an
Update has since re-claimed. A task ECS does not list yet, within two
minutes of the claim, is reported as not yet confirmed rather than dead.
If ECS cannot be asked at all (the grant not yet applied, no credentials
locally), the row is left alone and says so. While a task is alive the
row shows it, with the start time and the ECS status, and says what to
do if it looks stuck.

### Delete waits for a running publish

Previously, Delete was always enabled. An admin who clicked it during a
publish saw the teardown start while the task was still uploading into
the bucket being emptied. A re-populated bucket can push the stack
delete to DELETE_FAILED, which this same PR classifies as delete-only.

The fix: the delete route asks ECS about the row's task and refuses
while it is alive, with the task's ECS status in the message; a stopped
or forgotten task never blocks a delete. The Delete button is disabled
only when ECS confirms the task is alive. When the task cannot be
confirmed, Delete stays available and both the row and the delete modal
say that deleting now may race a live publish. The publish task also
re-reads its row after registering its ARN and stops before touching AWS
if the row has been deleted in the meantime. A row already in `deleting`
keeps Delete enabled, since clicking it retries a stuck teardown.
Refusals refetch the list, so a stale tab catches up.

## Decisions to review

- **Converging the stack is a hard gate on every republish.**
UpdateStack runs before the asset copy and bundle upload, and a rollback
or failure throws, so a template or IAM regression now fails routine
content republishes on every existing dashboard. The alternative is to
converge best-effort and upload the bundle regardless.
- **A rotated `MMGIS_DASHBOARDS_PASSWORD` reaches each dashboard on its
next republish, not all at once.** Update bakes the environment's
current secret into the edge function, so after a rotation dashboards
move to the new password one republish at a time; before this PR a
rotation reached no existing dashboard at all. The alternative is a
separate "push the password to every dashboard" action, which can be a
follow-up.
- **The admin lock is keyed on ECS task liveness, and fails open.**
Update and Delete are refused only when the row's task is confirmed
alive; if ECS cannot be asked, the row is left alone, Delete stays
available, and the page says the task could not be confirmed. The
alternative is to refuse on row status alone, which is simpler and would
lock a row forever after a task dies while the grant is missing.
- **The busy-wait retry inside the converge stays as a backstop.** With
the compare-and-set claim, a second UpdateStack on the same stack can
only come from an operation started outside MMGIS or from the
millisecond window before the ARN is stored; the retry loop (up to 11
attempts, each wait bounded by the 30-minute stack timeout) covers
those. The alternative is to remove it and fail such a run outright.
- **A row becomes `failed` only on a definitive ECS answer.** A task ECS
does not list within two minutes of the claim is reported as
unconfirmed, not dead, and the page shows a neutral "Starting the task"
note for the first minute after a click. The alternative is to also time
out a row with no recorded task after some interval, which was cut
because it could mark a live task's row failed.
- **A second dashboard for a mission needs a confirmed Publish.** The
publish route refuses when the mission already has a dashboard that is
not deleted or being deleted, and the page offers "Publish another" with
the existing one named. The alternative is a hard
one-dashboard-per-mission rule.
- **A hung task holds Delete until it is stopped in ECS.** The row and
the tooltip say so. The alternative is a Stop button backed by
`ecs:StopTask`, which is a separate grant and its own decision.
- **A Publish re-run against an existing stack waits for it to settle
rather than converging it.** Publish only needs a working bucket, so it
never runs UpdateStack; only the update action converges. The
alternative is to converge on both actions.
- **`cloudformation:DescribeStackEvents` is removed from all three IAM
layers.** Nothing calls it, and failure reasons now come from the
stack-level StackStatusReason carried across polls rather than from
per-resource events. Getting it back costs a boundary apply and a role
apply.
- **Two OriginAccessControl grants are added as headroom.**
`GetOriginAccessControlConfig` and `UpdateOriginAccessControl` go into
all three layers even though the template's OAC properties are constants
today; least privilege says drop them, and adding them later is one PR
across three files.

Post-merge, in order:

1. Apply `infrastructure/terraform/bootstrap` by hand so the permissions
boundary carries `UpdateStack`, the new CloudFront grants, and the admin
role's `ecs:DescribeTasks`. Skipping this makes every republish fail
with AccessDenied, because the boundary caps the role; until the
DescribeTasks grant lands, in-flight rows show as unconfirmed and Delete
stays available.
2. Let the environment deploy workflow (`iac-deploy.yml`) apply the
module's publish and admin task-role changes.
3. Click Update once on every existing dashboard so it picks up the edge
function from #330.
4. Do one live republish in the DEV account: it is the only end-to-end
check that can't be run locally.
Base automatically changed from 316-republish-stack-convergence to development September 9, 2026 15:29
@CarsonDavis
CarsonDavis merged commit d39a92f into development Sep 9, 2026
@CarsonDavis CarsonDavis linked an issue Sep 9, 2026 that may be closed by this pull request
6 tasks
@CarsonDavis
CarsonDavis deleted the 316-dashboard-cache-control branch September 9, 2026 15:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The final dashboard should be deployable to a subpath

2 participants