[316] Converge a dashboard's CloudFormation stack on republish - #348
Merged
Conversation
Terraform plan preview —
|
Terraform plan preview —
|
This was referenced Aug 27, 2026
CarsonDavis
force-pushed
the
316-dashboard-subpath
branch
from
August 27, 2026 16:15
1b2ccc3 to
9f3ad31
Compare
CarsonDavis
force-pushed
the
316-republish-stack-convergence
branch
from
August 27, 2026 16:30
60c6db9 to
6bb5bf6
Compare
A customer fronting a published dashboard with their own CloudFront forwards the full path and declares the matched prefix in X-Forwarded-Prefix. The viewer-request Function now validates that header (open-redirect and path-traversal shapes are ignored), strips the prefix from request.uri, and 302-redirects a slash-less entry to its trailing-slash form so relative asset URLs resolve. infrastructure/cloudfront-function.js becomes the deployed source: cfn-template.js reads and templates it at render time instead of carrying a second copy. For that to work every client-side URL under the dashboard must be relative to the document rather than the origin root: the upload route returns the S3 key with no leading slash, the Card adapter and the Configure preview classify that key shape instead of mission-prefixing it, the PDF worker resolves against document.baseURI, LayerGeologic drops its rooted /public, and a static build stops routing tile URLs through the server-side tile proxy. Adds the customer-facing integration contract at docs/infrastructure/serving-a-dashboard-from-your-domain.md.
CarsonDavis
force-pushed
the
316-dashboard-subpath
branch
from
August 27, 2026 18:02
9f3ad31 to
078adc9
Compare
CarsonDavis
force-pushed
the
316-republish-stack-convergence
branch
from
August 27, 2026 18:08
6bb5bf6 to
1f719bb
Compare
- docs: require query-string forwarding and https-only origin in the customer subpath guide - docs: drop the deprecated apache SERVER option from sample.env and ENVs - resolveImageUrl/buildPreviewSrc: guard non-string config values instead of throwing (blank SPA / blanked cards) - Viewer_: build the mosaic path from missionPath (folder name), matching the rest of the file, so missions whose folder != display name resolve - QueryURL: return '' for a valueless param instead of the literal "undefined" - tests: add a CloudFront-function 10KB size check; cut duplicate, impossible-input, dead-guard, and redundant tests - cfn-template: drop the redundant readFileSync try/catch (Node's ENOENT already names the path)
…e literals are runtime-supported; the real limit is no let/const plus the acorn ES5 test)
The `update` action only re-uploaded files, so a dashboard created before an
edge-function change could never receive it: the stack was created once and
never updated. Update now renders the current template, calls UpdateStack, and
polls to UPDATE_COMPLETE — treating "No updates are to be performed" as an
up-to-date no-op — and the publish action reuses that same describe so a
re-run over a stack mid-operation waits it out instead of dying on
AlreadyExistsException. The publish task role, the Terraform module, and the
permissions boundary each grant cloudformation:UpdateStack plus the
CloudFront read-back/update actions the converge needs; DescribeStackEvents,
which nothing calls, is dropped. A test pins the template's logical IDs and
the bucket's anonymity, since renaming either would make UpdateStack REPLACE
the distribution and mint a new domain.
The wait itself has to survive DescribeStacks' eventual consistency: the first
polls after an UpdateStack can still report the pre-update status, which for
the ordinary republish IS UPDATE_COMPLETE. waitForStack therefore takes the
prior { status, lastUpdatedTime } and treats a poll as stale only while it
matches BOTH — an advanced LastUpdatedTime is positive proof this update
landed, so an update that starts and finishes inside one poll interval is
recognized rather than polled to a false timeout, and a rollback back to the
resting status still fails with its reason. That reason is the last non-empty
StackStatusReason seen while polling, because CloudFormation puts it on the
in-progress rollback and leaves the terminal status empty. UPDATE_FAILED joins
the terminal statuses so a stuck stack throws instead of being polled for
thirty minutes, and CREATE_FAILED — where this code's OnFailure "DO_NOTHING"
leaves a failed first publish — joins the statuses that get the
delete-and-republish guidance.
Two republish clicks start two ECS tasks, and CloudFormation rejects the
loser's UpdateStack outright. Rather than mark that row failed after the other
task succeeded, the update path recognizes the rejection, waits the in-flight
operation out, and carries on to the upload, so a double republish stays
harmlessly last-write-wins. planStackWait() derives the wait parameters for
each of these cases, so the wiring between the two is table-tested rather than
implicit in the call sites.
Replaces the one-shot busy handler with a single bounded retry loop (convergeStackUpdate) that fixes four related concurrency bugs on the update/republish path: - seed `prior` on our own converge wait so an eventually-consistent pre-op read can't resolve the loser early (false "published") - target UPDATE_COMPLETE so a rollback throws instead of being reported as a successful publish - narrow isStackBusyError to *_IN_PROGRESS and move the wedged statuses (ROLLBACK_FAILED, UPDATE_ROLLBACK_FAILED, UPDATE_FAILED, DELETE_FAILED) into the preflight, so a stuck stack gets delete-and-republish guidance instead of being misread as busy — independent of CloudFormation wording - the loser retries its OWN UpdateStack after waiting the winner out, so this run's template actually converges Collapses planStackWait/REUSABLE_STACK_STATUSES into the loop and the publish branch, deletes the tests that duplicated settleStatusFor or pinned the old behavior, and folds in the review nits (update route comment, UpdateStack on the topology diagram).
CarsonDavis
force-pushed
the
316-republish-stack-convergence
branch
from
September 1, 2026 14:47
789a98b to
3600864
Compare
CarsonDavis
force-pushed
the
316-dashboard-subpath
branch
from
September 1, 2026 21:24
c8557fa to
5dfd36d
Compare
# Conflicts: # configure/src/core/upload.js # docs/infrastructure/serving-a-dashboard-from-your-domain.md # infrastructure/README.md # infrastructure/cloudfront-function.js # src/essence/Ancillary/QueryURL.js # src/essence/Basics/Layers_/Layers_.js # src/essence/Basics/Viewer_/PDFViewer.jsx # src/essence/Basics/Viewer_/Viewer_.js # src/essence/Tools/Card/adapters/buildCardData.ts # tests/unit/relativePublicPaths.spec.js
…republish-stack-convergence
…ribe create-or-converge
…republish-stack-convergence
…republish-stack-convergence
…republish-stack-convergence
…republish-stack-convergence
…ss rule to the timestamp
… route wedged statuses to the delete guidance
…republish-stack-convergence
…, and widen the in-flight window
…, and guard the route's failure writes
…republish-stack-convergence
…hape from stackAction
…publish tasks through one function
…, paced busy retry, preflight and template before the build, delete-safe terminal writes
…republish-stack-convergence # Conflicts: # docs/adr/deployment/lean/adr.md # tests/unit/cfnTemplate.spec.js
…infrastructure.spec.js
A republish that finds its stack in UPDATE_ROLLBACK_FAILED now points at the console's Continue update rollback, which keeps the dashboard URL, before suggesting a delete. The two delete states say to finish or retry the delete. The other five dead-end statuses keep the delete-and-republish line. The status list and message helper move into aws-provision so they can be unit-tested.
The Deployments page and its routes now know whether a dashboard's publish task is alive. Both routes store the task ARN on the row and the task registers its own at startup; the admin role gains ecs:DescribeTasks in all three IAM layers. Update claims the row with a compare-and-set and refuses a second click with a state-specific reason; Publish refuses a second dashboard for a mission unless confirmed; Delete refuses while the task is confirmed alive and stays available when it cannot be confirmed. A task ECS reports stopped or forgotten flips its row to failed with the reason and a next step, pinned to the ARN that was observed so a stale answer cannot land on a re-claimed row. Every refusal and disabled button carries visible text, and refusals refetch the list.
CarsonDavis
marked this pull request as ready for review
September 9, 2026 14:29
…republish-stack-convergence
slesaad
approved these changes
Sep 9, 2026
…republish-stack-convergence
CarsonDavis
added a commit
that referenced
this pull request
Sep 9, 2026
## Line breakdown | Category | Lines added | % | |---|---|---| | Tests | 673 | 61.7% | | Production code | 301 | 27.6% | | Docs | 75 | 6.9% | | Config/build | 41 | 3.8% | | Generated (package-lock.json) | 1 | 0.1% | | **Total** | **1091** | | ## Overview A published dashboard can now be served from a path on a customer's own domain — `site.gov/tools/dashboard/` — not just from its own CloudFront address. Part of #316 — PR 1 of 3 (with #348 republish stack convergence and #349 cache-control tiering). Serves dashboards published after this merges; existing dashboards pick up the new edge function via #348. ## What this does - **Strips the customer's prefix at the edge.** The per-dashboard CloudFront Function reads `X-Forwarded-Prefix`, validates it, and removes it from the request path. A missing or malformed header passes through untouched — a loud 403, never the wrong files — and the password gate still runs first. - **Redirects slash-less entry URLs** (`…/tools/dashboard` → `…/tools/dashboard/`) with the query string preserved, marked `no-store`. A character that would corrupt the rebuilt URL (a raw `&` or `#`) is percent-escaped so the pair survives intact. - **Single-sources the edge function.** The deployed code is read from `infrastructure/cloudfront-function.js` at publish time instead of a hand-copied string; tests pin it to ES5 and under CloudFront's 10 KB cap. - **Stores uploaded image paths slash-less** (`assets/…`, not `/assets/…`) so they resolve inside the dashboard wherever it's mounted. Both readers (Card renderer, Configure preview) rebase legacy `/assets/…` values on read, so existing configs keep working. - **Removes the other domain-root assumptions.** The PDF worker, geologic patterns, and static-build tile/viewer paths no longer emit `/…` or `../..` escapes; server and Docker modes are unchanged. - **Hardens the URL readers.** Query values containing `=` survive whole, malformed percent-encoding reads as absent instead of throwing at startup, a valueless param reads `''` instead of the string `"undefined"`, and bare `?_preview`/`?forcelanding` flags are presence-based at both entry points (LandingPage now agrees with modern.js). Pinned by a new spec. - **Theme icons resolve through the same rule as Card images.** Both `upload` fields share one resolver in `src/essence/Tools/_shared/content/uploadKey.ts`; the viewer's master image and model texture use the plain mission path instead of a `../../../../` climb. - **Ships the customer doc** — `docs/infrastructure/serving-a-dashboard-from-your-domain.md`: setup steps, a behavior table, and the why behind each setting. ## Decisions to review - **`X-Forwarded-Prefix` is validated by deny-list** (`//`, `:`, `\`, `/../`), not an allowlist. - **Old `/assets/…` values are supported forever instead of migrated once** - **Publish reads `infrastructure/cloudfront-function.js` from the ECS image at runtime** — nothing pins that directory into the image beyond today's `COPY . .`; slimming the image later breaks publish at runtime. - **`forceLanding` is accepted in both casings at every entry point**, though nothing in the app emits the camelCase form; the alternative is reverting the landing page to lowercase only.
6 tasks
CarsonDavis
added a commit
that referenced
this pull request
Sep 9, 2026
| Category | Lines added | % | |---|---|---| | Tests | 242 | 76.6% | | Production code | 67 | 21.2% | | Docs | 7 | 2.2% | Every object a publish writes now carries a Cache-Control header, so a republished dashboard reaches visitors instead of sitting stale on a cache we cannot reach. Closes #316. PR 3 of 3 for the issue, stacked on #348 (which stacks on #330), retargeted to `development` as the stack merges. ### A republish now reaches a customer's own cache Previously, an operator republished a dashboard, saw the new content on our CloudFront (the publish invalidates `/*`), and a customer fronting that same dashboard on their own domain kept serving the old `index.html`. Every object we wrote went up with no Cache-Control at all: the PUTs carried only Bucket, Key, Body, ContentLength and ContentType, and the shared-asset copies used CopyObject's default COPY directive, which carries over the source object's headers, and the upload router never set a Cache-Control on those either. With no instruction from the origin, a CloudFront cache policy falls back to its Default TTL, 86400 seconds in the managed policies and in most hand-rolled ones. The old entry page could stay pinned for a day in a distribution we have no ability to invalidate. The fix: `cacheControlForKey(key)` assigns one of three tiers to every object, applied on all four write paths (`uploadDirectory`, `uploadFile`, `copyPrefix`, `copyObjectIfExists`): - `no-cache`, revalidate on every request, for `index.html`, `build/index.html` and `Missions/<mission>/config.json`: the entry page and the baked mission config. - `public, max-age=31536000, immutable` for `build/static/{js,css,media}/…` and `assets/<mission>/<subdir>/uploads/…`: the two classes whose filenames are content-addressed, webpack content hashes on one side and `crypto.randomUUID()` names that are never overwritten on the other. - `public, max-age=300` for everything else, the keys that genuinely change in place on republish: `build/static/cesium/…`, `public/workers/…`, `Missions/…/Data/mosaic_parameters.csv`, and non-upload `assets/…`. A republish now reaches a header-honoring cache in stages: entry page and config at once, the rest within about five minutes. The two copy paths switched to `MetadataDirective: REPLACE`, because COPY cannot add a header the source never had; REPLACE drops the source's whole metadata set, so those copies restate ContentType as well, re-derived from the file extension. One test guards the premise of the immutable tier: a new spec reads `configuration/webpack.config.js` as text and asserts a content hash in all six output-name settings, so dropping one breaks the suite instead of quietly putting a stable filename into the year-long tier. ### Two un-rendered templates are no longer published Previously, anyone hitting `<dashboard>/build/index.pug` or `<dashboard>/public/index.html` got a raw template with `#{LINK_PREVIEW_TITLE}`, `#{AUTH}` and `#{user}` still in it. The publish uploads `build/` and `public/` wholesale, and it interpolates those placeholders into `build/index.html` only. The two siblings shipped exactly as they sit on disk, and `uploadDirectory` had no way to leave a key out. The fix: `uploadDirectory` takes an optional `filter(key)` and `publish-static.js` skips those two exact keys. They are left out of the upload and out of the reported file count. ### The customer doc sets a Maximum TTL and covers password rotation Previously, `docs/infrastructure/serving-a-dashboard-from-your-domain.md` asked the path owner for a custom cache policy with Authorization in the cache key, all query strings, and Minimum TTL 0. It said nothing about a maximum, so a customer who set a one-day Maximum TTL would silently cap our year-long tier down to a day. Nothing told an operator what happens to a fronting cache after a dashboard's password is rotated. The doc was written when we sent no Cache-Control at all, so there were no TTL ceilings and no revalidating tier to describe. The fix: the setup list adds "Maximum TTL 31536000 (one year) or more", and the rationale section explains the three tiers as behavior, including what a staged republish looks like from the customer's side. It also notes that AWS's managed `UseOriginCacheControlHeaders` policies get the TTLs right but omit Authorization from the cache key, so a custom policy is still required. A closing paragraph tells operators to invalidate their distribution after rotating a password: the cache is keyed on Authorization, so responses fetched under the old password keep being served to anyone still presenting it, never reaching us to be rejected, for up to a year on the immutable tier. ## Decisions to review - **One-year `immutable` on password-gated files.** `public` is the directive that lets a shared cache store a response to a request carrying Authorization, so cross-visitor safety rests entirely on the customer keeping Authorization in their cache key, which the doc requires and we cannot enforce. The alternative is a ceiling of hours or days, which still fixes the day-long pin and bounds the exposure after a rotation without a customer-side step. - **The immutable tier is decided by key prefix, not by evidence of a hash in the filename.** The webpack half of that claim is guarded by the new spec and the uploads half is only a code comment; the alternative is to require a hash segment in the key itself, which needs no cross-file guard. Being wrong is asymmetric, because we cannot purge a stale immutable object from a customer's edge for a year. - **`MetadataDirective: REPLACE` restates only ContentType and CacheControl.** Content-Encoding, Content-Disposition and `x-amz-meta-*` are dropped, which is exact for every type the upload router writes today; the alternatives are to HeadObject the source and carry its headers forward, or to set Cache-Control at upload time so a plain COPY suffices. - **Our own per-dashboard distribution stays on the managed CachingOptimized policy.** Its 1-second minimum TTL overrides `no-cache` at our edge, which does not matter because the publish already invalidates `/*`; the alternative is an owned cache policy in the template with minimum 0. - **The two raw templates are kept out by an exact-key skip on a publish that only writes.** A dashboard published before this keeps its copies until its bucket is emptied; the alternative is to delete those two keys on republish, which needs `s3:ListBucket` and `s3:DeleteObject` added in three IAM layers.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Clicking Update on a published dashboard now brings its CloudFormation stack up to the current template, so infrastructure fixes reach dashboards that already exist. The Deployments page also stops an admin from starting two publishes for one dashboard or deleting one while its publish task is still running, and says why when it refuses.
Part of #316. PR 2 of 3: stacked on #330 (serve a dashboard at a customer subpath) with #349 (cache-control) on top of this one; retargets to
developmentonce #330 merges.Republishing a dashboard now updates its AWS infrastructure, not just its files
Previously, an operator clicked Update on a published dashboard in /configure → Deployments, the row went
updating→published, the bundle was replaced, the URL was unchanged, and nothing about the dashboard's CloudFront changed. The entire update branch ofscripts/publish-static.jswas a DescribeStacks call and a "does not exist, publish before updating" throw.Each dashboard's bucket, distribution and edge Function are created once, by CreateStack at the first publish, and were never touched again. So a dashboard already sitting behind a customer path prefix kept serving the old edge Function and stayed broken, and a rotated
MMGIS_DASHBOARDS_PASSWORDnever reached it. The only escape was delete and republish, which mints a new CloudFront domain: the one thing path owners hardcode.The fix: Update renders the current template, which bakes in the current password, calls UpdateStack, and polls to UPDATE_COMPLETE. Same stack, same bucket, same URL. It logs
Converging stack 'mmgis-dashboard-N' to the current template — this re-bakes the current dashboards password into the auth Function.and thenStack '…' reached UPDATE_COMPLETE.When the template already matches, CloudFormation's "No updates are to be performed" is caught and logged asStack '…' is already up to date.rather than failing the republish.Publishing over an existing stack either fails fast with guidance or waits for it to settle
Previously, a publish over a stack that already existed could hang for the full 30-minute timeout, or fail on the stack's own success status. A stack left at ROLLBACK_COMPLETE by a failed create ended, after the whole build, at
Stack 'mmgis-dashboard-N' reached terminal status 'ROLLBACK_COMPLETE': accurate, but it doesn't say what to do, and CloudFormation cannot update such a stack at all. DELETE_IN_PROGRESS and UPDATE_FAILED were not in the terminal list, so the poller sat on a deleting stack until it vanished (Stack '…' does not exist (deleted or never created)) and on a failed update until the 30-minute timeout. And the wait always ran against exactly CREATE_COMPLETE, so a stack resting at UPDATE_COMPLETE, which the converge path now produces routinely, killed a publish re-run withStack '…' reached terminal status 'UPDATE_COMPLETE'.The cause is one list of statuses doing two jobs, "this stack has stopped moving" and "this stack is unusable", written back when a stack was only ever created.
The fix: a preflight, before the build, rejects the eight statuses only a delete can clear (CREATE_FAILED, ROLLBACK_COMPLETE, ROLLBACK_IN_PROGRESS, ROLLBACK_FAILED, UPDATE_ROLLBACK_FAILED, UPDATE_FAILED, DELETE_FAILED, DELETE_IN_PROGRESS) with guidance matched to the state. Five of them get
Stack '…' is in ROLLBACK_COMPLETE and cannot be used — delete the deployment and publish it again (this mints a new URL). UPDATE_ROLLBACK_FAILED is the one state where a live dashboard, whose URL a customer may have hardcoded, can be rescued, so its message points at the console's "Continue update rollback" first (the stack then reads UPDATE_ROLLBACK_COMPLETE and the next republish keeps the URL) and a delete second. The two delete states say to finish or retry the delete; those rows already belong to the delete flow, so that text only reaches the task log. UPDATE_FAILED also joins the terminal list, so it throws on the first poll instead of timing out. And publish now waits for whatever the status it finds settles at (UPDATE_* to UPDATE_COMPLETE, UPDATE_ROLLBACK_* to UPDATE_ROLLBACK_COMPLETE, otherwise CREATE_COMPLETE), so a stack already at rest resolves on the first poll and one found mid-operation is waited out. UPDATE_ROLLBACK_COMPLETE is deliberately not on the delete-only list: that stack still has a working bucket and distribution and stays reusable.A republish waits for its own update, and reports why it failed
Nothing waited on an update before, so the three problems here are hazards the converge introduces rather than bugs an operator has seen.
Two Update clicks start two ECS tasks, and CloudFormation rejects the loser with
Stack:arn:… is in UPDATE_IN_PROGRESS state and can not be updated., which would mark that republish failed. A rollback settles at UPDATE_ROLLBACK_COMPLETE, which nothing distinguished from success, and the terminal-status error read StackStatusReason off the terminal status, which CloudFormation usually leaves empty (the reason sits on the preceding rollback status), so an operator would get a bare status with no cause. DescribeStacks is eventually consistent, so the first poll after an UpdateStack can still return the pre-update UPDATE_COMPLETE, and status equality alone would resolve the wait on a stack that had not begun converging: the task would upload files and reportpublishedwhile CloudFormation was still working.The fix, in three parts:
Stack '…' is busy with another operation (UPDATE_IN_PROGRESS); waiting for it to settle, then retrying., and retried, re-running this run's own UpdateStack so this template converges rather than only the winner's. It re-reads the stack each attempt and sleeps one poll interval first, because a read taken at the instant of the rejection can still show the pre-operation status and would otherwise burn the retry budget instantly. After 11 attempts it gives up withStack '…' stayed busy after 11 UpdateStack attempts — another operation may be stuck; try again shortly.Only an*_IN_PROGRESSstatus counts as busy, even though CloudFormation words the wedged-status rejection identically, and DELETE_IN_PROGRESS is excluded because waiting out a teardown can only arrive at a stack that no longer exists.Failures surface before the build, with the real error
Previously, the password was only read inside the create branch, after
npm run build:themesandnpm run build. A first publish with the secret unset burned the entire bake-and-webpack build, minutes of it, and then failed withMissing required environment variable 'MMGIS_DASHBOARDS_PASSWORD' (lean publish flow; see sample.env). An update against a stack that doesn't exist reachedStack '…' does not exist — publish before updatingthe same way, only after the build.Separately, the stack read treated any ValidationError as "there is no stack here": the check was
err.name === "ValidationError" || message includes "does not exist", an or rather than an and. ValidationError is CloudFormation's name for a whole family of complaints, a rejected stack name or a malformed request among them, so publish could go on to CreateStack against a stack it had merely failed to read, and update would report the stack missing while hiding the actual error.The fix: step one renders the template and reads the stack, so a doomed run flips the row to
failedwithin seconds of the task starting. And only the ValidationError whose message names a missing stack reads as absence; anything else is rethrown with its own message. The delete flow's teardown calls the same helper, so it sees the real error now too.Deleting a dashboard mid-publish no longer resurrects the row
Previously, an operator who hit Delete while a publish or update task was running saw the row go to
deletingand the teardown start, and then the still-running task flipped that same row back topublished, complete with acloudfront_urland bucket for resources that were being torn down, or tofailed. The Deployments list showed a live dashboard that no longer existed.The task's two terminal writes were
deployment.update({ status: PUBLISHED, … })and the failed equivalent, called on the instance, unconditionally, with no notion that anyone else might have claimed the row.The fix: both writes now go through
Deployments.update(…, { where: { id, status not in [DELETING, DELETED] } }). A row the delete flow has claimed is left alone, and the delete flow decides its next status.The permissions for UpdateStack, in all three layers
Previously nothing called UpdateStack, so nothing granted it. The task-role JSON recipe, the Terraform module's
iam.tfand the bootstrapboundary.tfeach listed CreateStack, DescribeStacks and DeleteStack, and nocloudfront:UpdateFunction. Had the code shipped on its own, a republish would have died on AccessDenied, and because the permissions boundary caps everything, granting it on the role alone would still have failed.The fix: all three layers gain
cloudformation:UpdateStack,cloudfront:UpdateFunction,cloudfront:GetDistributionConfig,cloudfront:GetOriginAccessControlConfigandcloudfront:UpdateOriginAccessControl. The read grants are there because CloudFormation reads a resource's existing config before changing it. Unusedcloudformation:DescribeStackEventsis dropped from all three.A template edit that would silently mint a new URL now fails CI
Previously nothing stopped someone adding an explicit
BucketNameto the dashboard bucket or renaming a logical ID. Under the new UpdateStack path, either one makes CloudFormation replace the resource: the distribution loses its origin and the dashboard's domain changes, and that domain is hardcoded in the customer's CloudFront forwarding rule.The fix: a test asserts
DashboardBucket.Propertiescarries noBucketName, and together with the existing logical-ID test a rename fails the suite. A second test pinsDashboardAuthFunction.Properties.AutoPublishto true, because without it a converge would update the function's code while every request kept meeting the old LIVE stage, which is to say the old password.The admin can no longer start two publish tasks for one dashboard
Previously, the Update button on the Deployments page was always enabled and the update route took no lock. Two clicks, or two admins in two tabs, started two ECS tasks against the same stack. Save & Publish in the config editor took the same path: it picks the mission's existing row and calls Update on it whatever state the row is in, so a second Save & Publish during a first provision hit Update on a
provisioningrow. The Publish form had no guard either, so a second Publish for a mission that already had a dashboard quietly created a second stack with a second URL.Nothing on the server could tell a running task from a dead one. The routes discarded the ECS task identifier that RunTask returns, and the admin role could only start tasks, not describe them. So a task that died without writing its row left the row at
provisioningorupdatingforever, and any lock keyed on that status would have locked the row forever. That is why an earlier attempt at a server-side refusal was reverted.The fix: the row now remembers its task. Both routes store the task ARN in the row's settings, and the publish task registers its own ARN at startup from the ECS metadata endpoint, so a missing ARN means the task never reached its first line. The admin task role gains
ecs:DescribeTasks, scoped to the environment's cluster, in all three IAM layers. With that, the update route claims the row with a compare-and-set (from a resting status toupdating, in one statement) and refuses the loser with a message naming the state: a publish or update already running, a dashboard being deleted, or one already deleted. The publish route refuses a second dashboard for a mission that already has one; the page shows a dialog naming the existing dashboard and lets the admin publish another on purpose. The Deployments page disables Update on any row that is not at rest, with a tooltip per state, and Save & Publish shows the server's reason and watches the existing row instead of reporting a generic failure.A dead publish task is reported as failed, and a live one is visible
Previously, a task killed mid-run (container stopped, out of memory, a deploy during a publish) left its row at
provisioningorupdatingwith no error text, indefinitely. The page showed the status and nothing else.The fix: on every list read, a row in flight with a stored ARN is checked against ECS. If ECS says the task stopped, or no longer lists it, the row flips to
failedwith the stop reason and a next step ("Delete this row and publish again" for a provisioning row, "Use Update to retry, or Delete and publish again" for an updating one). That write is pinned to the row's status and to the ARN whose death was observed, so a stalled ECS answer about an old task can never land on a row an Update has since re-claimed. A task ECS does not list yet, within two minutes of the claim, is reported as not yet confirmed rather than dead. If ECS cannot be asked at all (the grant not yet applied, no credentials locally), the row is left alone and says so. While a task is alive the row shows it, with the start time and the ECS status, and says what to do if it looks stuck.Delete waits for a running publish
Previously, Delete was always enabled. An admin who clicked it during a publish saw the teardown start while the task was still uploading into the bucket being emptied. A re-populated bucket can push the stack delete to DELETE_FAILED, which this same PR classifies as delete-only.
The fix: the delete route asks ECS about the row's task and refuses while it is alive, with the task's ECS status in the message; a stopped or forgotten task never blocks a delete. The Delete button is disabled only when ECS confirms the task is alive. When the task cannot be confirmed, Delete stays available and both the row and the delete modal say that deleting now may race a live publish. The publish task also re-reads its row after registering its ARN and stops before touching AWS if the row has been deleted in the meantime. A row already in
deletingkeeps Delete enabled, since clicking it retries a stuck teardown. Refusals refetch the list, so a stale tab catches up.Decisions to review
MMGIS_DASHBOARDS_PASSWORDreaches each dashboard on its next republish, not all at once. Update bakes the environment's current secret into the edge function, so after a rotation dashboards move to the new password one republish at a time; before this PR a rotation reached no existing dashboard at all. The alternative is a separate "push the password to every dashboard" action, which can be a follow-up.failedonly on a definitive ECS answer. A task ECS does not list within two minutes of the claim is reported as unconfirmed, not dead, and the page shows a neutral "Starting the task" note for the first minute after a click. The alternative is to also time out a row with no recorded task after some interval, which was cut because it could mark a live task's row failed.ecs:StopTask, which is a separate grant and its own decision.cloudformation:DescribeStackEventsis removed from all three IAM layers. Nothing calls it, and failure reasons now come from the stack-level StackStatusReason carried across polls rather than from per-resource events. Getting it back costs a boundary apply and a role apply.GetOriginAccessControlConfigandUpdateOriginAccessControlgo into all three layers even though the template's OAC properties are constants today; least privilege says drop them, and adding them later is one PR across three files.Post-merge, in order:
infrastructure/terraform/bootstrapby hand so the permissions boundary carriesUpdateStack, the new CloudFront grants, and the admin role'secs:DescribeTasks. Skipping this makes every republish fail with AccessDenied, because the boundary caps the role; until the DescribeTasks grant lands, in-flight rows show as unconfirmed and Delete stays available.iac-deploy.yml) apply the module's publish and admin task-role changes.