Deployment

What to deploy, how many of each, and which environment each one needs.

The combined Main deployment and independent backend both require the same deployment-access policy; there is no legacy authorization mode. No production cutover has been performed. Complete the review and rollout gates first.

The app is two processes — the Next server and the step worker — and the initial release is operated with exactly one step worker.


Services

ServiceStart commandCount
Web — default processOSnpm start1
Web — each skinnpm start + NEXT_PUBLIC_FRONTEND_IDN
Step workernpm run job:stepexactly 1

The table above describes the combined-host option. For canonical backend/file custody with independent peer frontends, use the concrete service mapping. Only backend/worker hosts connect to the canonical database in that topology.

⚠️ Exactly one worker

The worker polls running processes and advances eligible closed-language continuations. The canonical store uses compare-and-set revisions and effect claims, but this release does not claim transparent crash recovery or approved multi-worker operations. Keep one worker and do not overlap old and new workers during cutover.

Whitelabel web services must not start the worker. npm start is Next only; the worker is a separate service with npm run job:step.

Zero workers is the other failure: the app looks fine, input steps work, and every automated step sits forever.


Environment per service

Web (default)Web (skin)Worker
APP_BASE_URL / NEXT_PUBLIC_APP_BASE_URLthat originthat origin—
Auth0 vars✅✅✅
MONGO_URLsamesamesame
NEXT_PUBLIC_FRONTEND_ID—✅—
Integration credentials——✅

Integration credentials belong on the worker — it executes the steps that use them.

Each web deploy's APP_BASE_URL must be its own public origin, or links built in expressions (currentProcess.url) point at the wrong service.

Full table: reference/environment.md.

Worker initialization fails for invalid required configuration, including missing protected-store keys when signing is configured. Missing provider credentials can instead fail at dispatch; a running worker alone does not establish that a workflow can execute. Run the configuration preflight with the explicitly intended capabilities before activation.


Auth0 per origin

Every web origin must be registered in the Auth0 application's Allowed Callback URLs, Allowed Logout URLs, and Allowed Web Origins. Adding a skin means adding its origin.

CORS_ALLOWED_ORIGINS is only needed when a separate SPA origin calls the API — not for the skins, which serve their own API.

See auth0-setup.md.


Storage

MONGO_URL set → MongoDB. Unset → JSON files in .process-platform/.

Canonical storage selection is shared in composition/storage-configuration.ts; operator reads in composition/application-stores.ts do not initialize or seed the runtime. There is no legacy runtime switch, and file storage is not viable for a real deployment — it is local-dev only, and multiple web services would each write their own copy.

All services must share one MONGO_URL.


Deploying a change

npm run verify

This runs the local regression/architecture gate, including a fresh build. See the verification guide for exact coverage and limits. Also validate the actual deployment configuration, storage and identity/provider flows before rollout; the local fixture does not certify them. If merging to main auto-deploys, these checks must precede that merge rather than being deferred until after it.

Redeploy the worker as well as the web services when the change touches the closed language/runtime, registration contracts, or generic effect host. It is possible to ship a new compiled definition to the web and leave the worker on the old build, which produces a template you can author and a step that never executes.


Rollback

Templates and processes live in the database, not the build, so rolling back code does not roll back data.

Two consequences:

  • A process started under the new build keeps its template snapshot. Rolling back the code does not change runs in flight.
  • A template saved with a new step type will still contain that step after a rollback. The old build does not know the type and will not execute it — the run parks.

Adding a step type is therefore effectively forward-only once templates use it.


Release step and readiness

A versioned template's new version (and its migration of every process) is published by the release step, once, before the new deployment takes traffic; see template versions. Exactly one service runs it. Settings per service (the repository has no Railway config-as-code, so these are service settings):

ServiceStart commandHealth check
process-platform (default web)npm run release && npm startpath /health, timeout 600 s
other web services built from the main application (skins)defaultpath /health
bespoke applications (spell-review, allocation-risk)their ownnone: they serve no /health
step workerdefault (npm run job:step)none (no HTTP)

Pre-deploy command: none on every service.

  • Release in the start command, not pre-deploy. Railway does not mount volumes in the pre-deploy container, and the release step reads the deployment access policy (PROCESS_ACCESS_CONFIG, a file on the service volume); as a pre-deploy command it fails with "Deployment access policy could not be read as JSON" (seen in production, 2026-10-01). So the release runs at the start of the process-platform start command: npm run release && npm start. The guarantee is the same: if the release exits non-zero the server never starts, /health never returns 200, the deployment fails and the old deployment keeps serving; nothing changed for the failing template (another template that committed first stays published), and the output names every process that blocked the publication. Keep TEMPLATE_PUBLISH_SETTLE_MS (default 60000, at most 150000: one budget for every wait in the step) well under the health-check timeout.
  • Changing these settings needs a fresh deploy. A Railway "redeploy" of an existing deployment reuses that deployment's saved settings; after changing the start command or health check, deploy the current commit (e.g. push, or deploy the latest commit from the dashboard) for them to apply.
  • Health check. GET /health returns 200 once the service is initialized and the platform has the template versions this build serves, and 503 with the reason until then. A web service deployed before the release step has published its versions stays unhealthy and takes no traffic; set its health-check timeout longer than a release step can take (for example 600 s), or redeploy it after the process-platform deploy succeeds.
  • Worker. The worker is not gated on template versions: it runs stored programs, and in-flight work must keep moving to settle during a release step. Until the versions are published it logs [process-worker] running; the release step has not published this build's versions yet. Keep exactly one worker running across the deploy, so in-flight work can settle.
  • Prerequisite. With MongoDB the publication is one multi-document transaction, so MongoDB must be a replica set: MongoDB as a single-node replica set.
  • Code rollback. Versions only move forward: a rolled-back build whose template version is lower than the published one publishes nothing and keeps serving existing processes. Undo a form change with a new version that reverses it.

Health checks

Usable signals:

  • Web: GET /health returns 200 when ready (503 with the reason while it waits for the release step); GET /login returns 200
  • Worker: startup logs the env presence check, then ticks silently (or logs that the release step has not published its versions yet)

The worker has no heartbeat and no alerting. A crashed worker looks identical to a quiet system until someone notices runs are not advancing. Worth adding a supervisor that restarts it.

Closed-runtime activation

Deployment now consumes compiled registrations and explicit capabilities. Read deployment bindings and migration before activating a worker or moving historical data. The old storage/runtime cannot share a cutover dataset with new writers. Local development and packed proofs are not authorization for a live rollout.

Non-mutating configuration preflight

Run in the intended service environment, explicitly identifying its role and required capabilities:

npm run check:deployment -- web none
npm run check:deployment -- worker safe,dune,slack,notion

Choose the capability list from the intended workflows, not from the example above. The command reports missing/invalid fields without printing values or connecting to storage/providers. It checks Mongo/durable-file configuration, web-login fields, selected worker credentials and persistent signing key requirements. Explicit PROCESS_APP_CONFIG provisioning needs separate validation. A passing shape check does not certify Auth0 tenant permissions, shared-key consistency, mounted volumes, credential validity or live provider behavior. See cutover.

On Railway, sealed variables are delivered to the service but omitted from CLI variable listings and railway run. Run this preflight inside the intended service (or a securely provisioned equivalent), not against a local export that silently lacks sealed values. Absence from a control-plane read is not evidence that a running service lacks a credential. Never print the values to diagnose this.