What to deploy, how many of each, and which environment each one needs.
The combined Main deployment and independent backend both require the same deployment-access policy; there is no legacy authorization mode. No production cutover has been performed. Complete the review and rollout gates first.
The app is two processes — the Next server and the step worker — and the initial release is operated with exactly one step worker.
| Service | Start command | Count |
|---|---|---|
| Web — default processOS | npm start | 1 |
| Web — each skin | npm start + NEXT_PUBLIC_FRONTEND_ID | N |
| Step worker | npm run job:step | exactly 1 |
The table above describes the combined-host option. For canonical backend/file custody with independent peer frontends, use the concrete service mapping. Only backend/worker hosts connect to the canonical database in that topology.
The worker polls running processes and advances eligible closed-language continuations. The canonical store uses compare-and-set revisions and effect claims, but this release does not claim transparent crash recovery or approved multi-worker operations. Keep one worker and do not overlap old and new workers during cutover.
Whitelabel web services must not start the worker. npm start is Next only; the
worker is a separate service with npm run job:step.
Zero workers is the other failure: the app looks fine, input steps work, and every
automated step sits forever.
| Web (default) | Web (skin) | Worker | |
|---|---|---|---|
APP_BASE_URL / NEXT_PUBLIC_APP_BASE_URL | that origin | that origin | — |
| Auth0 vars | ✅ | ✅ | ✅ |
MONGO_URL | same | same | same |
NEXT_PUBLIC_FRONTEND_ID | — | ✅ | — |
| Integration credentials | — | — | ✅ |
Integration credentials belong on the worker — it executes the steps that use them.
Each web deploy's APP_BASE_URL must be its own public origin, or links built in
expressions (currentProcess.url) point at the wrong service.
Full table: reference/environment.md.
Worker initialization fails for invalid required configuration, including missing protected-store keys when signing is configured. Missing provider credentials can instead fail at dispatch; a running worker alone does not establish that a workflow can execute. Run the configuration preflight with the explicitly intended capabilities before activation.
Every web origin must be registered in the Auth0 application's Allowed Callback URLs, Allowed Logout URLs, and Allowed Web Origins. Adding a skin means adding its origin.
CORS_ALLOWED_ORIGINS is only needed when a separate SPA origin calls the API — not
for the skins, which serve their own API.
See auth0-setup.md.
MONGO_URL set → MongoDB. Unset → JSON files in .process-platform/.
Canonical storage selection is shared in composition/storage-configuration.ts; operator reads in
composition/application-stores.ts do not initialize or seed the runtime. There is no legacy runtime switch,
and file storage is not viable for a real deployment — it is local-dev only, and
multiple web services would each write their own copy.
All services must share one MONGO_URL.
npm run verify
This runs the local regression/architecture gate, including a fresh build. See the verification guide for exact coverage and limits. Also validate the actual deployment configuration, storage and identity/provider flows before rollout; the local fixture does not certify them. If merging to main auto-deploys, these checks must precede that merge rather than being deferred until after it.
Redeploy the worker as well as the web services when the change touches the closed language/runtime, registration contracts, or generic effect host. It is possible to ship a new compiled definition to the web and leave the worker on the old build, which produces a template you can author and a step that never executes.
Templates and processes live in the database, not the build, so rolling back code does not roll back data.
Two consequences:
Adding a step type is therefore effectively forward-only once templates use it.
A versioned template's new version (and its migration of every process) is published by the release step, once, before the new deployment takes traffic; see template versions. Exactly one service runs it. Settings per service (the repository has no Railway config-as-code, so these are service settings):
| Service | Start command | Health check |
|---|---|---|
| process-platform (default web) | npm run release && npm start | path /health, timeout 600 s |
| other web services built from the main application (skins) | default | path /health |
| bespoke applications (spell-review, allocation-risk) | their own | none: they serve no /health |
| step worker | default (npm run job:step) | none (no HTTP) |
Pre-deploy command: none on every service.
PROCESS_ACCESS_CONFIG, a file
on the service volume); as a pre-deploy command it fails with "Deployment access policy could not be
read as JSON" (seen in production, 2026-10-01). So the release runs at the start of the
process-platform start command: npm run release && npm start. The guarantee is the same: if the
release exits non-zero the server never starts, /health never returns 200, the deployment fails and
the old deployment keeps serving; nothing changed for the failing template (another template that
committed first stays published), and the output names every process that blocked the publication.
Keep TEMPLATE_PUBLISH_SETTLE_MS (default 60000, at most 150000: one budget for every wait in the
step) well under the health-check timeout.[process-worker] running; the release step has not published this build's versions yet. Keep exactly
one worker running across the deploy, so in-flight work can settle.Usable signals:
200 when ready (503 with the reason while it waits for the release step);
GET /login returns 200The worker has no heartbeat and no alerting. A crashed worker looks identical to a quiet system until someone notices runs are not advancing. Worth adding a supervisor that restarts it.
Deployment now consumes compiled registrations and explicit capabilities. Read deployment bindings and migration before activating a worker or moving historical data. The old storage/runtime cannot share a cutover dataset with new writers. Local development and packed proofs are not authorization for a live rollout.
Run in the intended service environment, explicitly identifying its role and required capabilities:
npm run check:deployment -- web none npm run check:deployment -- worker safe,dune,slack,notion
Choose the capability list from the intended workflows, not from the example above. The command
reports missing/invalid fields without printing values or connecting to storage/providers. It checks
Mongo/durable-file configuration, web-login fields, selected worker credentials and persistent signing
key requirements. Explicit PROCESS_APP_CONFIG provisioning needs separate validation. A passing
shape check does not certify Auth0 tenant permissions, shared-key consistency, mounted volumes,
credential validity or live provider behavior. See cutover.
On Railway, sealed variables are delivered to the service but omitted from CLI variable listings
and railway run. Run this preflight inside the intended service (or a securely provisioned equivalent),
not against a local export that silently lacks sealed values. Absence from a control-plane read is not
evidence that a running service lacks a credential. Never print the values to diagnose this.