Develop against production data

Copying the production database locally, so you develop against real templates and real runs without touching prod.

Prod is only ever read. The local copy is replaced on every pull.


Setup

1. Local MongoDB

docker compose up -d mongo

mongo:7 on :27017, data in the pp-mongo-data volume.

2. Prod connection string

Put it in .env.prod (gitignored):

PROD_MONGO_URL=mongodb+srv://user:...@cluster/...

⚠️ Use a read-only database user. mongodump needs no write access, and a write-capable credential sitting in a local file is an unnecessary risk.

scripts/pull-prod-data.sh refuses to run if this points at localhost, so you cannot restore onto prod by mistake.

3. Point the app at the copy

In .env:

MONGO_URL=mongodb://localhost:27017
SKIP_TEMPLATE_SEED=1

SKIP_TEMPLATE_SEED stops the step worker filling in missing templates, so the copied production set is left exactly as pulled. See manage templates.

4. Pull

npm run data:pull

Restart afterwards — the storage backend is chosen once at startup.


⚠️ The worker will act on copied data

This is the part that matters, and nothing in the tooling guards it.

A prod copy contains live running processes parked mid-flight on slack_notify, safe_execute, notion_create_database_item and so on. The worker cannot tell a copy from the real database. It sees a running process, a due step, and executes it.

Side effects need credentials, so the protection is simply not to have them:

.env containsRunning the worker
No integration credentialsNothing leaves the machine. slack_notify records the failure and advances; dune / notion / sheets fail the run.
Real SLACK_BOT_TOKENPosts to the real channel, mentioning real people.
Real SAFE_EXECUTOR_PRIVATE_KEY + mainnet ETH_RPC_URLBroadcasts a real transaction.

To run the worker safely

The worker exits at startup unless SAFE_EXECUTOR_PRIVATE_KEY and ETH_RPC_URL are set — so running it at all means setting them. Set them to values that cannot reach anything:

SAFE_EXECUTOR_PRIVATE_KEY=0x0000000000000000000000000000000000000000000000000000000000000001
ETH_RPC_URL=http://127.0.0.1:1

A throwaway key with no funds and an RPC that refuses connections. The worker starts, a safe_execute step attempts, the RPC call fails, and the run is failed in the copy. No transaction is ever broadcast.

Leave every other integration credential unset.

If you would rather not think about it

npm run dev

Web only. No worker, so no automated step runs at all — see installation.


What the copy does and does not give you

Real: templates as authored in prod, real runs at every stage, real context values, real audit trails. Enough to reproduce a bug, read a record, or check a UI change against data that looks like production.

Not real: users. Identity lives in Auth0, not in Mongo, so the copy has no accounts. Sign in with your Auth0 tenant as usual; your permissions are whatever that org grants, not a local stub.

Diverging immediately: the moment a worker runs against the copy, it advances processes that are still mid-flight in prod. The copy is a snapshot, not a mirror. Re-pull when you need current data.


Handling the data

Process context carries Safe addresses, transaction hashes, operator emails and Slack ids. npm run data:pull puts an unencrypted copy of that on your disk.

  • .env.prod and the Mongo volume are gitignored — keep it that way.
  • docker compose down -v removes the volume when you are done with it.
  • Do not pull prod data onto a shared or unencrypted machine.

Known gaps

  • Nothing prevents the worker from acting on a copied database; the only protection is withholding credentials.
  • The pull is all-or-nothing. There is no subsetting, no anonymisation, and no way to exclude a collection.