Plan: staging + prod environments with a promote-to-prod flow
Stand up a staging environment alongside prod so all development happens on staging first and ships to prod only when validated. Decisions taken (2026-06-14): staging runs on the same Hetzner box as prod, and staging data is reseeded from a prod dump (staging does not run its own crons).
This doc is both the design and the runbook.
0. Current state (grounded, not assumed)
- One Hetzner box (
dexhq.io): 4 vCPU, 7.6 GiB RAM (~6 free), 150 GB disk (~90 free). - nginx + Certbot.
creddit.xyz→127.0.0.1:3001(theonchain-creditNext.js app under PM2,next start). Box also runsdexhq(:3000),creddit-indexer,rindexer. - One Postgres cluster (
127.0.0.1:5432): dbcreddit(~841 MB), ownerpostgres, app login roleonchain_credit. Holds PII: the006-newsletter-subscriberstable has subscriber emails + user agents. - Deploy:
.github/workflows/deploy.ymlon push tomain→ SSH →git reset --hard origin/main+npm ci+npm run build+pm2 restart(rolls back the working tree on build failure). Code only; migrations/refreshers manual. - Cron (root): 6 refresher jobs (
refresh-assets6h,refresh-vault-riskweekly,refresh-sofrweekdays,refresh-vault-capacity/-collateral-exposure/-lending-positionsevery 6h, offset). - Two gaps this plan closes: there are no DB backups, and migrations are hand-applied numbered SQL (
001–034) with no ledger.
1. Target topology (same box, fully isolated staging)
| Layer | Prod (exists) | Staging (new) |
|---|---|---|
| Domain | creddit.xyz | staging.creddit.xyz (basic-auth + noindex) |
| Directory | /opt/onchain-credit | /opt/onchain-credit-staging |
| Git branch | main | staging |
| Port / PM2 | :3001 / onchain-credit | :3002 / onchain-credit-staging |
| nginx vhost | sites-enabled/onchain-credit | sites-enabled/onchain-credit-staging |
| Database | creddit / onchain_credit | creddit_staging / onchain_credit_staging |
.env.local | prod | staging (own DATABASE_URL, PORT=3002) |
| Cron | active | none (reseed + on-demand refresher runs only) |
| TLS | Certbot (LE) | Certbot (LE), same --nginx flow |
| Deploy key | DEPLOY_SSH_KEY | separate DEPLOY_SSH_KEY_STAGING (sec. 5.3) |
Isolation boundaries that matter are present: separate database + role (a bad migration/write cannot reach prod), separate PM2 process + port, and separate directory + env.
Honest scope of "isolated" (same box, same Postgres cluster). Same-box staging shares the OS/kernel and the Postgres process, so a runaway staging query, an OOM, or a disk-fill can degrade prod. Same-box is the right call here, but only with the concrete guards in sec. 9 (role-level timeouts + connection cap, a PM2 memory ceiling, and disk alerting). A dedicated VM is the upgrade path the day infra-level testing (OS/nginx/PG-version) is needed.
2. Branching + promotion flow
Two long-lived branches, each bound to one environment:
feature/* ──PR──▶ staging ──auto-deploy──▶ staging.creddit.xyz (develop + validate)
│
└──"Release" PR──▶ main ──auto-deploy──▶ creddit.xyz (promote)- All feature PRs target
staging. Merge →deploy-staging.ymlrebuilds/opt/onchain-credit-stagingon:3002. - Validate on
staging.creddit.xyz. - Promote with a single
staging → mainPR. Its diff is the release. Merge → existingdeploy.ymlships prod. - Direct commits to
mainstop (branch protection: PR-only). After each release PR,mainandstagingare in sync.
Tradeoff stated explicitly (environment-branch promotion). A long-lived staging branch + "release PR" is the simplest model to operate, but it has two known costs: (1) the two branches drift and can conflict (this very PR's first revision was already dirty from a stale base — the failure mode is real), and (2) each box rebuilds from source, so staging and prod produce different build artifacts — you validate the code, not the exact shipped bytes. The modern alternative is build-once, promote-the-artifact (build a Docker image / .next bundle in CI, deploy the same artifact to staging then prod). Deferred on purpose: the current SSH+git reset+npm run build deploy is simple and working, and artifact promotion is a larger change. Revisit if drift or build-skew bites.
AGENTS.md / contribution rules update: "feature branch → PR into staging → validate → staging → main release PR → prod."
3. Database
3.1 Create the staging database + role (one-time)
-- as postgres
CREATE ROLE onchain_credit_staging LOGIN PASSWORD '<staging-pw>'
CONNECTION LIMIT 20; -- cap (sec. 9)
ALTER ROLE onchain_credit_staging SET statement_timeout = '30s'; -- guard (sec. 9)
ALTER ROLE onchain_credit_staging SET idle_in_transaction_session_timeout = '60s';
CREATE DATABASE creddit_staging OWNER postgres;
-- schema + grants are established by the first reseed (it carries them).A separate role means a leaked staging credential cannot touch creddit, and the role-level limits cap staging's blast radius on the shared cluster.
3.2 Reseed from prod — scripts/ops/reseed-staging.sh (new), PII-scrubbed
# 1. dump prod (this dump is ALSO the backup, sec. 6) -> encrypted off-box
pg_dump -Fc creddit > /opt/backups/creddit-$(date -u +%FT%H%MZ).dump
# 2. restore into a fresh staging db
dropdb --if-exists creddit_staging && createdb -O postgres creddit_staging
pg_restore -d creddit_staging --no-owner --role=onchain_credit_staging <latest dump>
# 3. SCRUB PII before anyone can reach staging (sec. below)
psql -d creddit_staging -f scripts/ops/scrub-staging-pii.sql
# 4. re-grant to the staging role; ANALYZE.PII scrub (required). The prod dump carries real subscriber emails (newsletter_subscribers) and user agents. Copying them into a basic-auth-only staging DB is a privacy/GDPR exposure. scrub-staging-pii.sql masks them, e.g. UPDATE onchain_credit.newsletter_subscribers SET email = 'sub'||id||'@staging.invalid', user_agent = NULL; (extend the script as any future PII-bearing table lands; a test asserts the known PII columns are covered). Everything else in the schema is public on-chain data and is fine to copy verbatim.
Migration-vs-reseed ordering. A reseed brings staging to prod's current schema, so it reverts any not-yet-promoted migration under test. Either reseed before a migration's testing window, or re-run migrate.sh (sec. 4) after a reseed to re-apply staging-only pending migrations. The nightly reseed is pause-able (a flag file) so it never clobbers an active test.
3.3 Why same cluster, not a second Postgres
A separate database + role is the isolation unit. A second cluster only buys PG-version/config-experiment isolation, not a current need. Same cluster keeps pg_dump/pg_restore, backups, and monitoring trivial.
4. Migration ledger (replaces hand-applied SQL)
A deterministic, idempotent runner so "apply pending migrations" is one command per environment, and the promotion checklist can't miss or double-apply.
onchain_credit.schema_migrations(filename text primary key, applied_at timestamptz default now()).scripts/ops/migrate.sh <DATABASE_URL>: listsscripts/sql/*.sqlin order, skips any already inschema_migrations, applies the rest, records them. Forward-only. Each file runs inside a transaction.- Concurrency + failure semantics (explicit): the runner takes a
pg_advisory_lockso two deploys can't migrate at once, sets alock_timeoutandstatement_timeoutso a blocked migration fails fast instead of wedging the cluster, and exits non-zero on any failure (the deploy job then aborts beforepm2 restart). Note the code-rollback path indeploy.ymldoes NOT roll back the database — migrations are forward-only and must be backward-compatible with the previous code (expand/contract), so a rolled-back build still runs against the migrated schema. - Retrofit: seed
schema_migrationswith every existing0xx-*.sqlmarked applied on BOTHcredditandcreddit_staging(already applied), so the runner starts from "nothing pending." - Additive vs destructive: additive migrations (
ADD COLUMN IF NOT EXISTS, new tables — the house style) are safe to auto-apply. Destructive ones (e.g. the032stress-column DROP) stay a gated manual step run after the deploy that stops referencing the object; the runner refuses a file tagged-- DESTRUCTIVEwithout--allow-destructive.
5. CI/CD
5.1 deploy-staging.yml (new) — clone of deploy.yml, parameterized
- Trigger:
on: push: branches: [staging]. - Same pinned-host-key SSH pattern, its own
DEPLOY_SSH_KEY_STAGING(sec. 5.3). - Server step targets
/opt/onchain-credit-staging, fetchesorigin/staging, builds, then runsmigrate.sh "$STAGING_DATABASE_URL"(additive auto-apply), thenpm2 restart onchain-credit-staging. Same build-failure rollback. - Concurrency group
deploy-staging(independent ofdeploy-production). - Prod's
deploy.ymlgains the samemigrate.shline later, once trusted.
5.2 Post-deploy health gate (new, both envs)
Deploy success today = "pm2 restarted," with manual validation. Add an automated gate: after restart, the job curls a health endpoint (a lightweight /api/health returning a 200 + a DB-reachable check) and fails the deploy if it is not 200 within a timeout. On staging, optionally a couple of smoke routes (/money-market-rates, /carries) returning 200. This turns a broken deploy into a red CI run instead of a silently-down site.
5.3 Per-environment deploy keys + branch protection
- Separate SSH deploy keys per environment from day one (cheap; a staging-deploy compromise must not equal prod-box access). Two keypairs, two
authorized_keysentries (optionallycommand=/from=restricted), two repo secrets:DEPLOY_SSH_KEY(prod),DEPLOY_SSH_KEY_STAGING. main: PR-only, no direct pushes; require deploy-staging green + the health gate on the head commit before a release PR merges.staging: PR-only; requiretsc+ tests green.
6. Backups (prerequisite + standalone win)
Today there are none. Establish before staging (the reseed dump is the backup):
- Nightly
pg_dump -Fc creddit→/opt/backups, then encrypted (age/gpg) and pushed off-box (Hetzner Storage Box / S3-compatible viarclone) with a checksum recorded, ~14-day retention. - Monthly restore-test: the reseed into
creddit_stagingcontinuously proves the dump restores; document the restore runbook. - RPO/PITR tradeoff (stated): nightly
pg_dump= up to 24h RPO, no point-in-time recovery. Acceptable here because the dataset is re-derivable from chain (the refreshers rebuild it) and the only non-derivable data is the newsletter table — so the deliberate decision is: nightly logical dump now; consider WAL-archiving/PITR only if non-derivable data grows. (DEX_HQ'sdexhqDB should get the same nightly dump, noted, out of scope.)
7. nginx + DNS + TLS + access control
- DNS:
staging.creddit.xyz→ the box (same provider ascreddit.xyz). - New vhost
sites-enabled/onchain-credit-staging:server_name staging.creddit.xyz;→proxy_pass http://127.0.0.1:3002;, 80→443 redirect. certbot --nginx -d staging.creddit.xyz(same flow that issued the prod cert).- Non-public + non-indexed (we have prior reputation/blocklist sensitivity):
auth_basic "staging"; auth_basic_user_file /etc/nginx/.htpasswd-staging;(or Cloudflare Access if/when DNS moves to CF).add_header X-Robots-Tag "noindex, nofollow" always;+ a stagingrobots.txtdisallow-all.
8. App instance on the box
git clonethe repo to/opt/onchain-credit-staging, checkoutstaging./opt/onchain-credit-staging/.env.local:DATABASE_URL→creddit_staging(roleonchain_credit_staging),PORT=3002,NODE_ENV=production. RPC/0x keys: start shared (reads are idempotent; cheapest), move to separate keys later if independent rate-limit/observability is wanted. Document that they're shared.npm ci && npm run build, thenpm2 start npm --name onchain-credit-staging -- startwithPORT=3002and a memory ceiling (--max-memory-restart 700M, sec. 9),pm2 save.- Resource check: staging app ~60 MB +
creddit_staging~0.85 GB copy + occasional refresher runs — trivially within ~6 GiB free / 90 GB disk.
9. Same-box safety guards (the price of not using a separate VM)
Concrete, low-effort guards so staging cannot take prod down:
- Postgres role limits (sec. 3.1):
statement_timeout=30s,idle_in_transaction_session_timeout=60s,CONNECTION LIMIT 20ononchain_credit_staging— a runaway staging query is killed, not left to saturate the shared cluster. - PM2 memory ceiling:
--max-memory-restart 700Mon the staging process so a leak restarts staging instead of OOM-killing prod. - Disk alerting: a simple cron that alerts (or the existing monitoring) when
/crosses ~80%, since staging's DB copy + nightly dumps consume disk. Retention on/opt/backups(14 days) bounds growth. - Postgres
shared_buffers/work_memstay prod-sized; staging's connection cap keeps it from inflating memory pressure.
10. Refreshers / cron on staging
- No cron on staging. Data comes from the reseed (sec. 3.2).
- Validate a refresher change by running it once on staging by hand:
cd /opt/onchain-credit-staging && scripts/run-cron.sh <refresher>.ts(itsrun-cron.shsources the staging.env.local, so it writescreddit_staging).
11. The indexer
The nightly reseed carries creddit-indexer's tables from prod, so staging needs no separate indexer to render. A staging indexer is only required when testing indexer changes themselves (a second rindexer pointed at creddit_staging) — defer to that work.
12. Docs (Cloudflare Pages)
Already has branch previews. Point the staging branch's preview at staging-docs.creddit.xyz (or use the auto preview URL). No new infra.
13. Rollout order (each step is additive; prod untouched until the cutover)
- Backups: nightly encrypted+checksummed
pg_dumpoff-box + restore runbook. (standalone win) - Staging DB + role (with role limits);
reseed-staging.sh+scrub-staging-pii.sql; first reseed. - Staging app: dir,
.env.local, PM2:3002with memory ceiling, build. - nginx vhost + cert + basic-auth + noindex for
staging.creddit.xyz; smoke-test. - Separate staging deploy key;
stagingbranch;deploy-staging.yml(withmigrate.sh+ health gate); push to confirm. - Migration ledger +
migrate.sh; retrofit existing0xxas applied on both DBs. - Same-box guards (sec. 9) + disk alerting.
- Cutover: branch protection on
main(PR-only); switch the contribution flow to targetstaging; updateAGENTS.md. (only now does the dev flow change) - (optional) nightly reseed cron + pause flag; separate staging RPC keys;
/api/healthroute if not already present.
Steps 1–7 are invisible to prod and to the current main-based flow; step 8 is the single switch that turns on "develop on staging, promote to prod."
14. Rollback / safety notes
- Staging is disposable: any breakage is fixed by a reseed.
- The promotion PR is the audit trail of every prod release.
- Migrations are forward-only and must be backward-compatible (expand/contract); the deploy's code-rollback does NOT roll back the DB (sec. 4).
- Destructive migrations keep the gated manual discipline, validated on staging first.
Open items to confirm before/while implementing
- DNS provider for
staging.creddit.xyz(same ascreddit.xyz); basic-auth (works today) vs Cloudflare Access (if DNS moves to CF). - Off-box backup target (Hetzner Storage Box vs S3-compatible) + encryption key custody.
- Auto-run additive
migrate.shinside deploy from day one, or keep migrations in the release checklist until the ledger is trusted. - Separate vs shared RPC/0x keys for staging (recommend shared to start).
- Whether to adopt build-once/artifact-promotion now or accept env-branch rebuild drift for the first iteration (recommend accept-for-now; revisit if it bites).