Skip to content

built Built. This is a decision record, not documentation.

What is still current: The design is how staging runs: same box as prod, reseeded from a prod dump, no refresher crons of its own. The runbook is Deployment, not this page, and the current-state section describes the pre-staging world (there are backups and a migration ledger now).

Landed: June to July 2026: staging.creddit.xyz and .github/workflows/deploy-staging.yml

Header updated 2026-09-14. The body below is frozen history. All plans.

Plan: staging + prod environments with a promote-to-prod flow ​

Stand up a staging environment alongside prod so all development happens on staging first and ships to prod only when validated. Decisions taken (2026-06-14): staging runs on the same Hetzner box as prod, and staging data is reseeded from a prod dump (staging does not run its own crons).

This doc is both the design and the runbook.


0. Current state (grounded, not assumed) ​

  • One Hetzner box (dexhq.io): 4 vCPU, 7.6 GiB RAM (~6 free), 150 GB disk (~90 free).
  • nginx + Certbot. creddit.xyz → 127.0.0.1:3001 (the onchain-credit Next.js app under PM2, next start). Box also runs dexhq (:3000), creddit-indexer, rindexer.
  • One Postgres cluster (127.0.0.1:5432): db creddit (~841 MB), owner postgres, app login role onchain_credit. Holds PII: the 006-newsletter-subscribers table has subscriber emails + user agents.
  • Deploy: .github/workflows/deploy.yml on push to main → SSH → git reset --hard origin/main + npm ci + npm run build + pm2 restart (rolls back the working tree on build failure). Code only; migrations/refreshers manual.
  • Cron (root): 6 refresher jobs (refresh-assets 6h, refresh-vault-risk weekly, refresh-sofr weekdays, refresh-vault-capacity/-collateral-exposure/ -lending-positions every 6h, offset).
  • Two gaps this plan closes: there are no DB backups, and migrations are hand-applied numbered SQL (001–034) with no ledger.

1. Target topology (same box, fully isolated staging) ​

LayerProd (exists)Staging (new)
Domaincreddit.xyzstaging.creddit.xyz (basic-auth + noindex)
Directory/opt/onchain-credit/opt/onchain-credit-staging
Git branchmainstaging
Port / PM2:3001 / onchain-credit:3002 / onchain-credit-staging
nginx vhostsites-enabled/onchain-creditsites-enabled/onchain-credit-staging
Databasecreddit / onchain_creditcreddit_staging / onchain_credit_staging
.env.localprodstaging (own DATABASE_URL, PORT=3002)
Cronactivenone (reseed + on-demand refresher runs only)
TLSCertbot (LE)Certbot (LE), same --nginx flow
Deploy keyDEPLOY_SSH_KEYseparate DEPLOY_SSH_KEY_STAGING (sec. 5.3)

Isolation boundaries that matter are present: separate database + role (a bad migration/write cannot reach prod), separate PM2 process + port, and separate directory + env.

Honest scope of "isolated" (same box, same Postgres cluster). Same-box staging shares the OS/kernel and the Postgres process, so a runaway staging query, an OOM, or a disk-fill can degrade prod. Same-box is the right call here, but only with the concrete guards in sec. 9 (role-level timeouts + connection cap, a PM2 memory ceiling, and disk alerting). A dedicated VM is the upgrade path the day infra-level testing (OS/nginx/PG-version) is needed.


2. Branching + promotion flow ​

Two long-lived branches, each bound to one environment:

feature/* ──PR──▶ staging ──auto-deploy──▶ staging.creddit.xyz   (develop + validate)
                     │
                     └──"Release" PR──▶ main ──auto-deploy──▶ creddit.xyz   (promote)
  • All feature PRs target staging. Merge → deploy-staging.yml rebuilds /opt/onchain-credit-staging on :3002.
  • Validate on staging.creddit.xyz.
  • Promote with a single staging → main PR. Its diff is the release. Merge → existing deploy.yml ships prod.
  • Direct commits to main stop (branch protection: PR-only). After each release PR, main and staging are in sync.

Tradeoff stated explicitly (environment-branch promotion). A long-lived staging branch + "release PR" is the simplest model to operate, but it has two known costs: (1) the two branches drift and can conflict (this very PR's first revision was already dirty from a stale base — the failure mode is real), and (2) each box rebuilds from source, so staging and prod produce different build artifacts — you validate the code, not the exact shipped bytes. The modern alternative is build-once, promote-the-artifact (build a Docker image / .next bundle in CI, deploy the same artifact to staging then prod). Deferred on purpose: the current SSH+git reset+npm run build deploy is simple and working, and artifact promotion is a larger change. Revisit if drift or build-skew bites.

AGENTS.md / contribution rules update: "feature branch → PR into staging → validate → staging → main release PR → prod."


3. Database ​

3.1 Create the staging database + role (one-time) ​

sql
-- as postgres
CREATE ROLE onchain_credit_staging LOGIN PASSWORD '<staging-pw>'
  CONNECTION LIMIT 20;                          -- cap (sec. 9)
ALTER ROLE onchain_credit_staging SET statement_timeout = '30s';   -- guard (sec. 9)
ALTER ROLE onchain_credit_staging SET idle_in_transaction_session_timeout = '60s';
CREATE DATABASE creddit_staging OWNER postgres;
-- schema + grants are established by the first reseed (it carries them).

A separate role means a leaked staging credential cannot touch creddit, and the role-level limits cap staging's blast radius on the shared cluster.

3.2 Reseed from prod — scripts/ops/reseed-staging.sh (new), PII-scrubbed ​

# 1. dump prod (this dump is ALSO the backup, sec. 6) -> encrypted off-box
pg_dump -Fc creddit > /opt/backups/creddit-$(date -u +%FT%H%MZ).dump

# 2. restore into a fresh staging db
dropdb --if-exists creddit_staging && createdb -O postgres creddit_staging
pg_restore -d creddit_staging --no-owner --role=onchain_credit_staging <latest dump>

# 3. SCRUB PII before anyone can reach staging (sec. below)
psql -d creddit_staging -f scripts/ops/scrub-staging-pii.sql

# 4. re-grant to the staging role; ANALYZE.

PII scrub (required). The prod dump carries real subscriber emails (newsletter_subscribers) and user agents. Copying them into a basic-auth-only staging DB is a privacy/GDPR exposure. scrub-staging-pii.sql masks them, e.g. UPDATE onchain_credit.newsletter_subscribers SET email = 'sub'||id||'@staging.invalid', user_agent = NULL; (extend the script as any future PII-bearing table lands; a test asserts the known PII columns are covered). Everything else in the schema is public on-chain data and is fine to copy verbatim.

Migration-vs-reseed ordering. A reseed brings staging to prod's current schema, so it reverts any not-yet-promoted migration under test. Either reseed before a migration's testing window, or re-run migrate.sh (sec. 4) after a reseed to re-apply staging-only pending migrations. The nightly reseed is pause-able (a flag file) so it never clobbers an active test.

3.3 Why same cluster, not a second Postgres ​

A separate database + role is the isolation unit. A second cluster only buys PG-version/config-experiment isolation, not a current need. Same cluster keeps pg_dump/pg_restore, backups, and monitoring trivial.


4. Migration ledger (replaces hand-applied SQL) ​

A deterministic, idempotent runner so "apply pending migrations" is one command per environment, and the promotion checklist can't miss or double-apply.

  • onchain_credit.schema_migrations(filename text primary key, applied_at timestamptz default now()).
  • scripts/ops/migrate.sh <DATABASE_URL>: lists scripts/sql/*.sql in order, skips any already in schema_migrations, applies the rest, records them. Forward-only. Each file runs inside a transaction.
  • Concurrency + failure semantics (explicit): the runner takes a pg_advisory_lock so two deploys can't migrate at once, sets a lock_timeout and statement_timeout so a blocked migration fails fast instead of wedging the cluster, and exits non-zero on any failure (the deploy job then aborts before pm2 restart). Note the code-rollback path in deploy.yml does NOT roll back the database — migrations are forward-only and must be backward-compatible with the previous code (expand/contract), so a rolled-back build still runs against the migrated schema.
  • Retrofit: seed schema_migrations with every existing 0xx-*.sql marked applied on BOTH creddit and creddit_staging (already applied), so the runner starts from "nothing pending."
  • Additive vs destructive: additive migrations (ADD COLUMN IF NOT EXISTS, new tables — the house style) are safe to auto-apply. Destructive ones (e.g. the 032 stress-column DROP) stay a gated manual step run after the deploy that stops referencing the object; the runner refuses a file tagged -- DESTRUCTIVE without --allow-destructive.

5. CI/CD ​

5.1 deploy-staging.yml (new) — clone of deploy.yml, parameterized ​

  • Trigger: on: push: branches: [staging].
  • Same pinned-host-key SSH pattern, its own DEPLOY_SSH_KEY_STAGING (sec. 5.3).
  • Server step targets /opt/onchain-credit-staging, fetches origin/staging, builds, then runs migrate.sh "$STAGING_DATABASE_URL" (additive auto-apply), then pm2 restart onchain-credit-staging. Same build-failure rollback.
  • Concurrency group deploy-staging (independent of deploy-production).
  • Prod's deploy.yml gains the same migrate.sh line later, once trusted.

5.2 Post-deploy health gate (new, both envs) ​

Deploy success today = "pm2 restarted," with manual validation. Add an automated gate: after restart, the job curls a health endpoint (a lightweight /api/health returning a 200 + a DB-reachable check) and fails the deploy if it is not 200 within a timeout. On staging, optionally a couple of smoke routes (/money-market-rates, /carries) returning 200. This turns a broken deploy into a red CI run instead of a silently-down site.

5.3 Per-environment deploy keys + branch protection ​

  • Separate SSH deploy keys per environment from day one (cheap; a staging-deploy compromise must not equal prod-box access). Two keypairs, two authorized_keys entries (optionally command=/from= restricted), two repo secrets: DEPLOY_SSH_KEY (prod), DEPLOY_SSH_KEY_STAGING.
  • main: PR-only, no direct pushes; require deploy-staging green + the health gate on the head commit before a release PR merges.
  • staging: PR-only; require tsc + tests green.

6. Backups (prerequisite + standalone win) ​

Today there are none. Establish before staging (the reseed dump is the backup):

  • Nightly pg_dump -Fc creddit → /opt/backups, then encrypted (age/gpg) and pushed off-box (Hetzner Storage Box / S3-compatible via rclone) with a checksum recorded, ~14-day retention.
  • Monthly restore-test: the reseed into creddit_staging continuously proves the dump restores; document the restore runbook.
  • RPO/PITR tradeoff (stated): nightly pg_dump = up to 24h RPO, no point-in-time recovery. Acceptable here because the dataset is re-derivable from chain (the refreshers rebuild it) and the only non-derivable data is the newsletter table — so the deliberate decision is: nightly logical dump now; consider WAL-archiving/PITR only if non-derivable data grows. (DEX_HQ's dexhq DB should get the same nightly dump, noted, out of scope.)

7. nginx + DNS + TLS + access control ​

  • DNS: staging.creddit.xyz → the box (same provider as creddit.xyz).
  • New vhost sites-enabled/onchain-credit-staging: server_name staging.creddit.xyz; → proxy_pass http://127.0.0.1:3002;, 80→443 redirect.
  • certbot --nginx -d staging.creddit.xyz (same flow that issued the prod cert).
  • Non-public + non-indexed (we have prior reputation/blocklist sensitivity):
    • auth_basic "staging"; auth_basic_user_file /etc/nginx/.htpasswd-staging; (or Cloudflare Access if/when DNS moves to CF).
    • add_header X-Robots-Tag "noindex, nofollow" always; + a staging robots.txt disallow-all.

8. App instance on the box ​

  • git clone the repo to /opt/onchain-credit-staging, checkout staging.
  • /opt/onchain-credit-staging/.env.local: DATABASE_URL → creddit_staging (role onchain_credit_staging), PORT=3002, NODE_ENV=production. RPC/0x keys: start shared (reads are idempotent; cheapest), move to separate keys later if independent rate-limit/observability is wanted. Document that they're shared.
  • npm ci && npm run build, then pm2 start npm --name onchain-credit-staging -- start with PORT=3002 and a memory ceiling (--max-memory-restart 700M, sec. 9), pm2 save.
  • Resource check: staging app ~60 MB + creddit_staging ~0.85 GB copy + occasional refresher runs — trivially within ~6 GiB free / 90 GB disk.

9. Same-box safety guards (the price of not using a separate VM) ​

Concrete, low-effort guards so staging cannot take prod down:

  • Postgres role limits (sec. 3.1): statement_timeout=30s, idle_in_transaction_session_timeout=60s, CONNECTION LIMIT 20 on onchain_credit_staging — a runaway staging query is killed, not left to saturate the shared cluster.
  • PM2 memory ceiling: --max-memory-restart 700M on the staging process so a leak restarts staging instead of OOM-killing prod.
  • Disk alerting: a simple cron that alerts (or the existing monitoring) when / crosses ~80%, since staging's DB copy + nightly dumps consume disk. Retention on /opt/backups (14 days) bounds growth.
  • Postgres shared_buffers/work_mem stay prod-sized; staging's connection cap keeps it from inflating memory pressure.

10. Refreshers / cron on staging ​

  • No cron on staging. Data comes from the reseed (sec. 3.2).
  • Validate a refresher change by running it once on staging by hand: cd /opt/onchain-credit-staging && scripts/run-cron.sh <refresher>.ts (its run-cron.sh sources the staging .env.local, so it writes creddit_staging).

11. The indexer ​

The nightly reseed carries creddit-indexer's tables from prod, so staging needs no separate indexer to render. A staging indexer is only required when testing indexer changes themselves (a second rindexer pointed at creddit_staging) — defer to that work.


12. Docs (Cloudflare Pages) ​

Already has branch previews. Point the staging branch's preview at staging-docs.creddit.xyz (or use the auto preview URL). No new infra.


13. Rollout order (each step is additive; prod untouched until the cutover) ​

  1. Backups: nightly encrypted+checksummed pg_dump off-box + restore runbook. (standalone win)
  2. Staging DB + role (with role limits); reseed-staging.sh + scrub-staging-pii.sql; first reseed.
  3. Staging app: dir, .env.local, PM2 :3002 with memory ceiling, build.
  4. nginx vhost + cert + basic-auth + noindex for staging.creddit.xyz; smoke-test.
  5. Separate staging deploy key; staging branch; deploy-staging.yml (with migrate.sh + health gate); push to confirm.
  6. Migration ledger + migrate.sh; retrofit existing 0xx as applied on both DBs.
  7. Same-box guards (sec. 9) + disk alerting.
  8. Cutover: branch protection on main (PR-only); switch the contribution flow to target staging; update AGENTS.md. (only now does the dev flow change)
  9. (optional) nightly reseed cron + pause flag; separate staging RPC keys; /api/health route if not already present.

Steps 1–7 are invisible to prod and to the current main-based flow; step 8 is the single switch that turns on "develop on staging, promote to prod."


14. Rollback / safety notes ​

  • Staging is disposable: any breakage is fixed by a reseed.
  • The promotion PR is the audit trail of every prod release.
  • Migrations are forward-only and must be backward-compatible (expand/contract); the deploy's code-rollback does NOT roll back the DB (sec. 4).
  • Destructive migrations keep the gated manual discipline, validated on staging first.

Open items to confirm before/while implementing ​

  • DNS provider for staging.creddit.xyz (same as creddit.xyz); basic-auth (works today) vs Cloudflare Access (if DNS moves to CF).
  • Off-box backup target (Hetzner Storage Box vs S3-compatible) + encryption key custody.
  • Auto-run additive migrate.sh inside deploy from day one, or keep migrations in the release checklist until the ledger is trusted.
  • Separate vs shared RPC/0x keys for staging (recommend shared to start).
  • Whether to adopt build-once/artifact-promotion now or accept env-branch rebuild drift for the first iteration (recommend accept-for-now; revisit if it bites).

Private documentation. creddit.xyz