Skip to content

Deployment & server ops ​

How creddit runs in production, how code ships to it, and how to operate the box by hand. The app is a Next.js App Router site backed by a local PostgreSQL database; data is produced by off-chain cron refreshers (see Data Pipeline) and the app reads only.

Convention reminder: every trailing APY across every venue is annualizeRatio in src/lib/data/apy.ts (the realised ratio of an on-chain compounding index between two blocks, annualised by actual elapsed time). Never average per-snapshot annualised rates. This is what lets /carries compare venues like-for-like, and it is why a refresher is just an upsert job, not business logic the app re-derives.

1. Server topology ​

Single Hetzner box, shared with the DEX HQ stack. One host runs everything.

ItemValue
Hostdexhq.io (Hetzner)
SSHssh root@dexhq.io (key ~/.ssh/hetzner_ed25519)
Repo/opt/onchain-credit — git checkout tracking origin/main
App processpm2 onchain-credit on localhost:3001
Edgenginx reverse-proxy → https://creddit.xyz, with Cloudflare in front
Databaselocal PostgreSQL; prod DB creddit, role onchain_credit, schema onchain_credit
Cron wrapperscripts/run-cron.sh <script>
Daily backupscripts/ops/backup-creddit.sh at 02:15 UTC
Log rotation/etc/logrotate.d/onchain-credit (daily, 14 days) — see Log rotation

pm2 processes co-resident on the box:

pm2 processPortWhat it is
onchain-credit3001this app (prod)
onchain-credit-staging3002staging copy of this app (see §3)
rindexer—shared indexer
dexhq—DEX HQ app
creddit-indexer—shared indexer
creddit-event-ingester—always-on event-ledger ingester (portfolio 100k scale) — a manual add, see below

nginx terminates the public hostname and proxies to localhost:3001; Cloudflare sits in front of nginx for TLS and CDN. The app binds loopback only — it is never reachable except through nginx.

scripts/ops/ (backup, reseed, migrate, PII scrub, nginx template) is in the repo and ships via the normal deploy. The server crontab and the box-local secrets (.env.local, /root/.staging-basic-auth) live only on the box.

.env.local is the only place a key lives, and git reset --hard preserves it (step 3 below). A key the repo has never seen is a key a deploy cannot leak. Each environment holds its own: DATABASE_URL, SESSION_SECRET, the RPC overrides, DUNE_API_KEY + the saved-query ids, and COINGECKO_API_KEY — the Demo key behind every live price, the hourly history that fills a hole the tape never printed, and the weekly reported-volume reading the pricing-category test is judged on. Every variable, with its default and what it costs when absent, is in Environment variables. A missing COINGECKO_API_KEY is not an outage: the same host answers keyless at a much lower rate limit and a refused call degrades to DefiLlama and then to the stored bar, so the symptom is more assets served from the backup vendor rather than an error. A new environment, and an environment this release reaches for the first time, needs the line added by hand (pm2 restart <app> --update-env to pick it up).

2. How code ships (promote-to-prod) ​

Two long-lived branches, each bound to one environment. You develop on staging and promote to prod:

feature/* ──PR──▶ staging ──auto-deploy──▶ staging.creddit.xyz   (develop + validate)
                     │
                     └──"release" PR──▶ main ──auto-deploy──▶ creddit.xyz   (promote)
  • Feature work targets staging. Merge a feature PR into staging → deploy-staging.yml rebuilds /opt/onchain-credit-staging on :3002 and it appears on staging.creddit.xyz (basic-auth).
  • Validate on staging.creddit.xyz.
  • Promote with one staging → main release PR. Its diff is the release. The PR bumps version in package.json. Merge → deploy.yml ships prod, and release.yml tags v<version> + cuts a GitHub Release (see §2.3).
  • main and staging are in sync after each release.

Design + rationale (same-box isolation, reseed model, tradeoffs): docs/plans/staging-environment-plan.md.

Branch protection is advisory, not enforced. This is a free-plan private repo, so GitHub cannot enforce PR-only main or required checks. ci.yml (tsc + tests on every PR into main/staging) and the scripts/ops/git-hooks/pre-push guard report but cannot block. Honor a red CI run; don't push to main directly. To make it enforced, upgrade to GitHub Pro (~$4/mo) and add a branch rule on main: require a PR, require the CI check green, and (optionally) restrict who can merge. Nothing else changes.

Shipping a change, step by step ​

The concrete loop, from an idea to it live on creddit.xyz:

  1. Branch off staging (never main): git fetch && git switch -c feature/x origin/staging.
  2. Do the work, then locally npx tsc --noEmit && npm test. (npm test now runs npm run typecheck:scripts first — tsconfig.json EXCLUDES scripts/, and the tests there run through tsx, which strips types without checking them, so every refresher and sync CLI was unchecked until tsconfig.scripts.json was added. That blind spot shipped a dead status guard that gone-marked the whole Morpho market universe; the gate would have failed it.) If you touched a page/metric/schema/refresher, update the matching docs/ page in the same change (the docs-in-the-PR rule).
  3. Open a DRAFT PR into staging. ci.yml runs tsc + tests on every push; its browser jobs wait until the PR is marked ready for review, which happens once the independent review has converged (AGENTS.md, "Shipping"). Merge it → deploy-staging.yml rebuilds staging and it appears on https://staging.creddit.xyz (basic-auth).
  4. Validate on staging. Load the changed page; if the change needs data, run the refresher/backfill once on the box against creddit_staging (cd /opt/onchain-credit-staging && scripts/run-cron.sh <script>.ts).
  5. Promote. Bump version in package.json on staging (minor for a feature, patch for a fix), then open a staging → main PR and merge it with a merge commit (not squash — that keeps staging an ancestor of main). This is the release; get explicit sign-off before merging, since it ships prod.
  6. Merge → prod ships. deploy.yml deploys creddit.xyz and restarts both the app and the event ingester; release.yml tags v<version> and cuts a GitHub Release. Then fast-forward staging back up to main so they match: git push origin origin/main:staging.
  7. Post-deploy, if the change touched data. Apply prod migrations by hand (scripts/ops/migrate.sh against creddit), run the refresher/backfill, then re-run the deploy so the ISR pages re-prerender against the new data (see the ISR note).

Hotfix shortcut: for an urgent prod-only fix you can branch off main, PR into main, and back-merge to staging after — but the default is always through staging.

2.1 Production deploy pipeline (.github/workflows/deploy.yml) ​

Triggers on every push to main (in practice, the merge of a release PR). Code only. It does not run DB migrations, backfills, or data refreshers; those are manual server steps after deploy (see below).

The GitHub runner only opens an SSH session. The server fetches the new code from this repo and rebuilds in place. Step by step:

  1. SSH in. Runner writes DEPLOY_SSH_KEY (repo secret) to a temp key and pins the server host key (StrictHostKeyChecking=yes, hard-coded known_hosts line), then ssh root@dexhq.io. The whole deploy script is passed as a single SSH argument so it runs from memory — a git reset rewriting files on disk cannot disturb the running script.
  2. Record rollback target. PREV=$(git rev-parse HEAD) in /opt/onchain-credit.
  3. Fetch + reset. git fetch https://x-access-token:$GH_TOKEN@github.com/<repo>.git main (ephemeral repo-scoped GITHUB_TOKEN) then git reset --hard FETCH_HEAD. .env.local and node_modules are gitignored, so the reset preserves the server's secrets and installed deps.
  4. Install + build. npm ci --no-audit --no-fund && PG_POOL_MAX=4 npm run build. The PG_POOL_MAX=4 caps the pg pool for the build only: next build prerenders the ISR pages across several worker processes, each of which opens its own pool (see postgres.ts), so the default max of 12 per worker can trip an intermittent too many connections for role build failure that then rolls the deploy back. Four keeps the build under the role's connection cap; the runtime pm2 process leaves the var unset, so serving keeps the full pool of 12. If a build still fails this way it is transient — gh run rerun --failed clears it, no code change needed.
  5. Restart (success path). pm2 restart onchain-credit --update-env, then scripts/ops/restart-ingester.sh, which restarts the two always-on processes, the event ingester and then the ledger worker (each loads its code once, at start; see the deploy restarts the ingester and the worker), then log the deployed short SHA + timestamp. The script leaves a deliberately stopped process alone with a ::warning::; if pm2 refuses a restart, or the ingester is missing, it prints an ::error:: and the run exits 1, since the app is then live on the new code and that process is not. A missing worker is a ::warning:: while the build still derives inline and an ::error:: once the worker owns derivation (that section says why). Deployed <sha> means pm2 accepted the restarts, not that the ingester's first cycle or the worker's first job succeeded: read their logs after a release (see the post-deploy check in that section).
  6. Rollback (failure path). If npm ci/npm run build fails: git reset --hard $PREV, reinstall, rebuild from the previous commit, run the same restart (both processes are back on the code they already ran, and the rebuild would otherwise read as a missed restart to the freshness alarm), and exit 1. A failed deploy never leaves broken code or a half-built .next staged for the next process restart.

Other guarantees:

  • One at a time. concurrency: group: deploy-production, cancel-in-progress: false — a second push waits its turn rather than racing.
  • Least privilege. Workflow permissions: contents: read; the deploy key only authorises SSH to the box; the GitHub token is ephemeral and repo-scoped.
  • Host-key pinning is verified out-of-band against /etc/ssh/ssh_host_ed25519_key.pub. Update the known_hosts line in the workflow if the server is ever rebuilt, or every deploy will fail host-key verification.

Manual post-deploy steps (when a change needs more than code) ​

The pipeline restarts the app, the event ingester and the ledger worker with new code and a fresh build, nothing else. If your change touches the database or data, do these on the box after the deploy lands:

  • Schema migration → apply the new scripts/sql/NNN-*.sql as the postgres owner (see §4, Migrations).
  • Backfill → run the relevant scripts/backfill-*.ts via run-cron.sh (one-off history seeds for newly added wrappers/markets; they reuse the same annualisation math as the live refreshers).
  • Refresher → run the affected scripts/refresh-*.ts via run-cron.sh so the new rows exist immediately instead of waiting for the next cron tick.
  • Carry registry → after a new yield adapter or DEX-pool snapshot lands, re-run scripts/sync-carries.ts then --approve <key> the newly-PROPOSED carries (replaces the retired sync-fluid-vaults.ts; see processes A.0).
  • Pendle PT payout assets (migration 102) → apply it after the deploy, not before, and nothing else is required. The reasoning and the one-off tape cost are in the release step.
  • Pendle redemption-index factor → after migration 101 nothing is required: the 6h refresher fills the factor going forward, and fills a newly discovered market's history itself (at most two markets a tick, after the snapshot loop; it also drains markets that still hold no pendle_api history, narrowed in-query through portfolio_held_pts and read only back to the ledger derivation floor, rotating so no market can hold the head of the queue, so a deferred or failed pass retries on its own). backfill-pendle-history.ts --market <address> stays the manual tool for a market discovered before that wiring shipped, and for a market whose automatic pass was short — the fill is idempotent, so re-running only ever adds (4 to 5 chain reads per row, so per market, never in bulk). The sequence that ran once, when 101 shipped, is in the release-steps log.
  • Money market fund registry → after migration 086, run scripts/sync-money-market-funds.ts (discovery + classification), then refresh-assets.ts for the first state read, then backfill-money-market-funds.ts for the return history. Approving a manager needs no deploy: the refresher, the share-rate snapshotter and the backfill all read their work-list from the registry. The sequence that ran once, when 086 shipped, is in the release-steps log.

Release steps are not in this runbook ​

A note written for the deploy of one migration, or a gated repair written for one release, is executed a single time in each environment and is then finished. None of them lives here. They are in the release-steps log, verbatim, each carrying the migration or release it belongs to and an executed: status, so an operator can tell a step that is still outstanding from one that is done — which is exactly what could not be told while they were interleaved with the standing procedure on this page.

This page holds only what is true on any day: topology, shipping, the deploy pipeline, staging, versioning, the migrations procedure, the crontab, backup, verification and operations. scripts/runbook-split.test.ts fails the build if a one-time step is written back into it.

ISR re-prerender note ​

Pages are App Router with per-page ISR (revalidate 1800 or 3600 seconds), prerendered at build time. The build snapshots whatever the DB held when npm run build ran. So:

If a refresher (or a manual backfill) writes new data after the deploy build has already run, the prerendered pages keep serving the stale snapshot until their revalidate window elapses (up to 30/60 min). To make fresh data appear immediately, re-run the deploy (push an empty commit or re-run the workflow) so the build re-prerenders against the now-current DB. This matters most for capacity/risk pages where a refresher and a deploy land close together.

The clean ordering is therefore: migrate → backfill/refresh the data → then deploy (or re-deploy), so the build prerenders against complete data.

2.3 Versioned releases ​

Every promotion to prod is a numbered release. .github/workflows/release.yml runs on push to main when package.json changed:

  1. Reads version from package.json.
  2. If a release for v<version> does not already exist, it tags v<version> at the merge commit and cuts a GitHub Release with auto-generated notes (the merged PRs/commits since the previous tag).
  3. If the version was not bumped, or the release already exists, it is a no-op.

So: the release PR bumps version (semver: minor for features, patch for fixes), and merging it both ships prod (deploy.yml) and records the release (release.yml). The tag is the durable "what shipped, when" record; deploy.yml is still what actually deploys. git tag --list / the repo's Releases page is the prod history.

3. Staging ​

A second, independent copy of the app on the same box, fed by a scrubbed copy of prod data. This is where all development lands and is validated before promotion (see §2).

ItemValue
URLhttps://staging.creddit.xyz (Let's Encrypt cert, HTTP→HTTPS)
AccessHTTP basic-auth (/etc/nginx/.htpasswd-staging; credential in /root/.staging-basic-auth) + noindex
Repo/opt/onchain-credit-staging, tracking origin/staging
App processpm2 onchain-credit-staging on :3002 (memory ceiling --max-memory-restart 700M)
Databaseseparate DB creddit_staging, role onchain_credit_staging (conn-limit 20, statement/idle timeouts)
Deploy keyseparate DEPLOY_SSH_KEY_STAGING (a staging-deploy compromise ≠ prod-box access)

3.1 Deploy (.github/workflows/deploy-staging.yml) ​

Auto-deploys on push to staging (the develop flow). Same SSH + git reset + build + pm2 mechanism as prod, pointed at /opt/onchain-credit-staging and :3002, using the separate staging deploy key. It builds with the same PG_POOL_MAX=4 cap as prod; the creddit_staging role's connection limit is only 20, so the smaller build-time pool matters most here (this is where the too many connections for role build flake was first observed). After build it runs scripts/ops/migrate.sh against creddit_staging (additive migrations auto-apply on staging) and a :3002 health check, with the same build-failure rollback.

workflow_dispatch (with a ref input, default main) refreshes staging from any branch on demand — useful to reset staging to prod's exact state.

Post-deploy smoke ​

Once the deploy job succeeds, a second job in the same workflow (smoke) runs the Playwright end-to-end suite from tests/e2e against the build it just put on the box, in two projects: desktop-1360 and zoom-sweep. What this job catches is a page rendering empty against staging's real data, and that is the same answer at every viewport width, so one width project answers it. The sweep runs here as well as pre-merge because it is the one project whose answer changes with the data: staging carries the long asset names and row counts a deterministic fixture cannot, and long names are what collide. The remaining interaction projects run pre-merge instead — narrow-900 whole and the @reduced-motion-tagged tests, in CI's e2e-interaction job against the fixture — and laptop-1140 is a local-only project (1360 and 900 bracket the two shells, and the sweep already covers 1140's layout). Running all five here had grown to 20 minutes. It is a job in this workflow rather than a workflow of its own on purpose: a separate file triggered by workflow_run fires only from the copy that sits on the default branch, so merging one into staging would arm nothing.

Gated on the diff. A third job, smoke-gate, decides in parallel with the deploy whether the smoke has anything to look at. It lists the files between the push's before and its head and lets the suite through when any of them is app source (src/app, src/components, src/lib, src/data), the e2e suite, playwright.config.ts, next.config.ts, postcss.config.mjs, anything under public/, or a dependency manifest. That is exactly the pattern ci.yml's e2e jobs use, minus scripts/fixture/, which belongs to the fixture this job does not render from. The config and asset paths are in it deliberately: headers and the CSP live in next.config.ts, the Tailwind pipeline in postcss.config.mjs, and public/ holds every logo, icon and font a page paints, so a diff touching one of them can break a page without touching src/. Since this job is the last line of defence for exactly those, its pattern is never narrower than what pre-merge CI gates on, with the single exception of scripts/fixture/, which it does not read. It stands down only when none of those changed — a merge of docs, crons, SQL or workflow edits alone — which is worth doing because a healthy smoke costs about six runner-minutes of a free plan's monthly allowance. Every uncertain answer runs it: a workflow_dispatch, a push with no usable before, a compare call that fails, a file list long enough to have hit the compare API's 300-file cap, a compare that names no file at all, and — at the job level — a smoke-gate that times out or loses its runner, since the smoke stands down for an explicit "no" from the gate and for no other answer of the gate's. (Its if does skip on two further things, neither of them a doubt about the diff: a deploy that did not succeed, and a cancelled run.) A stood-down smoke is a skipped job — grey, with its reason written into the run summary — never a red one, and the run stays green. To smoke a stood-down deploy anyway, use the by-hand smoke below: re-running the run replays the same push event and lands on the same answer, and a workflow_dispatch redeploys the box to its ref input first.

It needs no secret that does not already exist. The runner opens an SSH tunnel with the deploy key this workflow already uses (DEPLOY_SSH_KEY_STAGING) and points the suite at the app directly instead of at the public hostname. The same two lines are the manual smoke, from a checkout of whatever is currently on staging:

bash
ssh -f -N -L 3002:127.0.0.1:3002 root@dexhq.io   # staging's app port, firewalled from the internet
E2E_BASE_URL=http://localhost:3002 npm run e2e   # setting it also stops playwright.config.ts booting a local server
# CI narrows this to `npm run e2e -- --project=desktop-1360 --project=zoom-sweep`, and only when the push touched something the suite can see; by hand, run whichever projects you need.

Going through https://staging.creddit.xyz instead would mean sending the NGINX basic-auth credentials, and NGINX holds only a one-way $apr1$ hash of that password. There is no BASIC_AUTH_* repo secret, none can be derived from the box's NGINX config, and none is needed.

Report only. The job is continue-on-error: true, so a red suite never paints the deploy run red: by then the deploy has succeeded and the new build is already live, and the build-failure rollback above stays the only automatic revert here. The signal is an error annotation, a block in the job summary, and, on failure, the Playwright HTML report plus traces uploaded as the run artifact staging-smoke-<run-id>-<attempt> (14-day retention). That artifact is what to read before opening the release PR.

Everything that ran passed, with the signed-in specs skipped, is the healthy result. That skip floor is large and expected: every authed /portfolio spec mints its own session cookie and skips when the server does not sign with the same secret, and staging's SESSION_SECRET is deliberately not handed to CI. It is a floor rather than a fixed number — it moves whenever a signed-in spec is added, and it is counted once rather than once per viewport project. A higher skip count is not lost coverage either. requireRows skips an interaction test when the deployment serves fewer rows than the test needs, so a nightly reseed that leaves a screener thin shows up here as an extra skip. Read it as a note about staging's data, not as a broken smoke.

The smoke runs inside the deploy-staging concurrency group, so a second merge queues behind it instead of restarting pm2 underneath a running suite. The group is held for the whole run, and the smoke step's timeout-minutes: 20 (with a timeout-minutes: 25 backstop on the job around it) is therefore also the longest a wedged staging can keep the next deploy waiting. When a recovery dispatch cannot wait, cancel the in-flight run from the Actions UI: its deploy job has already finished by the time the smoke is running, so cancelling costs only the report. Specs, fixture DB and the per-surface checklist: Processes → Verifying a UI change.

3.2 Data: reseed, not cron (by design) ​

Staging runs no refreshers (the one scheduled job in its checkout is the minutely drain-portfolio-backfills.ts, which serves user-triggered backfills rather than refreshing market data). Its data is a nightly scrubbed copy of prod (03:00 UTC cron, after the 02:15 backup), plus on-demand, via scripts/ops/reseed-staging.sh:

bash
# on the box
scripts/ops/backup-creddit.sh       # fresh prod dump -> /opt/backups (also the nightly backup)
scripts/ops/reseed-staging.sh       # drop+recreate creddit_staging from newest dump, scrub PII, re-grant

The reseed masks PII (scrub-staging-pii.sql: subscriber emails hashed to @staging.invalid, user agents dropped; the whole user-data graph — the four chat tables, accounts, account_wallets, the two portfolio history tables, portfolio_wallet_index and portfolio_backfill_state — TRUNCATEd in one statement, since a reseed restores the whole prod DB with no table-exclusion and a wallet's real positions must not ride into staging; one statement because every one of those tables FK-references accounts (as of migration 056 the chat tables and portfolio_wallet_index do too), and Postgres refuses to truncate an FK-referenced table unless every referrer is in the same statement) and fails closed if any unmasked email survives, or any account/portfolio table (accounts, account_wallets, both history tables, portfolio_wallet_index, portfolio_backfill_state) still holds rows after the scrub, or any per-wallet cursor (chain_scan_cursors, scope portfolio:derive:v2:%, or the retired discovery watermark portfolio:discover:%) survives it. That last one is a CORRECTNESS hazard for the ledger prefix: chain_scan_cursors has no FK to accounts, so a surviving per-wallet cursor certifies work against a table the TRUNCATE has just emptied — the shape of the 2026-07-21 incident, which the discovery watermarks once reproduced on staging for every wallet in the dump. The watermarks are retired and migration 106 deletes them, but a pre-106 prod dump still carries them (prod's enrolled-wallet list), so the scrub keeps that prefix until a reseed from a post-106 dump has run. The scrub deletes those rows (prefix-scoped, so every other refresher cursor survives as it must). Because the four portfolio-workstream tables are emptied, re-run scripts/ops/seed-portfolio-fixtures.ts after each reseed to re-register the public test whales (or touch /root/.reseed-paused during a multi-day validation window so the reseed is skipped and the fixtures persist). It carries prod's current schema and its schema_migrations ledger, so after a reseed migrate.sh reports 0 pending. A reseed reverts any not-yet-promoted migration under test; pause it with touch /root/.reseed-paused. To validate a refresher change on staging, run it once by hand: cd /opt/onchain-credit-staging && scripts/run-cron.sh <refresher>.ts (its run-cron.sh sources the staging .env.local, so it writes creddit_staging).

3.3 Same-box safety guards ​

Staging shares the box and the Postgres cluster with prod, so the guards matter: the staging role's statement_timeout/idle_in_transaction_session_timeout/connection cap kill a runaway staging query, the PM2 --max-memory-restart restarts staging (not prod) on a leak, and hourly disk alerting bounds the backups + DB-copy growth. A dedicated VM is the upgrade path if OS/nginx/PG-version testing is ever needed.

4. Operating the server ​

SSH and pm2 ​

bash
ssh root@dexhq.io                       # key ~/.ssh/hetzner_ed25519

pm2 list                                # all processes + status
pm2 logs onchain-credit                 # tail prod app logs
pm2 logs onchain-credit --lines 200     # recent
pm2 restart onchain-credit --update-env # restart prod (re-reads .env.local)
pm2 restart onchain-credit-staging --update-env

Event ledger ingester (portfolio 100k scale) ​

The event-ledger ingester (scripts/ingester/ingest-events.ts) is the wallet-count-independent spine of the 100k-wallet plan. Unlike the refreshers it is not a cron: it is an always-on PM2 process with its own ~60s loop. Bringing it up is a manual server step, in order:

  1. Apply migration 064 (additive + idempotent; staging auto-applies on deploy, prod is the usual gated migrate.sh). It creates raw_events (range-partitioned)

    • event_coverage and hands raw_events to the onchain_credit role so the ingester can auto-create partitions at runtime (see database.md).
    bash
    cd /opt/onchain-credit && scripts/ops/migrate.sh "$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)"

    Ownership caveat on staging. The parent ends up owned by onchain_credit (the prod role), and migration 064 also GRANTs that role CREATE on the schema so on prod the ingester and backfill auto-create partitions at runtime. Staging connects as onchain_credit_staging, which neither owns the parent nor is the grantee, so its runtime partition auto-create fails. As of this rollout (the band has since grown down to p90; current state below) the pre-created band covered only the live tip (the ingester's forward inserts), NOT the one-time backfill: the 2025-05-21 floor (~block 22.5M, partition p90) sat below the band's p98 floor, so step 3's backfill needed auto-create even for the validation window and failed LOUD at the time (ensurePartitions never does a silent bad insert). Loud is no longer today's symptom, and that is the trap: 064's own backfill has since materialised p90–p112, so the floor lands in an existing partition and the missing privilege now shows up as a swallowed error per partition index, with the loud failure not due until the head crosses 28,250,000 (measured, see below). The step is still required and its absence is now silent. Before exercising the backfill on staging, give the staging role both halves on that box — ownership of the parent alone is not enough, since CREATE TABLE … PARTITION OF also needs CREATE on the schema:

    sql
    ALTER TABLE onchain_credit.raw_events OWNER TO onchain_credit_staging;
    GRANT CREATE ON SCHEMA onchain_credit TO onchain_credit_staging;

    Both repeat after every reseed. Current state of the band, and how to verify the privileges directly: Migrations 082 / 083.

  2. Start the PM2 process. run-ingester.sh sources .env.local and execs the long-running loop, so PM2's stop signal reaches the ingester's graceful-shutdown handler: it sets a stop flag, finishes the current cycle and exits 0. Every prod deploy restarts it, so that is routine: a live-scan cycle advances its cursor last, so an interrupted one re-scans its range, and the reorg repair is one transaction that either lands whole or not at all.

    --kill-timeout is what makes "finishes the current cycle" true. pm2's default is 1.6 seconds, so without the flag a cycle still running at a restart is SIGKILLed rather than allowed to finish. 120s covers an ordinary cycle and costs a restart nothing, since pm2 waits only until the process exits; a cycle in the middle of a heavy archive catch-up can still outlast it, which is why the repair that a kill used to corrupt is a transaction rather than merely a slower thing to interrupt. The value is stored in the process definition, so pm2 restart does not add it to a process that was created without one — an existing ingester needs the one-time step.

    bash
    cd /opt/onchain-credit
    pm2 start scripts/ingester/run-ingester.sh --name creddit-event-ingester --time --kill-timeout 120000
    pm2 save                              # persist across reboots (and the flags above)
    pm2 logs creddit-event-ingester

    It reuses ETHEREUM_RPC_URL (live head + live-window getLogs) and ETHEREUM_ARCHIVE_RPC_URL (catch-up history + block timestamps) from .env.local.

  3. GATED STEP: run the historical backfill, one stream at a time (attended; fills raw_events from the 2025-05-21 floor up to the ingester's live cursor; resumable, archive RPC). This is a prod write and needs explicit approval before it starts.

    Run the smallest streams first, so the harness is proven cheaply, and never two concurrently — the run is bounded by the archive endpoint's rate limits and by disk, and its wall-clock per stream is a figure the capacity model depends on. Between streams re-read df -h / and stop if free space has fallen below 25 GB.

    bash
    cd /opt/onchain-credit && ETHEREUM_ARCHIVE_RPC_URL=<archive> \
      DATABASE_URL="$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)" \
      npx tsx scripts/backfill-event-ledger.ts --stream <name> --dry-run   # plan + expected rows
    cd /opt/onchain-credit && ETHEREUM_ARCHIVE_RPC_URL=<archive> \
      DATABASE_URL="$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)" \
      npx tsx scripts/backfill-event-ledger.ts --stream <name>

    Run --help first when preparing a command. Help, unknown options and positional arguments exit before contacting the database or an RPC provider. Final completion now requires the measured count band, uninterrupted receipts, named anchors, and a sampled comparison from LEDGER_VERIFY_RPC_URL; raw events and receipts remain resumable when any check fails, but no completion record, rollout marker, or live handoff is written. The pending-withdrawal history band is now exactly 294,592 rows, enumerated from an independently indexed copy of the chain pair by pair, and replaces the earlier 280,000–325,000 range, whose four venue subtotals were counts of the rows it was meant to judge. Both of the old endpoints are refusals now. It applies only to that completed history sweep, from the campaign floor through its frozen original handoff at block 25,828,233. Events from the next block onward are live activity and are excluded even when acceptance is retried later; a later or extended range needs its own preflight samples.

    The Pendle Router V4 history band is also strict: 671,654 records from the same floor through its frozen original handoff at block 25,833,248. That total came from a complete attended preflight over 662 small windows; independent early, middle, and recent checks matched the second archive provider. Later live router activity is deliberately excluded from this one-time historical acceptance.

    --budget N bounds an invocation to N windows if the window has to be broken up; it stamps no completion and the next run continues from the resume cursor. A stream swept over several such runs still certifies from the floor: the certificate is a claim about the range, made by every run together. Recovering a known hole takes an explicit range — --from A --to B — which sweeps that range whatever the resume cursor says; a run with no --from resumes and would skip it. An explicitly ranged run, and any run capped by --to below the stream's target, is not a completion: it prints not a completion run; the rollout marker was not evaluated, leaves the live cursor alone, and exits zero. That line is expected on a repair and is not a stop.

    Then, before starting the next stream, take the verdict from the command and not by eye. --verify evaluates four numeric conditions (row count against the stream's band, receipt contiguity, a sampled digest re-check against the second provider, and the stream's named anchors together with a live-from-the-floor certificate for every address it declares) and exits non-zero with a [backfill-ledger/fail] … line if any misses. Condition 4 refuses outright when a stream would leave it nothing to check — no anchor and no declared address set is a refusal, not a pass. Any condition missing = stop.

    The band condition 1 uses must be independent of the rows it judges. Every stream carries either a density forecast from the data contract (compared with a 70%–200% allowance) or an enumerated count from a chain scan (compared exactly). A stream whose forecast is wrong takes a fresh enumeration, never a count of what was imported:

    bash
    cd /opt/onchain-credit && ETHEREUM_ARCHIVE_RPC_URL=<archive> \
      LEDGER_VERIFY_RPC_URL=<second provider> \
      DATABASE_URL="$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)" \
      npx tsx scripts/record-ledger-backfill-band.ts --stream <name> --to <handoff> \
        --note "<date> preflight" [--window 1000] [--dry-run] --confirm

    It refuses when the verify endpoint is the one the sweep's own scan receipts name as the provider the corpus was collected from (the environment is only consulted for a stream that has no receipt yet), when --to is above the finalized head, when --to is anything other than the frozen handoff for the two streams whose range is fixed (escrow → 25,828,233, pendle-router → 25,833,248), and for wallet-token (whose expectation moves with the enrolled wallet set). --dry-run measures and writes nothing. Recording a band approves nothing: --verify's four conditions still decide. A band recorded before this provenance rule existed is quarantined by migration 089 and named in the refusal; re-measure it with the command above rather than editing the row.

    Size these before you start one. The enumeration is a full eth_getLogs walk of the whole 3.3M-block range against the free second provider — the same endpoint that times out on erc4626 at a 5,000-block window — so run them one at a time, --dry-run first, and expect them to be long:

    streamlogs to enumeratesuggested --window
    escrow~295k over 17 declared (address, topic0) pairs5,000
    pendle-router~672k, whole stream, no address filter5,000 (662 windows in the 2026-08-25 attended run)
    erc4626~1.25M over 546 addresses1,000 — 5,000 times out on the address-set width
    fluid-dex~24k over 49 pools5,000
    aave-token + spark-token~4.2M, the two together5,000, and budget for the longest run of the campaign

    The pair is measured as a pair: --stream aave-token enumerates both and writes both rows, because their band is the pair's joint total and condition 1 sums their rows together.

    bash
    cd /opt/onchain-credit && ETHEREUM_ARCHIVE_RPC_URL=<archive> \
      LEDGER_VERIFY_RPC_URL=<second provider> \
      DATABASE_URL="$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)" \
      npx tsx scripts/backfill-event-ledger.ts --verify --stream <name>

    --verify is read-only. Pass a smaller --window for erc4626's verify leg: the free second provider times out on a 5,000-block range across its 546 addresses. Record, per stream, the rows written, the pg_total_relation_size delta and the wall-clock.

    On completion the runner stamps each marker-governed stream's rollout marker, and refuses to when a declared address is short of the floor or the receipts have a gap. That refusal happens inside the promotion transaction, after the coverage rows are written and before the marker is, and the address set it checks is re-read from the registry rather than taken from the run's own list — so an address added while the sweep ran is caught. A refusal rolls the whole promotion back: no coverage row, no marker, no live handoff, and the run exits non-zero. Every stream the rebuild adds is stamped this way, escrow included; the two whose address set never finishes (transfers, wallet-token) are not, by design. Details and the enable rule: the backfill runner.

    Do not answer that refusal by re-running the command. The run keeps its imported events, its receipts and its resume cursor, so a re-run scans nothing and goes straight to the promotion with the new address now in its own declared set — and the runner refuses again, by name, because it will not mark an address complete unless that address's history over the certified range was actually imported. The repair is to import it, and which half does it is the ownership split: wherever the tip loop owns new addresses — any stream that already carries its rollout marker, plus the two open sets (transfers, wallet-token) — its per-address catch-up sweeps the address from the floor unattended, so wait for caught up <address> in creddit-event-ingester. On a stream whose history this runner still owns (no marker yet, no address live), re-sweep it with an explicit --from <floor> --to <handoff>, which is logged as a repair and promotes nothing on its own. Then run step 3's plain command again.

    That repair has to run to completion, and this is the part that is easy to get wrong. The runner records the blocks it has actually committed, per address, as it commits them — so a re-sweep sized with --budget, killed by an RPC error (a [fail] line, and on a budget-bounded run the exit code is still 0), or stopped with Ctrl-C enrols the address over exactly the slice it covered, and the next promotion refuses on the rest. Read the plan line: N/M window(s) this run with budget-bounded, resumable means it did not finish. Re-issue the same --from <floor> --to <handoff> with no --budget — an explicit --from re-sweeps its whole range regardless of the resume cursor, and re-scanning covered blocks is free (event writes are conflict-ignored and the cursor only moves forward). The [floor, floor] row a started-but-unfinished pass leaves behind is an ownership marker, not coverage, and the promotion reads it as the nothing it is.

  4. Record the market impairments over the backfilled history, once (attended; archive RPC). Step 3 is what makes this possible and step 2 is what makes it necessary — see the impairment pass below for why the always-on pass will not reach down into history on its own. Run it over the range step 3 just filled:

    bash
    cd /opt/onchain-credit && ETHEREUM_ARCHIVE_RPC_URL=<archive> \
      DATABASE_URL="$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)" \
      npx tsx scripts/ingester/market-impairment.ts --from <floor> --to <live cursor>

    It is idempotent and resumable (the upsert is content-wise, so a re-run of a recorded range touches nothing) and exits non-zero if it refused anything.

A read-only live smoke (scripts/ingester/smoke-ledger.ts) checks the decode path against a live RPC without touching the DB. Phase A ships the ledger only — nothing reads it yet (the 6h cron keeps its own scans), so a missed step degrades nothing user-facing. Log-rotation note: the ingester writes PM2 logs like the app, but the /etc/logrotate.d/onchain-credit glob currently matches only the four creddit-app logs — add creddit-event-ingester-{out,error}.log to it (same copytruncate treatment) when the process is promoted to prod.

The Morpho market-impairment pass, and the one line to grep for ​

Inside the same ~60s loop, after the streams have committed their ranges, the ingester records every Morpho bad-debt liquidation into morpho_market_impairment (one row per socialised loss, market-level, naming no wallet). It reads no chain logs of its own: it walks the morpho stream's already-committed rows, on its own cursor (ledger:morpho-impairment), in windows of INGEST_IMPAIRMENT_WINDOW blocks (default 250,000). Nothing is served from it yet.

It only walks blocks the morpho stream has certified. The stream's event_coverage row is the floor, not the 2025-05-21 history floor, and the difference matters on a fresh box: this pass derives from raw_events, so a block nobody scanned reads exactly like a block with no bad debt in it. Step 2 above starts the ingester before step 3's backfill, and a fresh live stream starts at the tip — so floored at 2025-05-21 the first pass would sweep three million empty blocks, record nothing, and park its cursor at the tip, after which step 3 fills the history underneath a cursor that never looks down again. Hence step 4: the daemon owns the tip, the standalone pass owns the history.

When something cannot be recorded truly, the cursor stops rather than skipping. A refusal means one bad-debt event could not be written honestly — an archive endpoint that will not serve market(id), a loan token whose decimals() cannot be read, or supply shares that moved through something other than a Supply or a Withdraw (the market-fee tripwire). Advancing past it would turn a visible refusal into a permanent silent hole in a record nothing else re-derives, and a missing impairment does not read as missing: it reads as a market that had no bad debt.

There is no exit code for a daemon to alert on, so the log line is the alarm and it repeats every cycle until somebody acts:

bash
pm2 logs creddit-event-ingester --lines 200 --nostream | grep 'morpho-impairment.*HOLD'

The line names the held cursor, how far behind the morpho stream it has fallen, and the refusal's own reason (transaction, log index, market, and which block the read failed at). The hold is also queryable, since a held cursor stops moving:

sql
SELECT scope, last_scanned_block, updated_at
  FROM onchain_credit.chain_scan_cursors
 WHERE scope IN ('ledger:morpho', 'ledger:morpho-impairment');

What to do about one. Most holds are a transient archive failure and clear themselves on the next cycle. A hold that persists is one of three things, in the order worth checking:

  1. The archive endpoint. Re-read the named market(id) at the named block against ETHEREUM_ARCHIVE_RPC_URL. A provider that has stopped serving that depth is the common cause; pointing the ingester at one that does and restarting it clears the hold.
  2. A non-zero market fee, if the refusal names one. That is the tripwire firing for real: AccrueInterest can mint supply shares with no Supply event, so the reconstructed denominator would be short by the fee shares. The refusal is correct and the fix is a code change, not an ops action.
  3. Genuinely unrecordable, if neither holds. The cursor stays put by design.

The stall is about the cursor, not about the record: every run still writes the events it could prove, and the rest of the range can be recorded immediately while the cursor stays below the refusal.

bash
# record everything ABOVE the refused block; leaves the ingester's cursor alone
npx tsx scripts/ingester/market-impairment.ts --from <refused block + 1> --to <live cursor>

Running the standalone pass over a range that contains the refused event refuses it again and exits non-zero, deliberately: the same judgement, from the same inputs. Note also that a held cursor re-reads its whole window from the archive every cycle, so a hold left in place is a standing cost as well as a standing gap.

--reconcile is a repair flag, not a default. It deletes stored rows whose liquidation the canonical chain no longer carries, which is how a reorg is repaired here (the always-on pass does this automatically for the range it re-derives, and the reorg sweep rewinds this cursor with the stream cursors so the repaired range is walked again). Outside the morpho stream's coverage certificate the same predicate would erase true history, so the standalone pass refuses the flag there. The historical pass in step 4 does not want it.

The deploy restarts the ingester and the ledger worker ​

creddit-event-ingester is a separate always-on process that holds its code AND its environment from the moment it was started (run-ingester.sh sources .env.local and execs the loop; the stream registry is additionally cached in-process). So a release that changes scripts/ingester/ingest-events.ts, src/lib/portfolio/tracked-contracts.ts or anything they import does nothing until the process restarts. The ledger worker (creddit-ledger-worker) is the same kind of process, with the same launcher shape, and it holds the derivation code and the token and fund registries from its own start.

deploy.yml restarts both on every prod deploy, right after the app, the ingester first, and on the rollback path as well (that path rebuilds too), through scripts/ops/restart-ingester.sh. Each process gets the same treatment:

  • running or crashed → restarted;
  • stopped (pm2 stop, i.e. a deliberate quiesce such as the 082/083 migration window) → left stopped, with a ::warning:: on the run, because pm2 restart would start it mid-window. The freshness alarm pages hourly until it is started again (the derive-lag arm for the worker), so a quiesce nobody ended cannot go quiet;
  • pm2 refuses the restart → an ::error:: and a red run, because the app is then serving new code over a process that is not;
  • missing → for the ingester, an ::error:: and a red run, for the reason above. For the worker, a ::warning:: while the deployed code still derives inline (WORKER_OWNS_DERIVATION is false in src/lib/portfolio/derive-jobs.ts), and an ::error:: once it is true. Until that flip nothing a user sees waits on the worker (the page load, the 6h tick and the registration replay derive every wallet inline, and the derive-lag arm pages hourly for a worker that is not there), and a release that introduces the worker deploys before it is started, so a red run there would be noise. After the flip a missing worker means nothing derives. The script reads the constant from the checkout it is deploying, so the flip turns the warning into an error by itself.

The grace period is the process definition's, not the script's. pm2 restart waits for the ingester's cycle or the worker's job in flight only as long as that process's own kill_timeout says, and only the documented start command sets it (--kill-timeout 120000 for both); a restart neither adds nor changes it. The freshness alarm reads both back every hour.

This used to be a by-hand runbook step, and it was missed: from 2026-09-03 to 2026-09-11 prod served two releases over a v0.52.2 ingester, which therefore never requested the share-transfer event stETH and eETH movements are booked from. Every rebasing-token movement was missing from the ledger, a new wallet's history was certified complete without them, and its chart booked a 9 ETH swap as return. The restart is automatic now so that cannot recur through a forgotten step, and the freshness alarm pages for every path that bypasses the workflow.

The restart lands BEFORE the release's gated migration, and that is safe. The ingester isolates every stream and the cycle itself behind fail-soft handlers: new code that meets the old schema logs a per-stream (or per-cycle) error and holds that stream's cursor, so nothing is certified over a range it did not write, and the first cycle after the migration picks up where it stopped. A release whose stream set widens (a new event on a stream) wants exactly this order: the ingester already runs the new filter when the gated coverage reset makes it re-scan the history.

One case stalls every stream, not one. The stream set is loaded at the top of each cycle, outside the per-stream handlers, so a release whose registry read needs the new schema fails the WHOLE cycle ([ingester] cycle error (continuing)) every minute until the migration lands, while the freshness alarm reads ok (the process is fresh and online). Nothing is certified in that state, but nothing is collected either. So after every release deploy, read the log: per-stream scanned lines, and no cycle error. If there is one, apply the release's migration now rather than later in the window.

bash
pm2 logs creddit-event-ingester --lines 50 --nostream | grep -c 'cycle error'   # expect 0

The restart lets the current cycle finish, and that needs a start-time flag. The ingester's signal handler sets a stop flag and returns; the loop finishes the cycle it is in and exits 0. pm2 only waits kill_timeout for that, and its default is 1.6 seconds, so before the flag below every release SIGKILLed a cycle mid-flight. The process is fail-soft about it — a live scan advances its cursor last, so an interrupted one re-scans its range, and the reorg repair is one transaction — but a cycle that finishes needs neither. kill_timeout belongs to the process definition: pm2 restart does not add it to a process created without one, so an ingester started before this release needs the one-time step once.

Restart by hand only when the workflow did not: after editing a variable in .env.local, after a manual deploy or rebuild on the box, when a quiesce ends (the deploy leaves a stopped process stopped), or when the deploy run says a restart failed. The worker also after editing a registry row it should see (the ledger worker).

bash
pm2 restart creddit-event-ingester --update-env   # PROD ONLY
pm2 logs creddit-event-ingester --lines 50
pm2 restart creddit-ledger-worker --update-env    # PROD ONLY
pm2 logs creddit-ledger-worker --lines 50

Verify from the log and the data, not from the restart's exit code: the expected stream names appear, a chain_scan_ranges row lands within two cycles, and new raw_events rows carry a non-empty block_hash.

The ingester alarm: freshness, feed lag, derive lag and the ledger audit ​

scripts/check-ingester-freshness.ts runs hourly and asks four questions. Any one can page on its own; each prints its findings with its own tag (ingester-freshness, ingester-feed-lag, ledger-derive-lag, ledger-audit) so the page says which.

Arm 1 — freshness. It pages when the ingester started from this checkout is older than the checkout's current build (it started before .next/BUILD_ID was written, more than 15 minutes ago), is not online in pm2, or does not exist. Both deploy paths restart it seconds after their build, so the grace only covers that gap. Every page leads with its remedy, which for all but the missing process is the restart above. A by-hand rebuild of unchanged code pages too: the rule is that every rebuild restarts the ingester.

It also reads pm2's kill_timeout back every hour and prints a non-paging note when it is not the documented 120,000 ms — which is how the one-time step is confirmed from the log rather than over SSH, and how a pm2 resurrect from a stale dump that quietly returns the grace to 1.6s becomes visible. A note rather than a page because nothing is being lost while it is wrong: an interrupted cycle re-scans its range and the reorg repair is a transaction either way. What is wrong is the claim.

It reads the ledger worker's kill_timeout back the same way, from the same process list, and prints a note tagged ledger-derive-lag when an online worker of this checkout does not carry the documented 120,000 ms: every deploy restarts the worker too, and a job cut short by pm2's 1.6s default is requeued at the next start with an attempt spent. The worker needs no one-time step (it is created with the flag), so the note means it was re-created without it; the remedy is to re-create it with the documented command (the ledger worker).

The ingester is placed in a checkout by the script pm2 runs (pm_exec_path), not by the directory it was started from, so a process re-created from /root is still judged. A run from a checkout the ingester does not run (staging's) reports not evaluated and exits 0. A host with no pm2 daemon pages (the check does not call pm2 there, because pm2 jlist would start one), and so does a pm2 read that fails. A stopped ingester pages hourly until it is started, a deliberate quiesce included. The ok line carries pm2's restart count: a crash loop slower than pm2's min_uptime reads online and fresh, and the climbing count is its only sign. The job never prints the process list: pm2 jlist carries every process's environment.

Arm 2 — feed lag. A fresh, online process says nothing about its feeds: every stream is isolated behind its own fail-soft handler, so one can throw its scan away every cycle for days while the other twelve advance and pm2 stays green. The first symptom used to be a wallet stuck in the building state or a chart that stopped advancing. So the check also compares each followed stream's ledger:<stream> cursor against the chain head, reads the gap as elapsed time at 12s per block, and pages when any feed is more than 6 hours behind (INGEST_FEED_LAG_MAX_HOURS), or when a rolled-out stream has no cursor at all — one the derivation may rely on that no cycle has ever scanned. One [fail] line per lagging feed, worst last so it survives the page's eight-line cut, then an aggregate line that says whether it is one feed or all of them. A feed that is behind catches up on its own once its cause clears; the remedy on the page is to read the ingester's log for that stream, not to requeue anything.

The stream set is built from the ingester's own facts, never from a list in the alarm: a stream is watched when it is rolled out (the '*' markers in event_coverage, plus the six required by definition) or when it simply has a cursor, which is what a stream the ingester is following always has. So a stream added to the ledger is watched from the release that starts following it, and a rolled-out stream with no cursor is the "never scanned" finding. The ledger:morpho-impairment cursor is deliberately not in the set: it is a derived writer rather than a stream, and it holds below a refused block on purpose (its own pass prints a HOLD line every cycle instead).

This arm runs only behind a freshness arm that found nothing, and both halves of that matter. There must be an online ingester of this checkout — a feed can only be behind relative to something meant to be following it. And arm 1 must not be paging: the alert keeps the newest eight matching lines, so feed lines printed beside a freshness page push that page's line, the only one carrying pm2 restart, out of the alert entirely. Being online is not enough to rule that out, because an ingester running pre-build code is online and paging — and the two are correlated rather than independent, since an ingester on old code is exactly one that has stopped following a stream the release added. Nothing is lost: the restart clears the freshness page and the next hourly tick measures the feeds. If the ledger or the head cannot be read, that is a [fail] too (exit 1), never a quiet pass.

Its crontab line is a server step.

Run it by hand once BEFORE installing the line, and settle whatever the feed arm says. The hourly schedule has no throttle and no dedupe, so a standing finding is a page an hour until somebody acts — and the feed arm's first tick reports state that has been accumulating since long before it existed. The shape to expect is a rollout marker that outlived its population: wallet-token's scan surface is the enrolled-wallet list, and with that list empty the ingester emits no wallet-token stream at all, so its cursor stops advancing while its '*' coverage row still says the stream is live. The feed lag then grows without bound and no action the page names clears it — nothing is broken, nothing to restart, nothing to requeue, and widening INGEST_FEED_LAG_MAX_HOURS only postpones it. That is a one-time decision (retire the marker, or re-enrol), and it belongs before the schedule starts sending it rather than after. An alarm nobody believes is a slower version of an alarm nobody has.

Arm 3 — derive lag. The ledger worker turns ingested events into the ledger the pages read, and it is a pm2 process whose log no cron greps, so this arm is its pager. It reads the worker's pm2 entry and the ingester's uptime from the list arm 1 already parsed, and the database: the worker's heartbeat (portfolio:worker:heartbeat, stamped every 30 s whatever the worker is doing), the continuous producer's cursor row (portfolio:worker:continuous, stamped by every producer run that completes) and the ingested tip it follows, the longest-running job, the oldest queued job (a sweep's apart, an audit left out), and the failed jobs. It pages, in this order, when:

  • migration 115's queue does not exist (the release's migration was not applied before its deploy);
  • the worker runs old code: an online creddit-ledger-worker of this checkout that started before the checkout's build, past arm 1's 15-minute grace — the same rule as the ingester's, for the same reason (it would keep deriving every dirty wallet with the previous release's code, and its merges would delete, over each job's range, rows that code does not produce). Restart it;
  • the worker has never written a heartbeat (it was never started here: the page carries the start command), or the heartbeat is more than 10 minutes old (it stopped: restart it);
  • a job has been running for more than 30 minutes (a tick-range job: continuous, sync, sweep) or 60 minutes (a whole-window job: enrol, rederive) while the heartbeat is fresh: the worker is wedged on it (the heartbeat is a timer, so a hung job keeps it beating). Restart it; the job is requeued at the next start, having spent an attempt. A whole-window job is a registration replay's worth of work, and gets the registration drain's own reclaim bound (its child runs up to 30 minutes, reclaimed past 60), so a whale's legitimate job does not page and does not have a restart spend its attempts;
  • the continuous producer has stalled: its cursor row has not been stamped for more than 60 minutes while the ingester has been up at least that long, or it has never been written (the producer has never completed a run here). The producer runs at the end of every ingester cycle, so this is sixty cycles in a row that failed, behind a fresh worker heartbeat and an empty queue: continuous derivation has stopped (the 6h tick still derives every wallet, so it is freshness, not correctness). The page names how far behind the ingested tip it stands; read the ingester log for [ingester/derive-jobs] (a statement timeout, a population read that errors), and restart the ingester if its cycle itself is stuck;
  • any job is failed (a range some wallet's ledger lacks until somebody requeues or deletes it: the page names the oldest; a wallet's continuous failures fold into one standing row, and the worker closes it by itself once the ledger has caught up over its range, see the ledger worker);
  • the oldest queued job that is neither a sweep nor an audit is more than 30 minutes old (alive, and not keeping up), or the oldest queued sweep job is more than 6 hours old (one tick's period: the sweep is not draining and the tick window has stopped advancing). An audit job's age is not judged: it goes back to the queue, its attempt handed back, for as long as it waits (for its wallet's derivation, for the audit of an earlier reading of the wallet, for the write lock, or past a fence), so its age measures what it waits for rather than the worker's pace, and no wait outlives its six-hour horizon, past which it ends partial by itself (arm 4, below).

The bounds are LEDGER_WORKER_HEARTBEAT_MAX_MINUTES, LEDGER_DERIVE_RUNNING_MAX_MINUTES, LEDGER_DERIVE_WHOLE_RUNNING_MAX_MINUTES, LEDGER_PRODUCER_MAX_MINUTES, LEDGER_DERIVE_QUEUE_MAX_MINUTES and LEDGER_DERIVE_SWEEP_MAX_MINUTES. It prints exactly one line, after the feed arm's, and when several things are wrong the line names the one whose remedy clears the most (a restart clears a stale worker, a stopped one and a wedged one; a dead worker explains an old queue) and counts the rest: fourteen feed lines and a derive page still leave the feed aggregate, the worst feed and the derive remedy in the eight the page keeps. It runs behind the same gate as the feed arm — a freshness arm that found nothing and an online ingester of this checkout — because the worker runs from the checkout that runs the ingester: staging, which runs neither, is not judged. An unreadable queue is a [fail] too (exit 1).

Arm 4 — the ledger audit. Every stored reading is checked against the movement records by an audit job in the same worker (the reading audit), and what it books is durable: an adjustment row where a difference was left over after the automatic explanation, and the readings themselves. This arm reads those records since its last look (its cursor, portfolio:worker:audit-arm, starts one hour back on the first run), up to five minutes before it runs, so a correction whose write was still committing when it looked is the next look's rather than lost (PR #961 final review, SF-1), and so is a reading whose 6h commit queued behind the portfolio write lock across a look: both are stamped when written, after the lock (round 2, SF-1). It pages when:

  • an audit booked an unexplained correction (explain_status = unexplained): a holding moved by a quantity nothing on chain accounted for, stated in its asset's own units and in dollars where its book is the dollar (for a correction kept where its reading is gone, from the copy of that reading it carries). The remedy is to explain it (extend the decoder and re-derive the range: the re-derivation queues the re-audit that replaces the correction by the proper record) or to accept its cause in the accepted-cause registry by PR. A correction the audit MOVED from a superseded live tip to the checkpoint that replaced it (a recomposed checkpoint included), with its size unchanged, is not paged a second time once this arm has paged it; one it never saw pages at its new home. One the checkpoint cannot take (it did not read that holding, or the holding moved in between) stays at the tip's block and pages nothing new; the audit removes it once the movement records hold what it stated. A correction booked at a checkpoint the 6h tick committed below a tip already audited replaces the tip's (the audit compares the readings above it again, and the tip's own row where an empty reading retired the tip, on every run, so an interrupted one still does on its retry) and may page once more; it is one correction;
  • a venue whose unexplained corrections were zero for the trailing 7 days books its first new one (a clean record broken, named apart so it is not buried in a count);
  • a leg the movement records hold went unread at two readings in a row (a reader failure, or a closing movement the ledger never recorded: the page says which to suspect when the leg's newest record leaves it open). The audit reads such a leg on its own first, strictly (PR #961 final review, SF-3 and round 2 B1: one call for that leg's own count): where the chain returns zero, the exit is booked as an unexplained correction (above) and the leg is no longer held, so what pages here is a leg that read found held (the reader missed it) or could not read: a revert, a leg outside the registry the worker loaded, a degraded registry load, or a Fluid position, which has no such read yet. The worker's log line for the reading names which (read directly: held (…) or read directly: unreadable (…)). A smart pool's member (or an escrow claim) worth under a cent when last read does not page when a reading omits it: the Fluid reader drops a member that rounds to nothing, and the log line says dust the reader drops, not a reader failure.

An accepted correction never pages; the ok line counts them. It prints exactly one line, after the derive arm's, naming the first finding and counting the rest; the per-finding detail is in the worker's log. It runs behind the same gate as arms 2 and 3, and stores its cursor only after its line is printed, so a look that could not print is repeated rather than skipped. An unreadable audit is a [fail] too (exit 1).

An audit job that ended failed (three attempts that threw: an explanation request the archive endpoint kept refusing, a statement that kept failing) is never closed by the worker — unlike a failed continuous job, no later job derives through it — so arm 3 pages it hourly until it is dealt with. Read its error, fix the cause, and either requeue it (UPDATE … SET state = 'queued', attempts = 0, error = NULL WHERE id = <id>) or, where the next reading's audit has compared the same balances since, close it (SET state = 'partial', finished_at = now(), error = 'closed by hand: superseded by the audit of block <b>'). Requeued, it runs before the audits of the wallet's later readings, which wait for it: a wallet's audits apply in block order. Requeue rather than close one whose wallet's audited-through watermark lies above its to_block (a checkpoint the 6h tick committed below a tip already audited): its comparison of what lies above it is what takes the tip's correction away, and for a wallet that holds nothing no later audit reaches that row until the wallet is re-derived. An audit whose own write, or whose in-place re-derivation, only queued for the write lock, or met a fence, goes back to the queue without spending an attempt, as the derivation itself does, so contention never parks one failed. No audit waits past its six-hour horizon: a deferred one, one waiting behind the audit of an earlier reading of its wallet, and one still meeting the lock or a fence all end partial by themselves then and page nothing (the next reading's audit compares the same balances), so the wallet's later audits stop waiting behind them.

The wrapper sends every line to the logfile, so read it back from there. The file is keyed by checkout, so this is prod's own history; the env= anchor is kept because a line is read far more often out of a page than out of the file it came from:

bash
cd /opt/onchain-credit && ./scripts/run-cron.sh check-ingester-freshness.ts
grep 'env=onchain-credit$' /tmp/onchain-credit-cron/onchain-credit-check-ingester-freshness.log | tail -5

Then install it:

bash
5 * * * * /opt/onchain-credit/scripts/run-cron.sh check-ingester-freshness.ts

Verify the install two ways: the count proves the line is on this checkout, the log proves it runs. Both the count and the log path name the checkout, so neither can be satisfied by the other environment's ticks.

bash
crontab -l | grep -c '/opt/onchain-credit/scripts/run-cron.sh check-ingester-freshness.ts'
grep 'env=onchain-credit$' /tmp/onchain-credit-cron/onchain-credit-check-ingester-freshness.log | tail -3

A healthy run is four lines, one per arm, plus a note for each process whose grace period is not the documented one: the ingester's while it is still pm2's default, which the one-time step clears, and the worker's if it was re-created without the flag (re-create it). Here, with the ingester's note:

ingester-freshness: ok: started 2026-09-11 11:37Z, after the build of 2026-09-11 07:22Z, v0.59.0, 9 restarts env=onchain-credit
ingester-freshness: grace period is unset (pm2 default 1600)ms, not the documented 120000ms, so a restart can cut a cycle short. Take the one-time step: docs/ops/release-steps.md, ingester kill timeout env=onchain-credit
ingester-feed-lag: feeds ok: 3 followed feed(s) within the 6h bound, worst escrow 0h30m behind head 24601206 env=onchain-credit
ledger-derive-lag: ok: heartbeat 0h00m ago, derived through block 24601140; producer 0h01m behind the tip; 0 queued (oldest none), 0 running, 0 failed, 0 partial in 24h env=onchain-credit
ledger-audit: ok: no unexplained adjustment, no leg unread twice env=onchain-credit

The ledger worker ​

creddit-ledger-worker (scripts/worker/ledger-worker.ts) is the one process that drains the derivation queue: each job derives one wallet's ledger over one block range through the same writer the 6h tick uses (data-pipeline). Like the ingester it is not a cron: an always-on pm2 process with its own loop, started from the prod checkout only. Staging runs no worker, as it runs no ingester.

Start it (the migration that creates its queue, 115, must be applied first; a worker that meets no queue logs a failed claim every 30 s and the alarm pages for the missing table):

bash
cd /opt/onchain-credit
pm2 start scripts/worker/run-worker.sh --name creddit-ledger-worker --time --kill-timeout 120000
pm2 save                              # persist across reboots (and the flags above)
pm2 logs creddit-ledger-worker

run-worker.sh sources .env.local and execs the loop, so pm2's stop signal reaches the worker's own handler: it sets a stop flag, lets the job in flight finish and exits 0. The --kill-timeout is what makes that true, exactly as for the ingester: pm2's default of 1.6 s kills a merge mid-flight, and the flag belongs to the process definition, so pm2 restart does not add it to a process created without one. A job cut short anyway (a whale's whole-window enrolment can outlast 120 s) stays running; the next start requeues it, having spent one of its three attempts.

One worker at a time, enforced. Before it claims anything the worker takes a session advisory lock (onchain_credit.ledger_worker) on a connection of its own and holds it for its whole life. A second instance does not drain beside it: a pm2 start logs another ledger worker holds the worker lock; waiting for it and waits; a hand-run refuses and exits 1. If the lock's connection fails (a database restart), the worker finishes the job in flight, claims nothing more and pauses its heartbeat until it has taken the lock again on a new connection — it does not exit, so a database outage cannot use up pm2's restart budget and leave no worker at all. A lock that stays lost pages through the stale heartbeat. Every verdict is also fenced on the claim that produced it, so even a job two processes somehow both ran is recorded only by the claim that owns it.

The deploy restarts it, beside the ingester and in the same way (the deploy restarts both): it holds its code, and the token and fund registries, from the moment it started. (The audit's own chain reads are the exception: its explanation and its strict read of a missed leg reload the registries every minute, and a leg the loaded registry lacks, or a degraded load, is read as unreadable, never as a closed position: PR #961 final review round 2, B1.) Restart it by hand after editing a registry row it should see, after an .env.local edit or a manual rebuild on the box, and when a quiesce ends. A worker left on the old table is fail-safe rather than wrong (PR #961 final review, N5): a range that moves a token the table it loaded does not know is refused by the wallet-token universe guard, the job ends partial and the wallet's stamp holds, but nothing derives that token until the restart:

bash
pm2 restart creddit-ledger-worker --update-env   # PROD ONLY
pm2 logs creddit-ledger-worker --lines 50

Stop it with pm2 stop creddit-ledger-worker: it finishes the job in flight and exits (stopping it is safe, below, says what keeps deriving meanwhile), and pm2 save keeps it stopped across a reboot. A stopped worker stays stopped: the deploy leaves it alone with a ::warning::, and the derive-lag arm pages hourly until pm2 restart creddit-ledger-worker --update-env starts it again. The user wipe refuses to run while it is not stopped.

Re-create it when the hourly alarm notes that its grace period is not 120,000 ms (a pm2 resurrect from a stale dump, a pm2 upgrade, a re-create that forgot the flag). The flag belongs to the process definition, so only a new definition carries it, and deleting the worker loses nothing:

bash
cd /opt/onchain-credit
pm2 delete creddit-ledger-worker
pm2 start scripts/worker/run-worker.sh --name creddit-ledger-worker --time --kill-timeout 120000
pm2 save

What healthy looks like. One line per job, [worker] job #<id> <kind> <wallet> [<from>,<to>] (derived [<floor>,<to>]) attempt <n> -> <state>, <written> written, …, mostly continuous jobs a few seconds after each ingester cycle that saw a tracked wallet move, and silence otherwise. The derived range is where the job actually started: its wallet's tick floor (the tick cursor + 1, or the wallet's pending marker when that is lower), so it reaches back up to one 6h tick window behind the job's own range — that is by design (a job anchors on the ledger below where it starts). An occasional -> queued (attempt not counted: another writer moved its wallet's ledger while it derived, so it merged nothing) is healthy too: a whole-window merge (a registration's ledger half, the re-derivation trigger, a history build) or the 6h tick rewrote that wallet while the job derived, so the job merged nothing and runs again a minute later on the new ledger (the stale-merge fence). So is an audit job's -> queued (attempt not counted: the audit waited or was fenced, and its reason says which), while it lasts: the reason after the colon says whether it waits for its wallet's derivation (deferred:), for the audit of an earlier reading of the wallet (waiting:), or for the write lock, or whether another writer moved the ledger under its write (fenced:); none spends an attempt, and none outlives the audit's six-hour horizon (the ledger audit arm). The heartbeat row advances every 30 s, and the queue is short:

sql
-- the worker is alive: the heartbeat is seconds old
SELECT last_scanned_block, now() - updated_at AS age FROM onchain_credit.chain_scan_cursors
 WHERE chain_id = 1 AND scope = 'portfolio:worker:heartbeat';
-- the queue by kind and state
SELECT kind, state, count(*), min(created_at) FROM onchain_credit.portfolio_derive_jobs
 WHERE chain_id = 1 GROUP BY 1, 2 ORDER BY 1, 2;

A failed job is a wallet and a range the worker tried three times and could not derive, and the hourly derive-lag arm pages while it stands. Its reason is on the row and on the [fail] ledger-worker: line in the log. (A job that only queued for the write lock past its budget is not a failure, and neither is one whose merge the stale-merge fence stopped because another writer moved its wallet's ledger while it derived: each goes back to the queue without spending an attempt, and a queue that stays busy, or a wallet whose ledger keeps moving under every attempt, pages as an old queued job instead. A history build rewrites the wallets it builds, so their jobs can be fenced while it runs; the worker does not need stopping for one.) A wallet's failed continuous jobs fold into ONE standing row: a later failure is folded into it. (The one exception: a job the worker parks at its own start, having found it running with its attempts already spent, is parked without the fold, so it can stand beside the wallet's row; the close below takes both.) It closes by itself once the ledger has caught up over its range — the tick cursor passes its top, or a later continuous job for the wallet finishes done through it — with a [worker] closed … failed continuous job(s) line; the row turns partial and keeps what it failed on in error. The ingester keeps enqueueing the wallet meanwhile, and the 6h tick keeps deriving it, so a fault that has cleared stops paging at the wallet's next movement or the next tick. One that PERSISTS keeps paging (the tick pages for the same wallet through its own exit code, too): fix the cause (an archive provider, a registry row). A failed job of any other kind stays until a person acts: requeue it once its cause is fixed, and delete it only when the range no longer matters (the wallet was removed, or a later tick has derived it):

sql
SELECT id, kind, wallet, from_block, to_block, attempts, error, finished_at
  FROM onchain_credit.portfolio_derive_jobs WHERE chain_id = 1 AND state = 'failed' ORDER BY finished_at;
UPDATE onchain_credit.portfolio_derive_jobs
   SET state = 'queued', attempts = 0, claimed_at = NULL, finished_at = NULL, error = NULL
 WHERE chain_id = 1 AND id = <id> AND state = 'failed';

Enqueueing by hand. An operator may insert two kinds, and the table fixes each kind's priority:

  • continuous: one wallet over a range, with the settle line -1 if in doubt (nothing it writes is labelled settled that the next pass does not re-derive). It starts at its wallet's tick floor like any other (or at its own from_block, when that is lower): a range above the tick cursor cannot be derived alone, because its running balances are anchored on the ledger below it.
  • rederive: one wallet's WHOLE window, re-derived and re-certified (no settle line). It runs the re-derivation trigger's three launch guards, so it can certify nothing under conditions the trigger would not (data-pipeline). A wallet whose registration replay is parked or still pending (portfolio_backfill_state.statuserror, queued or running) is held: the job ends partial having derived nothing, and the wallet stays excluded until that replay completes (re-queue a parked one; its own ledger half derives and stamps the wallet). The job opens at the wallet's own history floor and certifies from there, so from_block must be that floor: a job naming any other block derives nothing and ends partial with the floor in its reason. And it never stamps below the recorded arm block (portfolio:v2-campaign:arm-block), where a derive cursor reads as a short build for good and the trigger stops coming back for the wallet: while no arm block is recorded, or the ingested tip is still below it, the job is held the same way (the trigger declines then too, and derives a cursor-less wallet by itself once the tip passes the arm; enqueue anything else again then), and a to_block below the arm derives nothing — so <to> is the chain head: any explorer's latest block, or on the box the newest block the ingester has scanned, SELECT max(last_scanned_block) FROM onchain_credit.chain_scan_cursors WHERE chain_id = 1 AND scope LIKE 'ledger:%' (the writer stops at the ingested tip either way). The INSERT below reads the floor exactly as the worker does (no stored pair is the indexed floor, 24136053; a stored one is never below it). This is also how the comparator re-derives a wallet (plan §7.3). A sub-range is a continuous job's.

Never insert enrol or sync by hand: they are the registration replay's (after its own gate) and the page load's, and each, run as its kind, trusts its producer for a guard a typed row does not carry (a sync the page load's floor, an enrol the replay's gate). In this release nothing produces either kind, so the worker refuses one: partial, nothing derived, the reason naming this section. A sweep job enqueued by hand in this release is derived like a continuous one, as a plain range, and moves nothing: the tick cursor and the pending markers belong to the inline 6h tick until WORKER_OWNS_DERIVATION flips.

sql
-- one wallet over a range
INSERT INTO onchain_credit.portfolio_derive_jobs (chain_id, wallet, from_block, to_block, priority, kind, settle_line)
VALUES (1, '<lower-cased wallet>', <from>, <to>, 20, 'continuous', -1);
-- one wallet's whole window: from its own history floor through <to> (the writer stops at the
-- ingested tip, so the chain head is the right <to>; a <to> below the arm block derives nothing)
INSERT INTO onchain_credit.portfolio_derive_jobs (chain_id, wallet, from_block, to_block, priority, kind)
SELECT 1, a.uid,
       CASE WHEN a.history_floor_ts IS NULL OR a.history_floor_block IS NULL THEN 24136053
            ELSE GREATEST(a.history_floor_block, 24136053) END,
       <to>, 30, 'rederive'
  FROM onchain_credit.accounts a
 WHERE a.uid = '<lower-cased wallet>';

A bounded hand-run (a check after a manual step, never a replacement for the process): LEDGER_WORKER_EXIT_WHEN_IDLE=1 drains what is claimable and exits. It refuses while the pm2 worker runs, so stop that first and start it again after:

bash
pm2 stop creddit-ledger-worker
cd /opt/onchain-credit && LEDGER_WORKER_EXIT_WHEN_IDLE=1 bash scripts/worker/run-worker.sh
pm2 start creddit-ledger-worker

Stopping it is safe. Nothing is lost: the ingester keeps enqueueing, and the 6h tick, the registration replay and the page load keep deriving inline for as long as WORKER_OWNS_DERIVATION is false (this release). What stops is the one-cycle freshness; the derive-lag arm pages hourly until the worker is back. The backlog is one job per wallet (the producer widens a wallet's queued job rather than adding another), and it is harmless when the worker returns, however many 6h ticks it waited across: each job starts at its wallet's tick floor as it stands when the job RUNS, so it anchors on the ledger the tick has derived since; it reads the tick's pending markers and leaves them to the tick; and it cannot relabel a row the tick settled.

The health check, after a deploy and whenever the queue looks wrong. Four numbers, each from what the pieces already log (the ledger-first release reads them on the worker's first day):

bash
# 1. the page load's skipped merges (the worker now holds the write lock for every merge it runs;
#    a page load only TRIES it, and a skip means its reading was not stored). Compare with the
#    same window before the deploy; a handful a day is the expected order.
grep -c 'writer lock busy; v2 persist skipped' /root/.pm2/logs/onchain-credit-out.log
# 2. the worker's verdicts: mostly `-> done`; no `-> failed`. A `-> queued` is a retry; the ones the
#    stale-merge fence made say `another writer moved its wallet's ledger` (count them separately)
pm2 logs creddit-ledger-worker --lines 2000 --nostream | grep -oE -- '-> (done|partial|queued|failed)' | sort | uniq -c
pm2 logs creddit-ledger-worker --lines 2000 --nostream | grep -c "another writer moved its wallet's ledger"
# 3. the continuous producer: a `[ingester/derive-jobs] … continuous job(s)` line on the cycles that
#    saw a tracked wallet move, and no `continuous enqueue error` (a run over its 20 s budget, say)
pm2 logs creddit-event-ingester --lines 2000 --nostream | grep -c 'continuous enqueue error'
# 4. the queue: one open continuous job per wallet at most, nothing failed that is not closing
sql
SELECT kind, state, count(*), count(DISTINCT wallet) AS wallets, min(created_at)
  FROM onchain_credit.portfolio_derive_jobs WHERE chain_id = 1 GROUP BY 1, 2 ORDER BY 1, 2;

A page-load skip count that climbs with the worker's job count, rather than staying rare, is the contention the flip of WORKER_OWNS_DERIVATION removes (the page load becomes a sync job that queues for the lock); until then it costs those refreshes their stored reading, never a derived row (the skip marks the wallet pending and the next tick derives it).

The escrow registry check ​

The escrow stream scans a set of queue and manager addresses pinned in code. They can move — a pool manager is upgradeable, a queue can be redeployed — and if one does, the stream keeps scanning an address that emits nothing while the certificate for those blocks still says they were covered. Re-read them from the chain before a release that touches the registry, and after any Ethena, Lido, EtherFi or Maple upgrade:

bash
cd /opt/onchain-credit && ETHEREUM_RPC_URL="$(grep ^ETHEREUM_RPC_URL= .env.local | cut -d= -f2-)" \
  npm run verify:escrow-registry

Exit 0 is clean; 1 is drift (a pinned address, or a cooldown duration, disagrees with the chain — fix the registry in the same release, and re-check whether the class still escrows at all); 2 means the check could not run (no readable endpoint), which is not the same thing and must not be read as a pass. It is one multicall round, so it costs one request whatever the registry grows to.

The ingester runs the same check itself, before it scans the escrow stream, and holds that stream's cursor on a disagreement it can read — so a drift that slips past this step stops the escrow stream rather than corrupting its certificates, visible as a repeating [ingester/escrow] registry drift line. A clean pass is cached for the life of the process; an unreadable one is retried on a 10-minute interval rather than every cycle, so a provider having a bad ten minutes cannot slow the loop's cadence for every other stream.

The ledger's writers ​

The ledger had one offline builder and three live writers, and the writers used to be off by default — turned on before the build, so the two overlapped rather than leaving a gap between the block the build stopped at and the block the writers started from.

None of that is a choice any more. Three environment variables governed it and all three are gone: PORTFOLIO_LEDGER_WRITE (whose default was write nothing), PORTFOLIO_LEDGER_MODE (which chose between the old chain scan and the ledger), and PORTFOLIO_LEDGER_SOURCE (which chose which ledger the reader served). There is one ledger, one writer set and one reader. Delete all three lines from .env.local when you next edit it; they are inert either way, and the runbook for the release that retired the last of them is the release-steps log.

Arming was an edit plus a restart — no rebuild, no migration — and it is now nothing at all: the writers run unconditionally, so a fresh environment needs no arm step. What a fresh environment does still need is the arm block, which is the campaign's upper bound, what complete is measured against, and the last resort of the tick window's bootstrap chain (below):

sql
INSERT INTO onchain_credit.chain_scan_cursors (chain_id, scope, last_scanned_block, updated_at)
VALUES (1, 'portfolio:v2-campaign:arm-block', <the block at the restart>, now())
ON CONFLICT (chain_id, scope) DO UPDATE
  SET last_scanned_block = EXCLUDED.last_scanned_block, updated_at = now();

Confirm the writers are running, on all of them: one 6h tick logs a v2 dual-write: line naming the wallets and the range, the drain logs a [rederive] candidates=… line (on a half-hourly heartbeat when there is nothing to re-derive, so allow thirty minutes rather than one), and one page load leaves no [v2-partial] … skipped behind.

All five writers run the derivation composition: the 6h tick, the page-load persist, the registration replay, the re-derivation trigger, and the offline history builder.

The refusal is per wallet and per range, and that distinction is what an operator reads in the log. A wallet is refused, with its reason named on the line, when this run cannot honestly derive it:

  • it holds a position at a venue the run's adapter set did not cover — a stated refusal rather than a partial write, because writing the covered venues would let the range be certified as derived while one venue was never read;
  • it holds a position the run covered the venue for and still cannot name — a vault, a reserve, a market or a PT the venue's registry does not list. The line names the position. This is the same rule as the one above, applied at the unit coverage is actually measured in; without it the wallet would be reported as fully derived over a position nothing looked for. No tracked wallet trips it today — replayed against production on 2026-08-28, all of them resolve — so a line of this kind means a registry lost an entry, and the entry it lost is on the line;
  • a venue registry read failed, so the run cannot state what it covered;
  • a liquidation landed in the range on a vault whose position states could not be read. That receipt has no event of its own — the change in position across the block is the receipt — so a run that cannot read the state cannot see the liquidation at all.

A refused wallet writes nothing, stamps nothing and is retried by the next pass, which is the same behaviour the whole-file refusal had; what changed is that the wallets a run can serve are now served. Seeing a handful of refusal lines beside successful wallets is the expected shape. Seeing every wallet refused for the same reason is a genuine fault, and the reason on the line names it.

The composition prices every row it writes. Each row is valued at its own block on the three stored lines: the market mark in the position's book, the redemption mark in the same book, and the dollar figure that lets capital net across books. The marks come from the same price mirror, the same wrapper composition and the same block-anchored redemption rates the existing flow ledger is valued with, so a movement cannot be worth one thing on the old ledger and another on the rebuilt one.

A mark that cannot be resolved stays empty and says so. It is never zero and never a guess: a token whose decimals cannot be read, a price bar the mirror does not hold, a young Pendle market whose oracle reverts — each leaves that row's market figure empty and raises a mark-unresolved anomaly naming the asset and the block, and the resolver logs one line per asset and per empty figure, carrying how many of that asset's rows came out that way and the lowest block, so a repair can target the asset instead of re-deriving every wallet that touched it. Two figures are empty by design rather than by failure: a holding outside the yield books has no redemption claim, and a Pendle principal token's redemption mark is the pull-to-par curve the read path derives from the position's own entry.

A half-valued row counts as unvalued, and that is what makes the count worth reading. The redemption mark needs no price bar — for a par holding it is the identity — so a row whose bar is missing comes back carrying a redemption figure and no market figure, and looks populated. The reconciliation is computed from the market mark, so such a row is exactly as unusable as an empty one; it is counted as one, and the redemption figure it did resolve is still stored.

The dollar figure is counted the same way, and it can be empty on its own. It is the market figure carried into dollars by the book's own unit, so in the ether book it needs an ETH price at the same instant. A holding of ether itself — or of any asset whose claim is one unit of its own book — is worth a real number of ether with no price lookup at all, and at an instant where the ETH price is missing that row comes back with a book figure and no dollar figure. Cross-book capital netting is computed from the dollar figure, so it withholds a whole transaction on the first such leg. Both figures are counted, each on its own line naming what to load: the asset's own coverage for a missing market figure, the book's unit for a missing dollar one. The redemption figure is the only one of the three with a by-design empty, which is why the count is never taken on it.

How precise those marks are, and the one case where they are less so. Every mark divides two quotes taken from the same bar, so the exchange level cancels and only the slow secondary-market basis survives — that holds whatever the bar's age. What can be stale is the bar itself. The existing flow ledger pays for one Dune execution per run cluster to true its bars up to the exact minutes it values at; the composition deliberately does not, because its anchors are every receipt block in a range rather than a handful of window anchors, and a whole-history build would spend thousands of executions to move a mark by the basis drift of one hour (~1-3bps in calm markets). Inside the 6h tick this costs nothing: the flow ledger runs first, writes the exact bars for the same blocks, and the composition reads at the same timestamps. It shows only where the old ledger never valued a flow at that block, and on a standalone history build, where a mark can sit on a bar up to an hour old.

Where you read the count. The offline history build prints it per range in its anomalies= total. The 6h tick carries it as an N unpriced token on its own summary line ([portfolio] v2 dual-write: … 0 failure(s), 0 partial, 2 unpriced in …), which is now the count of wallets carrying at least one row with an empty money figure — either the market one or the dollar one. Its healthy value is 0; a persistent non-zero names an asset, not a release, and the resolver's own per-asset lines say which asset, which figure is empty, how many of its rows, and the lowest block. The two remedies it points at are getting the missing series covered by the price mirror — the asset's own for an empty market figure, the book's unit for an empty dollar one — and then re-deriving the affected range, in that order, since a re-derivation against a mirror that still cannot price it changes nothing. The first tick after a change that widens what gets valued can legitimately show a non-zero here where the same rows were previously empty in silence; the per-asset lines are what tells the two apart.

THE PRICED REBUILD'S PRECONDITION: load the price mirror back to the ledger's first event before you start. This is a correctness boundary, not an optimisation, and it is the one step that decides whether the rebuild below is worth its wall clock. Rows are valued at their own block; below the mirror's earliest bar for that token there is no bar to value against, so every row of it down there stores a redemption figure and no market figure — populated-looking rows the reconciliation cannot evaluate. The gap is not self-healing: a missing bar falls through to a per-window fetch from the pricing provider, which is credit-guarded per run and, once the guard trips, returns nothing for every remaining block without failing the build.

Coverage is per token, and the number to compare against is the WORST one. The mirror's earliest bar overall is a property of whichever token has been tracked longest, not of the token your rebuild is about to price. On prod those differ by more than a year: the deepest series starts 2025-07 and roughly three-quarters of the tracked tokens have their first bar in 2026. A check that reads the overall earliest bar and calls the range covered is answering the wrong question.

And "which tokens" is the registries, not the log stream. The tokens a rebuild has to price are the ones a row can be denominated in — everything in the tracked-token registry, every lending-reserve underlying, and every Pendle market's underlying. That is deliberately wider than "tokens seen emitting a log under their own address": a log carries the emitting contract, which is the token only on a plain transfer and is the venue on every pooled, vaulted or wrapped stream. Asking the log stream which tokens matter therefore drops exactly the assets a lending ledger is made of — an ETH liquid-staking token held only through a lending venue has no log of its own and plenty of rows. Measured on prod 2026-08-29: 127 tokens in that universe, 71 of which the mirror holds no bar for at all, and 19 already short against a log of their own. A token with no bar anywhere is the worst case there is, not an absent one.

The build prints the answer before it does any work ([shadow-build] price mirror: worst-covered token 0x… has NO BAR AT ALL …), and --dry-run prints it too. That line is the authority: it takes the worst coverage over the tokens this ledger prices from a bar of their own — excluding the wrappers whose mark composes from another token's series, the par-pinned assets that consult no bar at all, and the principal tokens marked off their underlying — it names the tokens already short, and it prints the exact load command for them. To see the list by hand, use the build's own line rather than a hand-written query: the hand-written one is where the "which population" mistake gets made.

Run the command the line prints, verbatim. The tape command below is one example of what the line can say, not its only shape. It names its tokens explicitly, and that matters more than it looks: the loader's default set is its own tracked list of 27 addresses, and on prod 2026-08-29 75 of the 87 tokens this check names are outside it — including the worst-covered one, and 7 of the 19 already short against a log of their own. Running the bare script would exit 0 without touching them and the check would print the identical line forever. The named set also governs the loader's resume gauge (the emptiest named token decides whether the window is already loaded), which is the right yardstick here and the reason --force is rarely needed. One long-window execution, a few credits against a ~4k/month budget:

bash
# exactly as the build printed it — start instant and token list included
/opt/onchain-credit/scripts/run-cron.sh backfill-token-price-bars.ts 2026-01-01T00:00:00.000Z \
  --tokens=0xd3fd63209fa2d55b07a0f6db36c2f43900be3094,0x056b269eb1f75477a8666ae8c7fe01b64dd55ecc,…

Then re-run the check, and read a second identical line as an answer, not a failure. Whether the load actually filled a thin token's history is not knowable from its exit code — it skips a window it judges complete, and whether the provider holds bars that deep for a thinly-traded token is not something the preflight can see in advance. Re-reading the build's own line (--dry-run is enough) is what turns this from a hope into a precondition. A token still short after a load that ran is one the provider has no series that deep for — a matured principal token, a vault share that never traded, a token first quoted last month. That is a coverage note to record (its rows will carry a redemption figure and no market figure, and the reconciliation will exclude them), not a load to re-issue forever. Re-issue only for a token you have reason to believe the provider does cover.

The loader is idempotent and resumable, so re-running it after a partial load is safe. --force disables the resume gauge entirely; with an explicit --tokens set the gauge is already the emptiest named token, so reach for --force only when you have named a token that does have bars in the window and want them re-fetched anyway. For the record on the default (WETH) gauge, that margin is thin and shrinking: measured on prod 2026-08-29, over [2025-05-21, now] the mirror holds 9,649 on-hour WETH bars against 11,158 window hours (0.86, under the 0.9 the yardstick calls complete), and the ratio rises on its own as the mirror ages.

A range derived before the resolver shipped is worth nothing and must be rebuilt. There is no re-mark repair for a range derived unpriced — so the remedy is the ordinary one: re-derive it. That is exactly what the re-derivation trigger is for, and it is why the composition never reads a cursor. Budget one full history rebuild if any range was built unpriced, before C6 — and spend it after the mirror's per-token coverage has been loaded and re-checked, or it buys a ledger that still cannot reach C6 and a second rebuild.

The second precondition is that every declared stream is actually being followed, and it has a visible failure mode rather than a silent one. Each writer stops at the slowest stream's ingestion progress, and a stream that is live in the coverage table with no ingester cursor counts as progress zero — so one such stream holds the whole dual-write at the floor. What you see is a tick that logs "the slowest stream this tick relies on is ingested only to block 0" and derives nothing, on every tick, rather than quietly writing half a history. If that line appears, the fix is the ingester; the ordering the runbook already prescribes (restart the ingester onto the new streams, then backfill each to the floor) is what keeps it from happening.

The rebuilt ledger's own 6h window ​

The 6h tick's rebuilt-ledger pass used to borrow its block window from the old pipeline's flow scans. Those scans are deleted, and the pass keeps a cursor of its own — portfolio:v2:tick in chain_scan_cursors, one row for the whole tracked population — and computes its window from it:

  • the floor is that cursor plus one;
  • the cursor stops at the chain's finalised head, exactly as every old flow scan's cursor did. Everything the tick derived above that line is reorg-exposed, so it sits above the cursor and the next tick re-derives it. Nothing re-derives below a cursor, so a cursor saved at the top of the range would strand that tail;
  • the ceiling is unchanged: the tick's own scan bound, clamped to the slowest ingestion progress across the streams the derivation reads. A tick that cannot reach the ceiling keeps its floor rather than stepping over the gap, which is what the borrowed window could never do;
  • a page load derives from the LOWER of that cursor and its own lookback floor, up to its own anchor — min(last snapshot block + 1, cursor + 1) — so the two passes agree on where the ledger is current and the window can only ever widen relative to what shipped. The min is the safety half rather than decoration: a wallet with no snapshot at all takes a fixed-lookback floor that can sit below the population's cursor, and that wallet was not in the population the cursor speaks for. A page load never advances the cursor, because it derives one wallet.

The first tick after the deploy prints which rule produced its floor, and that is the one line to grep:

bash
grep 'v2 tick window' /tmp/onchain-credit-cron/onchain-credit-refresh-portfolio.log | tail -3
# [portfolio] v2 tick window: floor 23001235 from the lowest per-wallet derive cursor (bootstrap: no v2 tick cursor yet)
# [portfolio] v2 tick cursor -> 23003010 (settle line 23003010)
# ...and every tick after it:
# [portfolio] v2 tick window: floor 23003011 from the portfolio:v2:tick cursor

The bootstrap chain, in order, for a tick that finds no cursor. That includes the first tick of the release that introduced the cursor, not only a restored or rebuilt database: portfolio:v2:tick is written by the tick and by nothing else, so it does not exist anywhere until one has run.

  1. the lowest per-wallet derive cursor over the tick's own population, plus one — the highest block every wallet in it is already derived to. A wallet with no cursor is left out rather than dragging the floor down: nothing of it is served, and its registration replay derives its whole history anyway.
  2. the recorded arm block, plus one: the top of the one-time historical build, so everything above it is what the live writers owe.
  3. nothing. The tick derives no rebuilt-ledger range at all and says so on the same line. Record an arm block (the INSERT under the writers section), or let one wallet's registration replay stamp a derive cursor, and the next tick picks it up.

What rung 1 costs on the first tick, stated so it is not a surprise. A tick never stamps a per-wallet derive cursor (only a whole-history run does), so those cursors sit where the last replay or historical build left them — for most of the population, at the recorded arm block. The first tick therefore derives from the arm block to the ingested tip: on the order of a week of blocks (tens of thousands) for the whole population, against ~1,800 on a steady tick. It is correct and it is bounded — the composition segments its own log load, so this is a handful of segments per wallet rather than one enormous read, and run-cron.sh's flock is non-blocking, so a tick that overruns its slot is skipped rather than stacked. No range is missed either way: every wallet with a cursor was derived at least that deep by its own whole-history run, and the tick covers everything above the minimum for the whole population.

What it does do is re-derive a week of already-written rows and restate their marks — a few basis points on value_market (the accepted same-bar drift), more on the unserved value_usd (see the true-up note).

Optional insurance: seed the cursor before the first tick, and before migration 095. It skips the wide first tick entirely by handing the new cursor the window the deleted rung would have taken, read from the old flow scopes' own cursor rows. Those rows exist only until migration 095 runs: its second half deletes exactly the seven scopes the seed reads from, so after 095 the INSERT below is a silent no-op (HAVING count(*) > 0) and the first tick falls back to the bootstrap window, which is safe but wide. It needs no migration (so it is unaffected by prod migrations being a manual step) and it is a no-op once a tick has run:

sql
INSERT INTO onchain_credit.chain_scan_cursors (chain_id, scope, last_scanned_block, updated_at)
SELECT 1, 'portfolio:v2:tick', min(last_scanned_block), now()
  FROM onchain_credit.chain_scan_cursors
 WHERE chain_id = 1 AND scope IN ('portfolio:transfers','portfolio:transfers:wallet',
       'portfolio:events:morpho-blue','portfolio:events:aave','portfolio:events:sparklend',
       'portfolio:events:fluid-operate','portfolio:transfers:fluid-nft')
HAVING count(*) > 0
ON CONFLICT (chain_id, scope) DO NOTHING;

HAVING count(*) > 0 is what makes it a no-op rather than an error on a box where those scopes are absent — a reseeded staging, or any environment that never ran the old scans. Without it, min() over zero rows returns one row of NULL and last_scanned_block is NOT NULL, so the statement fails with a constraint violation. It is fail-closed either way (nothing is written), but a constraint violation reads as a broken runbook rather than as "you do not need this step". Verified both ways: INSERT 0 0 when the scopes are absent, the correct min when they are present, and a second run inserts nothing.

Take it or leave it deliberately, and grep the v2 tick window line either way: seeded, the first tick reads from the portfolio:v2:tick cursor and covers ~1,800 blocks; unseeded, it reads from the lowest per-wallet derive cursor and covers the week described above. Both are safe; the seed only buys the wall clock and the restatement.

The acceptance reconciler ​

scripts/ops/reconcile-ledger.ts answers one question: does every figure the rebuilt engine books follow from the stored rows alone? It recomputes each figure a second time, from the stored snapshots and the stored movements, with arithmetic that shares no code with the engine, and reports every disagreement. The same job also runs on a schedule, so a disagreement that appears later is not silent.

It is read-only by default — every statement is a SELECT — so it is safe to run at any time, against any environment:

bash
cd /opt/onchain-credit && DATABASE_URL="$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)" \
  npx tsx scripts/ops/reconcile-ledger.ts
# one wallet, one block range, and the machine-readable form
npx tsx scripts/ops/reconcile-ledger.ts --wallet 0x… --from 24136053 --to 24200000 --json

--from and --to restrict what is reported, never what is computed: the engine always runs over the wallet's whole history, because a leg's figures depend on where its position opened. Both ends are required — a half-specified range would read as a full run. The range is echoed on the report for the same reason.

Reading the report. A clean run is quiet. The one line it always prints is the coverage line:

ledger-reconcile: wallets=3 reconciled=3 excluded=0 uncertified=0 checked=4128 book-checked=1032 withholds=2 notes=3 env=onchain-credit
  • excluded — wallets the reconciler did not check, because their rebuilt history is not yet provably complete. Checking a half-built history against a full curve reports a difference that is not a defect, so excluding them is right — but the figure is printed on every run, at zero as well, and each excluded wallet is named with which of three cases it is, because the remedies differ: no cursor (the wallet's replay is in flight, never ran, or ran and correctly declined to stamp), cursor short (the history build stopped part way — resume it), and coverage not certified (that wallet's own enrolment backfill has not reached the floor yet — wait, and read its progress). An alarm whose population silently shrinks reports green forever, which is the failure this whole programme exists to remove.
  • uncertified — a second, independent count of wallets whose required coverage does not reach the floor, read from the coverage relation with the derive cursor never consulted. Both sides of the reconciliation read the same stored rows, so a row that was never written is invisible to it; this figure is what can see that. Both are read inside one REPEATABLE READ snapshot, so a wallet that certifies mid-run cannot be reported as a defect.
  • checked — how many (leg, window, line) comparisons actually ran. A window is normally one interval; it widens only across points nobody could value, which is why the field is not counted in intervals. A number that collapses between two runs means the population or the coverage narrowed, not that the ledger got cleaner.
  • book-checked — the same comparisons rolled up per book, which is the scope the specification states the identity at. It is the only check that reaches a composite position (a Fluid smart pair whose two sides were re-split against each other): per leg those figures are a share of a joint answer rather than a derivation from one leg's own evidence, so the sum across the book is where the arithmetic is actually testable.
  • env — which checkout produced the line. Prod and staging share this script's logfile (it is keyed on the script name, while the lock and the toolchain marker are keyed on the checkout), so without this field a line in that file names no environment.

When it pages. Exit 1 means the run could not be performed at all (no DATABASE_URL, no recorded arm block, an empty reserve registry). Exit 2 means it ran and found something, and each finding prints a [fail] ledger-reconcile … line naming its own cause:

LineWhat it meansFirst thing to do
<wallet> <interval> <line> residual=… tau=…the engine's figure and the independent recomputation disagree by more than float precisionread the leg named at the end of the line, then that interval's stored movements
book-residual <book> <interval> <line> residual=…the same disagreement, one level up: the book's total does not follow from its legs' stored rows. position-attributed on the line means a composite position is inside the sumread the per-leg lines for the same interval first; if there are none, the finding is about a composite position and only this line can see it
measured-nothing …the run compared nothing at all — no tracked wallets, or figures that should have been compared and none of them wasa run that measured nothing is not a clean run; find out why the population is empty before reading anything else as green
excluded=<n> <wallet> <case> [<remedy>]the same wallet has been excluded on two consecutive scheduled runs, so it is not clearing itselfact on the <case>: coverage, a stalled build, or a replay at its retry cap. The optional <remedy> narrows it further: rederive-at-cap/rederive-failing names a re-derivation that is stuck, and replay-deferred <age> stalled-enrolment/stalled-derivation names a registration whose history was never derived — the first says look at the ingester, the second that a parked replay needs a re-queue. This is the only alarm that pages for a deferred replay; the minutely drain reports one and exits zero
stamp-defect <wallet> complete=true certified=falsea wallet is marked complete over blocks nothing coversdo not wait: this is a defect in the marking, not a backlog
shape …the engine declined to book something, and the reason is not one of the nine the specification allowsan engine question, not a data one
budget …too many declined figures for one wallet-day, or any occurrence of the two classes whose production budget is zero (W6, W8)read the named class. A stored W9 ghost-row or W10 ledger-blind-move row is a shape finding: both were retired by the ledger-first release, whose reading audit books the opening or correction that settles a ledger/reading disagreement. A ghost itself (a snapshot row for a holding the ledger says is empty) is the engine's cross-check and the audit's page now, and the ghost adjudicator still asks the chain which record is wrong. If the wallet was registered after the campaign's arm block and before the #756 release, check that shape first: its plain token movements above the arm were never derived, and the remedy is the ranged re-derivation, not a restatement
unexplained …a figure that is in none of the three accounted-for buckets, or in two of themthe loudest structural finding: nothing in the system explains this pair

A run can exit 2 with zero disagreements. The exit code answers does this need a human?, not did the arithmetic disagree?.

Lines that are printed and never page, deliberately: the bridge cross-check (the engine's figure across a stretch the data could not value, stated anyway so a run whose only imperfection is a reading gap still produces evidence); the tip-boundary tally, printed at zero because a later decision is made against it; the values line described next; and the coverage block after it.

The values line. Both sides of the identity take the two money columns computed at read, from one computation over the rows the engine is served (the right-hand side values first, so the engine's side is answered from the same result and cannot disagree with it about an input). Each wallet prints one line, and the population line prints the sums:

ledger-reconcile: 0x… values=computed readings=13216 differ=412 receipts=25 differ=3 statements=2 late-tokens=0 book-overlays=0 book-withheld=0 factless=0/0/0 (reported, not paged)
  • differ, per relation: rows whose computed pair (and, on a PT movement, the PT mark the lot book strikes it against) is not the stored one on the wire. It is a statement about how the stored column was produced (a price that arrived after the row was written, a rate its writer read on chain, a value the read withholds), never an alarm, and it moves on its own as late bars land. scripts/ops/value-parity.ts says which rows and why.
  • statements: the series reads the wallet's valuation cost, both sides together. Two per wallet (the bars at the hours its rows are valued at, and the rates), plus one or two where some hour's own bar cannot stand for its whole 48h walk-back (a gap in the bars, or a bar from another price source than WETH's), which read those walk-backs. A higher figure means the two sides were handed different rows.
  • late-tokens: tokens the valuation had to read on their own, at an hour its batched read did not foresee. Each costs a statement, not a number.
  • book-overlays / book-withheld: readings valued in the book they were stored in because the registry has since re-booked their asset, and movements withheld because the registry now prices them in a book their leg is not stated in. Both are zero until an operator re-books an asset that has history.
  • factless = readings / movements / pre-window PT fills served their STORED columns because their value reads a rate no fact answers (the loader rule): rows written before migration 116 that the rate-facts backfill has not filled (or could not: no fact reproduces the value they, or another leg of their reading that reads the same rate, were struck at), a reading whose value reads a second chain-read rate, a rate two legs of one reading state two ways, a PT movement with no stamp, a fill stored without its consideration. Such a row is never compared, so differ does not count it. The backfill alone does not bring it to 0/0/0: a reading over a two-level chain-rated composition (wFalconX) and a PT leg over a composed payout asset have no fact their one column can hold, and the writers keep producing them. The contract release that drops the stored columns requires 0/0/0, so it owes the decision the loader rule names first.

Filling the rate facts of rows written before 116 ​

scripts/ops/backfill-rate-facts.ts fills, IN PLACE, the rate facts a row written before migration 116 lacks, so the row is computed instead of served stored, and only where the fact reproduces the value the row stores. A writer before 116 did not always strike a row at the chain's answer: where its archive read failed it took its fallback (the six-hourly series at or before the block on the enrolment replay and a movement, the newest observation on the 6h tick and the live tip) and stored a value struck at that rate. So for each missing rate the run has CANDIDATES, in order: the chain's answer at the row's block (the writers' reader, on the archive endpoint ETHEREUM_ARCHIVE_RPC_URL), then the fallback the row's writer had, as the series stood when the row was written (updated_at); a listed fund's NAV and a per-share rate have none. It fills the first candidate under which the served read reproduces the row's stored columns (every line equal on the wire, or reproduced by the series as its writer saw it, which is late or corrected series data; at least one stored number reproduced), labelled with where it came from (chain, series or live).

A reading's rate is filled only where every leg it reaches reproduces. The read pools a reading's facts over its legs, so a fact filled on one leg also answers every other leg of that reading that reads the same rate, whether the run filled that leg or not: a leg whose writer struck it at another rate (a series write landing between two legs' reads of one 6h tick), or a PT leg over a composed payout asset, whose value reads the payout asset's rate. So the run values every leg of each reading it fills and holds each leg the fill turns computed to the same check; where one does not reproduce, the rate is filled nowhere in that reading and every leg of it that reads the rate stays served stored. A leg the run does not fill that the read computes through another leg's fact is counted apart (computed through another leg's fact), not among the rows left. So every row the run turns computed serves what it served before, up to late or corrected series data, and every other row is served exactly as before the run.

It writes the fact columns only (rate_raw / rate_source, a movement's meta.rateFacts, a fill's consideration_* and rate_facts), only where they are empty, and never updated_at; a fact already there is never replaced, not even one a writer lands while the run is deciding. Idempotent and resumable: each wallet's rows are written as they are decided, and a re-run skips what is filled and tries again what was left.

bash
cd /opt/onchain-credit   # the checkout of the environment being filled
set -a; . ./.env.local; set +a
# 1. Dry run: reads every candidate, checks it against the stored value, logs each row it would
#    fill, writes nothing.
npx tsx scripts/ops/backfill-rate-facts.ts --db-url "$DATABASE_URL" | tee /tmp/rate-facts-dry.log
# 2. Apply.
npx tsx scripts/ops/backfill-rate-facts.ts --db-url "$DATABASE_URL" --apply | tee /tmp/rate-facts-apply.log
# 3. Rows it left are listed, each with the candidates it tried and the line each could not
#    reproduce. A row whose chain read failed ("no candidate rate at its block") is retried by
#    running step 2 again.
# 4. A second --apply must report 0 write statements.

Its last line is the summary: rows scanned, rows lacking a fact, rows filled (and how many facts came from each source), the readings computed through another leg's fact (each one the run turned computed reproduces its stored value, and its log line says so; one the read already computed before the run says as before this run), and why each row it left was left:

  • its value reads a rate its one fact column cannot hold: a wrapper over a chain-rated fund (wFalconX), or a PT leg over a composed payout asset, where no other leg of its reading holds that rate (where one does, the leg is counted as computed through another leg's fact instead). A PT leg's own rate is a reference no value reads, and is not read at all.
  • no candidate rate at its block: the chain did not answer, and the row's writer had no fallback there (a per-share rate, a listed fund, an asset with no series).
  • no candidate rate reproduces its stored value: the row was struck at a rate neither the chain nor the series answers now (a fallback observation since rewritten, or a live tip that stored a vendor level before decision c1). It stays served stored, and so do the legs of its reading that read the same rate (next).
  • another leg of its reading that reads the same rate would turn computed and does not reproduce its stored value: the rate reproduces this row, but a fact on it would also compute that other leg (named in the row's log line, with its stored and computed numbers), which was struck at another rate, or is a PT leg over a composed payout whose value the rate does not reproduce. Every leg of the reading that reads the rate stays served stored.
  • the read would still serve it stored: a PT movement with no stamp, or a rate another leg of its reading states another way.
  • no router log on file for the fill.

Flags: --wallets a,b, --concurrency N (archive reads in flight, default 8), --page N (rows per write, default 500). There is no --series-fallback any more, and the run refuses it: a row its writer struck at the series is filled series by the check itself, and one the series does not reproduce is never filled from it.

What the run could and could not see. Every run also prints, at zero as well:

  • which classes of declined figure it derived for itself, and which it could not evaluate at all with the reason and the change that closes each. Today one class is named there: the check that a wallet's coverage certified at the tip needs a certificate read the served data layer performs, and this job does not perform it. Nothing else in the output tells "no such case arose" apart from "this run could not have seen one", and a reader of the acceptance evidence would otherwise draw a conclusion the run does not support.

  • how many comparisons left the check and under which rule — the run-level excluded figure covers wallets; this covers individual figures, so checked is not the only visible measure of the narrowing.

  • the number of declined figures beside the sum of their sizes. The budget is enforced on the count, and the sum is printed beside it, unbudgeted: the nine classes record nine different kinds of quantity (one of them records none at all, by design), so their sum is not a quantity anything can sensibly be compared against.

    Two of the nine cover a RUN of periods rather than one, and are filed once for the run they cover: an unread stretch, and a stretch where the stored readings hold a value for a holding the movements say was never there. Filing once is deliberate — otherwise the budget would count days of a single reading gap instead of counting reading gaps — and each row carries the period it closed at so the whole run is accounted for. Before R9 the row claimed only the period it was filed at, so every other period of the same stretch reported as "the engine declined and the reason has no shape", which is what 2,214 of production's 3,134 such findings were. The second of the two is NEW in R9 and its production budget is zero: the readings and the movements contradicting each other about whether a holding existed is never a legitimate steady state, so giving it a name must make the decline legible without retiring it. The reasoning is written out in the code that enforces it. Whoever sets the budget after the first full run sees both figures and can say which one they are choosing.

  • the row counts each wallet's read returned, including movements that arrived without a quantity (whose production budget is zero) and movements whose type is not one the ledger should be able to hold at all. Both are counted where the read refuses them, so a zero on either is a measurement rather than a field that is always zero.

Two quiet lines that mean "nothing to compare", and neither is a defect. A wallet that holds only assets outside the modelled perimeter, and a wallet whose history is too short to contain a single period — a signup whose enrolment landed between two refreshes — produce no comparison at all. Both are healthy and neither pages, but both are named, because the only thing that distinguishes them from a run whose every figure was declined is that line:

  • not comparable yet: the spine holds N grid point(s) — that wallet has one reading or none, so there is no period to compare over. It enters the check on the next refresh, by the ordinary path. Nothing to do.
  • nothing to compare: 0 (leg, interval) pairs … — the whole run had nothing comparable in it. Expect this on a hand run narrowed to a single new wallet, or to a block range that contains no reading. On the scheduled run over the full population it should not appear, and if it does the question is why the population has no comparable holdings rather than why the run is quiet.

--marker is for the scheduled run only. It maintains the "excluded twice in a row" state, and a hand run that wrote it would either give a genuinely stuck wallet a free pass or page the next scheduled run for a wallet somebody already looked at. A hand run says so on its own report.

The permanent reconciliation alarm ​

scripts/refresh-ledger-reconcile.ts is the job above on a 6h cadence, so a disagreement that appears after the cutover is not silent. It is a thin wrapper: every page condition, every line and every exit code is the reconciler's, described just above. What the wrapper adds is the schedule, one remedy field, and the budget.

Its crontab line is a server step. The root crontab lives only on the box and is in no repository, so no deploy installs it and no rollback removes it. Install it once, after the release that adds the script has deployed:

bash
30 1,7,13,19 * * * /opt/onchain-credit/scripts/run-cron.sh refresh-ledger-reconcile.ts

Verify the install two ways, not one. Reading the next tick's line proves a checkout is running the job; counting the crontab line proves it is this one. Only the second carries the path, and only the first proves it actually runs:

bash
crontab -l | grep -c '/opt/onchain-credit/scripts/run-cron.sh refresh-ledger-reconcile.ts'
tail -5 /tmp/onchain-credit-cron/onchain-credit-refresh-ledger-reconcile.log
# ledger-reconcile: wallets=10 reconciled=10 excluded=0 uncertified=0 env=onchain-credit

A line still has to say which environment printed it

run-cron.sh keys the logfile on the checkout, as it does the lock and the toolchain marker, so /tmp/onchain-credit-cron/onchain-credit-refresh-ledger-reconcile.log is this checkout's history and the alert's why: slice is read back out of this checkout's file. That makes the file one environment's, not the LINE's: a line read out of a page, a paste or a screenshot carries no path with it. The env= field is what tells the two apart: it is derived from the checkout root, it is on every line this job prints, and it is the same label the Telegram alert prints as env:. A line without it names no environment. If the count above returns 0 while the log shows this job's lines, the line you installed is on the other checkout.

Every tick connects, reconciles and reports, whatever the environment says; there is no gate in front of it. scripts/ops/reconcile-ledger.ts is the narrowed hand run whose output is the acceptance evidence.

A brand-new environment with no recorded arm block now pages every 6h

The reconciler refuses to measure a population it cannot bound, so with no portfolio:v2-campaign:arm-block row it exits 1 and pages — where the dormant tick used to print one quiet line. Prod and staging both carry the row (it is outside the scrub's prefixes, so a reseed keeps it), so this is a fresh-database concern: record the arm block, with the INSERT under the writers section, as part of standing an environment up. It is the same row the rebuilt ledger's 6h window falls back to.

When it pages, read the reconciler's table above — the lines are the same. Two additions:

  • excluded=<n> <wallet> <case> <remedy>: <reason>. A wallet excluded on two consecutive ticks pages, and the <case> alone cannot separate two states an operator acts on differently. When the wallet has a re-derivation attempt row, the line carries which: rederive-failing <k>/5 check-archive-provider (still retrying, on a doubling backoff — the fault is upstream) and rederive-at-cap <k>/5 needs-a-human (it has stopped retrying and nothing will clear it on its own). A wallet with no attempt row carries the reconciler's own reason unchanged, which for the ordinary case already reads wait on that wallet's enrolment backfill. Without that field the page names a wallet while the fault is a loop, and all three read as "wait".
  • PORTFOLIO_RECONCILE_WITHHOLD_BUDGET in .env.local — the declined-figure budget, as a whole count per wallet-day, never a percentage. Set it from the first full run over the rebuilt history. Unset is legitimate and the run says so on its own report; the two classes whose budget is zero are enforced either way. A value that is not a whole number is fatal rather than a silent fall back to unset, because an unbudgeted alarm reports green on exactly the condition the budget exists to catch.

A wallet stuck at the re-derivation cap now pages from here

It did not always. The re-derivation trigger shipped before this alarm woke, and a rebuilt-ledger fault never reaches the minutely drain's exit code by design — so in that window a wallet that had failed five re-derivations produced no page from here, no [fail] line from the drain, and no alert at all: it was visible only in the drain's log. With the gate deleted this alarm carries it on the first tick after it happens. The drain's rederive-attempts figure is still the fastest way to tell "still waiting on coverage" from "the derivation itself keeps failing": at 5 it has stopped retrying and needs a human.

To look at the ledger by hand, run scripts/ops/reconcile-ledger.ts, not this one. This job maintains the "excluded twice in a row" state; a hand run of it would either give a genuinely stuck wallet a free pass or page the next scheduled run for a wallet somebody already looked at. It also takes no arguments — a flag meant for the hand reconciler is refused rather than silently ignored, so a narrowed run cannot be mistaken for the full one.

The ghost adjudicator ​

scripts/ops/adjudicate-ghost-rows.ts answers the one question a ghost leaves open. The engine's ghost cross-check (until the ledger-first release, the reconciler's W9 ghost-row finding) says the stored snapshot holds a valued row for a holding the movement ledger says was not there. It does not say which of the two is wrong, and two opposite faults produce the identical row:

  • the snapshot is stale — the holding really went to zero and the 6h refresher kept republishing the last balance it had (the promotion that stops it). The stored rows are wrong and can be restated.
  • the movement ledger has a hole — the wallet really held it and the derivation is missing inbound receipts, so its running balance sits at zero or has gone impossibly negative. Here the stored snapshot is the only correct record, and the fault is upstream.

A third cause used to produce the same shape on the leveraged-vault venue and no longer does: a two-token position whose running total drifted from the venue's own reading of it. Those are now settled from that reading directly (why), so a rerun after that release should not present the adjudicator with any of them. If one appears, it is a hole or a stale snapshot like the other two and the chain still decides.

The same release also gives the other half of every two-token position its own record, so a holding that was paid nothing at a settlement now closes instead of staying open forever (why). After the re-derivation that follows it, expect on production:

  • Eight new records and six corrected balances, across four positions and three accounts. Nothing is deleted.
  • Three reported figures change, and only three — for THIS release. One account's fixed-income day 2026-05-14 goes from +$29,424.91 to +$187.42 (its whole reported lifetime for that book, from +$29,769.87 to +$532.38); a second account's 2026-05-03 goes from −$29,927.29 to +$49.81 (lifetime −$29,491.89 to +$485.20); and the first account's ether book on 2026-07-21 goes from +41.36 ETH to −0.17 ETH (lifetime +41.57 to +0.04). Every other day on every account is unchanged to the cent: 2,564 of the 2,567 day-and-book figures the accounts hold. A fourth figure moving is a shape nobody has measured: stop and report it.
  • All three are restated AGAIN by the Fluid tick-padding re-mark, which is a later release and a separate prod step, and the figures above are only the reading between the two. After that step: 2026-05-14 reads −$14.91 (lifetime ≈ +$280.04), 2026-05-03 reads −$10.47 (lifetime ≈ +$375.92), and the ether book on 2026-07-21 reads −0.20 ETH (lifetime ≈ +0.02). It also moves more cells than three: 13 across the corpus by more than $0.02, because the padding is released at every Fluid close and not only at these three. Its own section carries the full table, and the "stop and report a fourth" rule above applies to this release taken before that one lands.
  • The three DAY figures are exact; the lifetime figures are only exact to the tip. A lifetime total sums every day up to and including the live one, so it moves on its own as the day advances, with no code change: the same reading taken a day apart moved one of the figures above by about $1.60. Reconcile the release on the three DAYS, which are fixed, and read the lifetimes as approximate.
  • No holding is left open at a settlement. The sweep for a position holding a non-zero balance with no closing record, beside a sibling that did close, returns ten rows today and zero after. It is the cheapest single check that the release landed.
  • Withheld-day counts fall and none rise: the "endpoint could not be read" class goes from 21 to 14 across the corpus (12 to 7 on one account, 7 to 5 on another, 2 unchanged on a third). The "composition could not be resolved" class keeps its 4 existing rows and gains 4 more that are RESOLVED rather than withheld, one per closing position. The ghost-row class and the adjudicator's populations do not move at all; if they do, the change is on a plain token holding and is a different shape — stop.
  • Both write-time refusals fire zero times, and so does the new "the venue did not answer for this side" report. Any of them firing means a block where the reading failed, and the remedy is to re-run that range with the archive endpoint present, never to write the row anyway.
  • The withhold budget must be re-set from the post-release run, not carried forward: the re-derivation changes rows on three accounts, so any figure taken before it describes a ledger that no longer exists.

What a restatement moves. It writes zero into the row's quantity AND its two money columns. The money columns are no longer what the page reads (values are computed at read), so that half moves no served number; the quantity half does, and a zero quantity is valued at zero on both lines.

Getting these the wrong way round is unrecoverable in one direction: restating a hole overwrites a correct record with a zero the chain contradicts, the finding goes quiet, and the gate passes on a ledger that is still wrong. So nothing decides from the shape, from the sign of the ledger's balance, or from which record looks more plausible. One third party decides: the chain, read at the contradicted row's own block. A bare token balance is an integer, and the question is any at all? rather than how much?, so the rule carries no tolerance to tune.

bash
cd /opt/onchain-credit && set -a; source .env.local; set +a
# investigate one wallet (read-only, and the only mode --wallet allows)
npx tsx scripts/ops/adjudicate-ghost-rows.ts --wallet 0x…
# adjudicate the whole tracked population and restate what the chain has falsified
npx tsx scripts/ops/adjudicate-ghost-rows.ts --apply

It is read-only without --apply, and it refuses to run at all without an explicit ETHEREUM_ARCHIVE_RPC_URL — the verdict is the archive read, so falling back to the shared default would decide a restatement on the provider this programme convicted of serving incomplete history. The endpoint's host is printed on the run, so the evidence names who answered.

Naming the provider is not the same as checking it, so --apply requires a second one: LEDGER_VERIFY_RPC_URL, the same variable the ingester's scan-receipt cross-check uses, on a different host. Every contradicted row is then read on both. A disagreement is treated as this run cannot judge this row rather than as a tiebreak — there is no third opinion to prefer with, and what is known is only that one of them is wrong — so a single disagreement anywhere refuses the whole run. A dry run is happy with one endpoint, because it decides nothing.

The population it acts on is the one the finding names, and that is checked rather than assumed. The tool runs engine v2 over each wallet, from the same input the served reader builds, and compares the ghost points the engine raised against the rows it enumerated, point by point. Any difference in either direction refuses the run: a row the engine never named would be a restatement behind no finding, and a named point the run did not look at would let a repair report nothing left while the reconciler still fails.

When it refuses. Exit 2, and it writes nothing at all. Every applicable line is printed, not just the first:

LineWhat it means
REFUSED … the chain agrees with the SPINEat least one row is a derivation defect. Fix the derivation and re-derive; do not restate the record that is right
REFUSED: … the adjudicated rows are not the ghost points engine v2 raisedthe run's population and the finding's population differ; the set-drift lines above name the points. An engine question, not a data one
REFUSED: … wallet(s) are not certified completea wallet's coverage does not certify, so W7 withholds its whole grid and it has no ghost finding to repair. Certify it, or use --wallet to investigate
REFUSED: … not backed by exactly one stored rowzero rows means the relation moved under the run — re-run. More than one means the leg exists under two venues at one window, which re-running will not change: that one needs a look
REFUSED: … stored at a block the verdict was not read atthe row describes a different moment from the evidence. Investigate before anything is written
… provably wrong and this run is read-onlythere is work to do and nobody asked for it. Re-run with --apply
compare-and-set missed …the row no longer matches what the run read — its quantity, its block or a mark moved between the verdict and the write, so it was left alone. Re-run

The tracked population is the wallet-enrolment universe — every address any ledger arm has been told to scan, which is wider than the reconciler's certified set. As accounts sign in it will include wallets whose history is knowingly truncated; each of those refuses on its own certification line, so a whole-population --apply needs the enrolled set to be certified, not merely quiet.

The refusal is run-wide, not per row, and --apply is refused together with --wallet for the same reason. The reason is not that a partial repair would turn a gate falsely green — the finding still fires on every pair the holed wallets carry. It is that the acceptance figures were taken against the stored snapshot as it is now: restating rows moves it, and re-opens them. One re-take should cover the repair and the re-derivation together rather than paying for two. Narrowing the population would reach the same place by another route, so investigating one wallet is free and writing is all-or-nothing.

That wait is not free. Until the derivation is fixed and the wallets are re-derived, the product keeps serving the stale rows: on the population measured in R10-A, one wallet's $534 holding across 25 windows that the chain says was never there. The exposure is stated rather than discovered, and the step that clears it is the re-derivation.

What a restatement writes. The row's quantity and both marks go to zero, never to NULL and never deleted: NULL means the price read failed and would trade one finding for another, and deleting would assert the refresher never read the leg — a reading gap, which is not what happened, and which can silently remove a grid point when that leg was the only one at its window. A zero-quantity row is a shape no ordinary writer produces (the refresher emits nothing for a zero balance), and it is the intended one: it says read, and there was nothing there, which is exactly what the chain answered.

Each write happens under the portfolio write lock, as a compare-and-set against every value the run read: the quantity, both marks, and the block the evidence was taken at. The chain reads take minutes on a real population and cannot be held inside one transaction, so this is what carries the run's checks across them — the 6h refresh rewrites a row's block while leaving a stale quantity identical, and a re-mark rewrites a mark while leaving both alone, so neither is visible to a guard on the quantity by itself. A row that moved updates nothing, is reported, and exits 2; nothing is overwritten on a value nobody looked at, and the before/after printed is the row that was actually replaced — on the console line, and in the restatements array under --json. Re-running is a no-op: a zeroed row no longer contradicts anything.

Staging runs the ingester as a BOUNDED HAND-RUN, never as a process ​

There is no creddit-event-ingester-staging, and there must not be one: staging runs no refreshers, and an always-on process nobody deletes becomes the permanent staging refresher that policy exists to prevent. pm2 restart creddit-event-ingester is a PROD command — the only ingester process on the box runs from /opt/onchain-credit — so it never appears in a staging procedure. The prod deploy runs it on every release; deploy-staging.yml must never, and a test asserts it does not. The inverse trap is worse than the prod one: on staging the new code does not run against an old binary, it does not run at all, and the silence reads exactly like a clean cycle.

Exercise the tip path with an attended, foreground, time-boxed run from the staging checkout. Pre-flight by READING the target, never assuming it:

bash
grep ^DATABASE_URL= /opt/onchain-credit-staging/.env.local     # must name creddit_staging
cd /opt/onchain-credit-staging
# (a) N discrete cycles. The sleep matches the loop's own cadence: without it, cycle 1
#     consumes the backlog and the rest scan a head that has moved by seconds.
for i in 1 2 3 4 5; do INGEST_SINGLE_CYCLE=1 scripts/ingester/run-ingester.sh; sleep 60; done \
  2>&1 | tee /tmp/staging-ingester-$(date -u +%FT%H%M%SZ).log
# (b) a time-boxed soak, when the evidence needs wall-clock (receipt contiguity).
timeout -s TERM 60m scripts/ingester/run-ingester.sh \
  2>&1 | tee /tmp/staging-ingester-$(date -u +%FT%H%M%SZ).log

It is never added to PM2 and never left running. Record the command, the start and end times and the number of cycles in the run log — a soak whose length nobody wrote down is not a measurement. Teardown doubles as proof prod was untouched: zero PIDs with a cwd of /opt/onchain-credit-staging, and pm2 list showing the same processes as before.

bash
pgrep -af 'scripts/ingester/ingest-events\.ts' | while read -r p _; do echo "$p $(readlink /proc/$p/cwd)"; done

TWO lines with cwd /opt/onchain-credit is the NORMAL prod state. run-ingester.shexecs the tsx launcher and tsx runs the script in a CHILD process, so every invocation shows two PIDs. If a PID with the STAGING cwd survives, kill that PID — never by name, which would match prod's two.

A new wallet's enrolment backfill: how long, and when it is a symptom ​

Registering a wallet enrols it in the wallet-token stream, and the ingester backfills that wallet's bare-token history from that wallet's own history floor — the rolling 30-day window's start (accounts.history_floor_block, migration 103), or 2026-01-01 for a wallet that has none recorded. The budget is 200,000 blocks per wallet per ~60s cycle, given to each of up to three wallets at once — nothing is split between them. So the wait is ceil(span / 200,000) cycles, and the span is what the window decides:

the walletspan to walkthe wait
added under the rolling window (the ordinary case)~216,000 blocks (30 days)~2 cycles ≈ ~1–2 minutes
the 2026-01-01 floor (tracked before the rolling window, recorded by migration 104) or no recorded floor (a failed boundary resolution)~2.1M blocks (2026-01-01 to head)~11 cycles ≈ ~11 minutes, growing by a day's blocks per day

Before the rolling window it was the full ~3.25M blocks from the 2025-05-21 ingestion floor, i.e. ~17 cycles ≈ ~17 minutes, for every new wallet. Beyond three wallets it is ceil(N/3) rounds, because the slots are handed out in a fixed order and a fourth wallet makes no progress until one ahead of it finishes. That order is the wallet address, ascending — not the order people signed up in — so a fresh registration can sit behind older wallets that are still catching up. It is sorted rather than arbitrary so that "which wallets are advancing" reads the same way from one cycle to the next in the log. Each figure is a floor: a cycle is max(60s, work).

The certificate clears when the catch-up FINISHES, and from_block is not what to watch. What the wallet's completeness certificate asks is two things at once: that the row's status is live, and that its from_block reaches that wallet's own history floor. The second is satisfied from the very first committed window and never moves: planCatchUpWindows walks upward, starting at the wallet's floor and climbing toward the live cursor ~200,000 blocks a cycle, and every committed window stamps the row {from_block: the wallet's floor, to_block: <window top>, status: 'backfilling'}, which stampCoverage merges with LEAST(from_block). So from_block reads the floor from cycle one and crosses nothing. What actually clears the certificate is the flip to status = 'live', written once, only when the walk reaches the live cursor. There is no earlier point at which the wallet is publishable.

An existing coverage row's bottom is the pass's floor, in BOTH directions.catchUpPassFloor hands back the row's own from_block whenever there is a row, and the floor it was given only for an address that has none. Both directions are the same defect from opposite sides, and only the row's own bottom avoids both. A floor ABOVE the row's bottom (the wallet's floor rose while its catch-up was in flight) would skip the blocks in between, which LEAST(from_block) leaves certified and unscanned. A floor BELOW it (the only way to be handed one is a degraded read falling back to 2026-01-01, since a stored floor is written once) would be worse: planCatchUpWindows only ever walks UPWARD from the row's confirmed tip, so nothing beneath the row's bottom would be read, while LEAST would happily LOWER the row's claim to that floor — permanently certifying, on prod today, up to ~1.76M blocks of bare-token history nobody scanned, after which a derivation opening there books missing deposits and withdrawals as return. Deepening an existing row is an ops campaign (backfill-event-ledger.ts, record-ledger-backfill-band.ts), which scans what it stamps. The only wallets that see the fast path are ones with no coverage row yet.

A wait past ~5 minutes for a rolling-window wallet (or ~20 for one with no recorded floor), with three or fewer wallets in catch-up, is a symptom. With more, expect ceil(N/3) rounds and read the ingester log for which addresses are advancing — [ingester/wallet-token] catch-up progress <address> -> <scannedTo>/<liveCursor> per advancing wallet, and N wallet(s) pending catch-up; advancing 3 this cycle when the queue is over capacity. Healthy progress on other wallets while the one you are waiting on has not started is the shape that matters. Until the backfill lands, that wallet's coverage certificate withholds rather than publishing partial numbers; no operator action moves it along, and none is needed.

Tune the slot count with INGEST_CATCHUP_MAX_WALLETS (default 3). It is separate from INGEST_CATCHUP_MAX (the token catch-up) on purpose, so a signup never queues behind a newly-listed vault. Do not change INGEST_CATCHUP_WINDOW or INGEST_CATCHUP_WINDOWS_PER_CYCLE without re-deriving the figures above: every published number for this wait comes from their product.

Reorg repair (what the ingester does by itself, and what an operator does) ​

Each cycle, before any scan, the ingester re-reads the headers of the blocks it holds that the chain has not finalized, and repairs any height whose hash has changed: orphaned raw_events rows deleted by predicate, the canonical header stored, the covering chain_scan_ranges receipts deleted, and the affected cursors (plus the cron's ledger dirty tip) rewound to just below the lowest repaired height. It logs [ingester/reorg] REORG at block(s) …. Nothing is required of an operator for an ordinary reorg; the next cycle re-scans and the next 6h tick re-derives the wallets in the range.

Two things to check if that line ever appears repeatedly for the same height: that the provider serving ETHEREUM_RPC_URL is not flapping between forks, and that INGEST_REORG_SWEEP_MAX (default 256 headers/cycle) is not being hit — the sweep repairs lowest-height-first and leaves the remainder to the next cycle, so a persistently truncated sweep shows as slow progress rather than as an error.

Optional, and recommended before the ledger is relied on: set LEDGER_VERIFY_RPC_URL to a SECOND archive-capable provider. Every live range is then re-scanned against it and the two log digests compared before the receipt is written; a match stamps verified_by, a mismatch holds the cursor and logs loudly. Without it the receipt digest only proves the ledger is self-consistent, which does not catch a truncated eth_getLogs that returned HTTP 200 with half the range — the ingester logs that state once per process at startup.

Deleting portfolio users: every account, or some wallets ​

A user delete removes accounts and everything that hangs off them. It has two shapes: every account, with the tool (scripts/ops/reset-portfolio-users.ts), and some wallets, by hand (the tool deletes every account or none). Both are prod writes that need explicit approval, a data-only backup first, and a slot between 6h ticks: the refresher runs at 50 */6 UTC and reads its population minutes before it writes, so an account deleted in between aborts the rest of that tick's ledger work (it heals on the next tick, but there is no reason to cause it). The dated runs are in the release-steps log.

What a delete must reach. Deleting an accounts row cascades (migration 056) through the tracked-wallet rows, the readings, the ledger, the wallet index, the backfill state, the pre-window PT fills and the assistant's tables. Four things are NOT reached by the cascade, and each takes its own statement:

  • the per-wallet cursor rows in chain_scan_cursors: the family portfolio:derive:v2:<chain_id>:<wallet>[:<suffix>] (the derive cursor, the pending marker :jit-pending, the replay's deferral marker :replay-deferred, the reading audit's :audited-through watermark, the continuous producer's :followed-after record) and the retired portfolio:discover:<wallet>. A survivor claims a wallet is derived or discovered over rows that no longer exist. This is the wipe-on-reset family: it holds per-wallet state only, and a campaign-global value lives under portfolio:v2-campaign:, which nothing here touches (nor the worker's chain-wide rows, portfolio:worker:*);
  • the derivation queue, portfolio_derive_jobs (migration 115, no foreign key by design): a deleted wallet's job would fail its merge on the ledger's foreign key, page, and write the wallet's pending marker back;
  • the wallet's wallet-token and native-eth coverage rows in event_coverage (one per wallet each; the native-ether block feed's since migration 118), from which the ingester re-enrols a wallet every cycle (the streams' '*' markers are chain-wide and stay);
  • portfolio_held_pts, which is keyed by PT, not by wallet: cleared on a full wipe only.

The ledger worker is stopped for the whole delete. A job it has in flight would write back what the delete removes, after the delete has committed. The tool refuses --execute while a worker runs (it checks the worker's own lock, and holds that lock through its transaction so none can start mid-wipe); the hand recipe checks the same lock. And the ingester's continuous producer holds its wallet population for up to five minutes (INGEST_CONTINUOUS_POPULATION_TTL_MS), so until the ingester restarts it can still queue a job for a deleted wallet: both recipes restart the ingester, clear the queue once more, and only then start the worker.

Every account:

bash
# PROD, between 6h ticks. 0. stop the worker (the tool refuses --execute otherwise), and save it
#    stopped, so a reboot inside the window does not bring it back running.
pm2 stop creddit-ledger-worker && pm2 save
# 1. back up what is about to be deleted. --table-and-children (pg_dump 16), because -t on a
#    partitioned table dumps the parent, which holds no rows.
sudo -u postgres pg_dump -d creddit --data-only \
  --table-and-children=onchain_credit.accounts \
  --table-and-children=onchain_credit.account_wallets \
  --table-and-children=onchain_credit.portfolio_position_snapshots \
  --table-and-children=onchain_credit.portfolio_flow_events_v2 \
  > /root/backup-users-$(date +%F).sql
# 2. the wipe: a dry run (it deletes, reports, verifies and rolls back), then --execute, which
#    requires --confirm-db (matched against current_database()). Through the socket: without a
#    host, node-pg dials TCP and the postgres role has no password there.
sudo -u postgres bash -lc 'cd /opt/onchain-credit && \
  DATABASE_URL="postgres://postgres@/creddit?host=/var/run/postgresql" \
  npx tsx scripts/ops/reset-portfolio-users.ts'
sudo -u postgres bash -lc 'cd /opt/onchain-credit && \
  DATABASE_URL="postgres://postgres@/creddit?host=/var/run/postgresql" \
  npx tsx scripts/ops/reset-portfolio-users.ts --execute --confirm-db=creddit'
# 3. rotate SESSION_SECRET in /opt/onchain-credit/.env.local, so no old cookie resolves to a
#    deleted account, then restart the app (it picks the secret up) and the ingester (it drops
#    its in-flight enrolment backfill and its producer's population).
pm2 restart onchain-credit --update-env
pm2 restart creddit-event-ingester --update-env
# 4. what the tool does not own. AFTER the restart, or the still-running ingester re-writes the
#    coverage rows underneath the DELETE. Then the jobs its producer queued for the deleted
#    wallets before the restart.
sudo -u postgres psql -d creddit -c \
  "DELETE FROM onchain_credit.event_coverage WHERE stream IN ('wallet-token', 'native-eth') AND address <> '*';"
sudo -u postgres psql -d creddit -c "DELETE FROM onchain_credit.portfolio_held_pts;"
sudo -u postgres psql -d creddit -c "DELETE FROM onchain_credit.portfolio_derive_jobs j
  WHERE NOT EXISTS (SELECT 1 FROM onchain_credit.accounts a WHERE a.uid = j.wallet);"
# 5. start the worker and save it running (step 0 saved it stopped), then confirm the ingester
#    enrolled nothing. WAIT ONE 60s CYCLE: the line prints only when the count CHANGES, so the
#    pass is the drop (`enrolled wallets: 11 -> 0`).
pm2 restart creddit-ledger-worker --update-env && pm2 save
sleep 70
pm2 logs creddit-event-ingester --lines 50 --nostream | grep 'enrolled wallets'

With no wallet enrolled, the wallet-token stream is not emitted, so the ingester alarm's feed-lag arm pages about 6h later until a wallet enrols, and the 6h reconciler pages measured-nothing until one is tracked. Neither is a fault.

Some wallets. First check that no account you are KEEPING tracks one of them, or the tracker re-creates that wallet's account at its next load:

sql
-- read-only: must return no rows
SELECT w.account_uid, w.wallet FROM onchain_credit.account_wallets w
 WHERE w.wallet IN ('<lower-cased uid>', …) AND w.account_uid NOT IN ('<lower-cased uid>', …);

Then stop the worker and save it stopped (pm2 stop creddit-ledger-worker && pm2 save, so a reboot inside the window does not start it), take the backup above, and delete in one transaction. The first statement refuses while a worker holds its lock, and then holds that lock itself until the commit:

sql
BEGIN;
DO $$ BEGIN
  IF NOT pg_try_advisory_xact_lock(hashtextextended('onchain_credit.ledger_worker', 0)) THEN
    RAISE EXCEPTION 'a ledger worker is running: pm2 stop creddit-ledger-worker, then start again';
  END IF;
END $$;
CREATE TEMP TABLE doomed (uid text PRIMARY KEY) ON COMMIT DROP;
INSERT INTO doomed VALUES ('<lower-cased uid>'), ('<lower-cased uid>');
DELETE FROM onchain_credit.event_coverage
 WHERE stream IN ('wallet-token', 'native-eth') AND address IN (SELECT uid FROM doomed);
DELETE FROM onchain_credit.portfolio_derive_jobs WHERE wallet IN (SELECT uid FROM doomed);
DELETE FROM onchain_credit.chain_scan_cursors
 WHERE (scope LIKE 'portfolio:derive:v2:%' AND split_part(scope, ':', 5) IN (SELECT uid FROM doomed))
    OR (scope LIKE 'portfolio:discover:%' AND split_part(scope, ':', 3) IN (SELECT uid FROM doomed));
DELETE FROM onchain_credit.accounts WHERE uid IN (SELECT uid FROM doomed);   -- cascades
COMMIT;

The cursor rows are selected by the wallet's own part of the scope, whatever suffix a record carries: the fifth colon-separated part of the derive family, the third of the discover family. Matching the scope's last 42 characters instead, as the 2026-09-22 hand delete did, matches the bare derive cursor alone and misses every suffixed record. scripts/ops/scrub-staging.test.ts pins this selector against every per-wallet scope the code writes. Then, as for every account: rotate SESSION_SECRET (a still-valid cookie would re-create a deleted wallet's account through the assistant's route), restart the app and the ingester, clear the jobs the producer queued before its restart (the NOT EXISTS delete above), and start the worker and save it running (pm2 restart creddit-ledger-worker --update-env && pm2 save).

Running a cron/refresher by hand ​

scripts/run-cron.sh is the wrapper for every scheduled TS job: it sources .env.local, resolves node_modules/.bin/tsx, runs scripts/<script> from the repo root, and appends output to /tmp/onchain-credit-cron/<checkout>-<script>.log. Extra args are forwarded to the script. It takes a non-blocking per-script flock under the same name ($CRON_LOG_DIR/<checkout>-<script>.lock), keyed by the checkout: a second run of the same script in the same checkout while the first is still going logs "skipping this tick" and exits 0 (so a slow 6h job never stacks a copy on itself), while the prod and staging checkouts get distinct lock files and never block each other. The script's real exit code is preserved for cron. NODE/TSX/CRON_LOG_DIR fall back to the box defaults (/usr/bin/node, the checkout's tsx, /tmp/onchain-credit-cron) but are env-overridable for local testing (scripts/run-cron.test.ts runs the wrapper end-to-end through them). The log, the lock and the toolchain marker below are all keyed by checkout, so a tail of a job's log is one environment's history. (The log was keyed by script alone until the release that carried this change — prod and staging interleaved in one file per job — and each environment switches on its own next deploy, staging first. The first tick of each job after that opens a new file; the old shared ones stay in the directory, unwritten, until /tmp is cleared.)

Toolchain guard (deploy window). A deploy rebuilds the checkout's node_modules in place (npm ci removes the tree before repopulating it) while the crontab keeps firing, and both checkouts drain the backfill queue every minute. A tick landing in that window finds no tsx and dies with MODULE_NOT_FOUND: a job that never started, not a job that failed. The wrapper skips such a tick and exits 0, exactly like the flock skip, logging "deploy in flight, skipping this tick" (and, when stderr is a terminal, saying so on it too, so a hand-run is never mistaken for a job that succeeded).

This removes the common case, not every deploy-window page: npm ci links .bin before the whole tree is reified, so a tick landing after tsx appears but before its dependencies do still fails on a different module, and there is an unavoidable race between the check and the exec.

No log line carries a credential. Node providers put the API key in the URL itself, and several jobs print the endpoint they are about to use as a start-of-run breadcrumb; the registration backfill printed one per wallet. On 2026-08-06 a validation agent found the archive provider's key in plain text in /tmp/onchain-credit-cron/drain-portfolio-backfills.log (the shared name the log carried then), which was mode 0644 on the prod box. The permissions were tightened, and that is a mitigation, not a fix: the wrapper also forwards matching log lines into the Telegram alert, so a leaked line need not stay on the box. Every such site now prints through redactUrl/redactedEndpoint (src/lib/redact.ts), which keeps the scheme, host and port (the part that answers "which provider ran this?") and replaces anything credential-shaped: a key in the path, a key in a query parameter, userinfo before the host. It over-redacts by design. A NEW job that logs an endpoint must go through the same helper — the value should never reach the log in the first place, whatever the file's permissions are. The 2026-08-06 key still wants a rotation decision (it was never reproduced anywhere, including in the campaign's own report). So does the key the event ingester's start line printed, in its live and archive URLs, on every start before the ledger-first release (#954 follow-up 1): that decision is Fred's, and it is that release's step 10, with the count of the lines that still hold the key and the log they sit in.

The skip is bounded, so a checkout whose install genuinely never came back cannot go quiet forever. $CRON_LOG_DIR/<checkout>-toolchain-missing.since holds two stamps, the episode's first miss and its most recent one, and the wrapper escalates only when misses were observed continuously for longer than TOOLCHAIN_GRACE (default 900s, ~15x the measured npm ci): it then exits non-zero with a [fail] line leading with the fix command, which is both halves the alert below needs. Any tick that finds tsx clears the marker. A gap longer than the grace since the last observed miss also restarts the window, which is what keeps a stamp left behind by a paused or not-yet-added cron from making the next healthy deploy page instantly.

Cron-failure alerting (WS8). On a non-zero exit, the wrapper best-effort POSTs a Telegram message naming the checkout, the script, and the exit code, so a failed job is not silent. It is gated on ALERT_TG_BOT_TOKEN + ALERT_TG_CHAT_ID (both must be in .env.local; if either is unset it skips silently, so an un-provisioned box never posts). The alert runs with errexit off and a time-bounded curl, so it can never change the job's real exit code or block the wrapper (a failed POST is logged as "ignored"). Because prod and staging share the box, the checkout label ($(basename "$DIR"), or the ALERT_ENV override) tells you which environment failed. This covers every cron job. Separately, the portfolio refresher posts its own unknown-asset alert to the same bot when an accounting asset it holds with value is unmapped in buckets.ts (see Data pipeline → refresh-portfolio.ts). To exercise the alert path locally, set ALERT_TG_API_BASE at a stub endpoint (defaults to https://api.telegram.org).

bash
cd /opt/onchain-credit
scripts/run-cron.sh refresh-assets.ts
scripts/run-cron.sh refresh-vault-capacity.ts
scripts/run-cron.sh refresh-sofr.ts --since=2018-04-03   # args pass through
tail -f /tmp/onchain-credit-cron/onchain-credit-refresh-vault-capacity.log

Server crontab (all UTC):

ScheduleCommandCadence
0 */6 * * *run-cron.sh refresh-assets.tsevery 6h
15 */6 * * *run-cron.sh refresh-vault-capacity.tsevery 6h
30 */6 * * *run-cron.sh refresh-collateral-exposure.tsevery 6h
45 */6 * * *run-cron.sh refresh-lending-positions.tsevery 6h
50 */6 * * *run-cron.sh refresh-portfolio.tsevery 6h (WS4; manual crontab add)
5 * * * *run-cron.sh check-ingester-freshness.tshourly (manual crontab add). Four arms: pages when the event ingester is older than the checkout's build, not online, or missing; and, behind a process that is healthy so its own page is not displaced, when any followed event feed has fallen more than 6h behind the chain, when the ledger worker has stopped or fallen behind, and when a reading audit booked an unexplained correction or a leg went unread twice. See the ingester alarm.
* * * * *run-cron.sh drain-portfolio-backfills.tsminutely (WS5 queue drain). Exit 2 (= Telegram alert) when a wallet is parked as error after exhausting its reclaim budget. The 6h portfolio tick warns if anything sits queued >30min (drain missing/broken).
*/10 * * * *run-cron.sh refresh-portfolio-discovery.tsevery 10 min (portfolio taxonomy T4 §3.6; manual crontab add — apply 050 FIRST). Discovery enrollment drain: one FULL-universe current read per eligible wallet with no FRESH completeness certificate (accounts.discovery_scanned_*, migration 061), capped PORTFOLIO_ENROLL_MAX (default 10). Quiet no-op once every eligible wallet holds one. Selection is freshness-based, not presence-based, so it is also the REPAIR half of the staleness gate: a wallet demoted for a frozen certificate is re-certified within one run instead of paying the full-universe read tax indefinitely. Un-scanned wallets force the whole 6h batch to a full-universe read, so this is what makes the discovery bound ENGAGE for new signups.
20 3 * * *run-cron.sh refresh-morpho-universe.tsdaily (T4 §3.5; manual crontab add — apply 041 FIRST). Exhaustive Morpho Blue market universe via a cursor-driven CreateMarket scan. First run: set MORPHO_UNIVERSE_FROM_BLOCK to the Morpho Blue deploy block so it walks the full history; steady-state scans only new blocks.
40 3 * * *run-cron.sh refresh-metamorpho-factory.tsdaily (T5 §3.5; manual crontab add — apply 052 FIRST). Exhaustive MetaMorpho vault universe via a cursor-driven CreateMetaMorpho scan over BOTH factories. First run: set METAMORPHO_FACTORY_FROM_BLOCK to the v1.0 deploy block; steady-state scans only new blocks.
30 4 * * 0run-cron.sh refresh-portfolio-reconcile.tsweekly, Sun 04:30 (T5 §3.6; manual crontab add — apply 050 FIRST). Discovery reconciliation safety net: reads each SCANNED wallet against the COMPLETE universe and adds any membership the index missed, with a WS8 drift alert. Capped PORTFOLIO_RECONCILE_MAX (default 50), least-recently-reconciled first. Scheduled AFTER the two universe ingests so "the complete universe" really is complete. The one-time post-ingest step is in the taxonomy release step. Wallets rotate least-recently-RECONCILED first (accounts.discovery_reconciled_at, migration 061 — a column only this sweep writes, so the 6h tick's certificate advance cannot flatten the rotation). Exits non-zero on four conditions, each printing a [fail]-tagged line so the cron alert has both halves it needs: STARVED (eligible wallets present but an empty batch = the backstop checked nothing), an AGED DIVERGENCE (a membership the index lacked while the ledger already carried its flow — a membership the byproduct simply has not reached yet is repaired quietly, so a brand-new position never pages), a DANGLING LINK (counted as FOUND, not repaired, so a run whose repairs all failed pages more loudly rather than less), and an all-errors run. The dangling-link check runs BEFORE the batch selection, so an empty cohort — the situation a bulk account delete creates — cannot hide it.
30 4 * * 1run-cron.sh sync-portfolio-tokens.tsweekly, Mon 04:30 (T6; manual crontab add — apply 049 FIRST). Portfolio token-registry sync: propose + alert only, never writes on the cron. A human applies an entrant with --approve <SYMBOL>.
50 3 * * *run-cron.sh sync-money-market-funds.tsdaily (money market fund coverage, processes.md F; manual crontab add - apply migration 086 FIRST). Discovers every Morpho mainnet vault of both generations, confirms the listing facts on chain, and PROPOSES anything under an unapproved manager. Never lists a new house on its own.
0 5 * * 1run-cron.sh refresh-token-market.tsweekly, Mon 05:00 (manual crontab add — apply migrations 111 and 112 FIRST). The pricing categories' market limb: per tracked yield-bearing token, how many of the last 30 UTC days it traded on a DEX, the median daily volume over those same 30 days as the price vendor reports it (exchanges and DEXes together, with the on-chain median stored beside it), and the deepest single pool holding it against an asset of its own book or against a recognised counter of that book (a dollar stable at par, or ether and its liquid wrappers), never counting a side that is the pool's own share token. ONE Dune execution over every token (DUNE_QUERY_ID_DEX_MARKET) plus one CoinGecko market-chart call per measured row plus a pool-reserves walk per token (up to three pages of twenty, stopping at the first pool that clears the $1M bar or at the end of the vendor's list, because the endpoint will not sort by reserves), paced to the endpoint in use (with COINGECKO_API_KEY set, 30 a minute against CoinGecko's /onchain mirror; without it, four a minute against the keyless GeckoTerminal endpoint, which is what the vendor actually allows). A 429 is waited out and retried twice rather than costing the token its reading. Scheduled 30 minutes after the registry sync, so a row approved that morning is measured in the same hour rather than a week later. It DECIDES nothing: the six-hourly refresh-assets.ts reads the readings, and a category flip is a registry edit in a pull request. Exits non-zero when a whole LIMB wrote nothing — no Dune row for any token, no reported volume for any token asked about (a 404 is an ANSWER and not silence), or no pool reading for any token — each with its own [fail] line. A row counts as measured on any ONE of its readings, so the row count cannot see a limb that has stopped: an unset query id, a thrown execution or a pass throttled end to end all leave every verdict null for the week while two thirds of the rows still look fine. A single token the vendor could not answer for is not a failure; that token's verdict stays null rather than turning into a wrong answer.
30 3 * * 1run-cron.sh refresh-vault-risk.tsweekly, Mon 03:30
0 13 * * 1-5run-cron.sh refresh-sofr.tsweekdays 13:00
10 4 * * 1run-cron.sh chat-retention.tsweekly, Mon 04:10 (chat retention, audit B6; manual crontab add). Prunes chat conversations idle >365d (messages cascade via the migration 038 FK) and chat_usage rows >400d, in one transaction. Portfolio history is untouched (retained in full by design). Prod only.
15 2 * * *ops/backup-creddit.shdaily DB backup, 02:15
0 3 * * *ops/reseed-staging.shdaily staging reseed from the 02:15 dump (pause: touch /root/.reseed-paused)
hourlydisk-usage alerthourly

The staggered :00 / :15 / :30 / :45 on the 6h jobs spreads RPC/API load. Note ops/backup-creddit.sh (02:15) runs ahead of any data job, so the daily dump is a quiet-state snapshot.

psql ​

bash
psql -U onchain_credit -d creddit                 # prod app role
psql -U onchain_credit_staging -d creddit_staging # staging

# inside psql
SET search_path TO onchain_credit;
\dt onchain_credit.*
SELECT max(snapshot_ts) FROM onchain_credit.assets;

Migrations ​

DDL lives in scripts/sql/NNN-*.sql, applied in numeric order as the postgres owner (the app role is read-only against the data and lacks DDL rights). Applied files are tracked in onchain_credit.schema_migrations, and the runner scripts/ops/migrate.sh is the canonical way to apply pending ones:

bash
# on the box — apply any pending migrations to a database, idempotently
cd /opt/onchain-credit
scripts/ops/migrate.sh "$(grep ^DATABASE_URL= .env.local | cut -d= -f2-)"        # prod
scripts/ops/migrate.sh "postgresql://onchain_credit_staging@127.0.0.1:5432/creddit_staging"  # staging

The runner lists scripts/sql/*.sql in order, skips any already recorded in schema_migrations, applies the rest inside a transaction, and records them. It is flock-serialized with a lock_timeout/statement_timeout so a blocked migration fails fast, and it refuses a file tagged -- DESTRUCTIVE unless passed --allow-destructive (expand/contract discipline; the deploy never auto-drops). A file needing to run outside a transaction (e.g. CREATE INDEX CONCURRENTLY) tags itself -- NO-TRANSACTION.

  • Staging auto-applies additive migrations inside deploy-staging.yml.
  • Prod migrations stay a gated manual step for now: after a release deploys, run migrate.sh against creddit by hand (it is a no-op if nothing is pending). Wiring migrate.sh into deploy.yml is the planned next step once the ledger has proven itself over a few releases.
  • Migrations are forward-only and must be backward-compatible with the previous release (expand/contract): the deploy's code-rollback path does not roll back the DB, so a rolled-back build still runs against the migrated schema.

After applying, run the matching refresher/backfill so the new columns/tables fill, then (if data changed) re-deploy to re-prerender (see §2.1 ISR note).

Database backup ​

scripts/ops/backup-creddit.sh (in-repo) runs daily at 02:15 UTC from cron: pg_dump -Fc creddit → /opt/backups, integrity-checked (pg_restore --list), 14-day retention. The same custom-format dump is what reseed-staging.sh restores into staging, so the backup is continuously restore-tested. On-box only for now; an off-box encrypted copy is a documented follow-up (needs a storage target + key). RPO is up to 24h with no PITR, accepted because the data is re-derivable from chain (the refreshers rebuild it); the only non-derivable data is the newsletter/capacity email tables.

Log rotation ​

pm2 does not rotate its own logs, and pm2 restart does not truncate them. Left alone, /root/.pm2/logs/onchain-credit-error.log accumulates forever — it reached 62k lines / 4.1 MB spanning every deploy since the box was built before this was fixed.

Two independent pieces, both needed:

PieceWhat it fixesWhere
/etc/logrotate.d/onchain-creditdates and caps the files (daily, 14 days, maxsize 50M, compressed)reference copy: scripts/ops/logrotate-onchain-credit.conf
pm2 --time on both creddit appstimestamps the linespersisted in the pm2 dump via pm2 save

Rotation alone is not enough: it dates the files but every line inside is still anonymous. --time is what makes a line answerable. It survives deploy.yml's pm2 restart because pm2 save writes it to /root/.pm2/dump.pm2.

Reading the error log after a deploy: don't trust the tail. Lines have no inherent ordering guarantee relative to now, and (pre---time) months-old lines sit directly above fresh ones. Three known-benign residents cost triage time repeatedly:

  • Failed to find Server Action "<hash>" — a browser tab on the old build POSTs an action id the new build doesn't have. Expected on every deploy; self-heals on reload. Hundreds of distinct hashes = many deploys, not one incident.
  • Failed to load external module pg-<hash> — alarming (pg is the Postgres driver) but historical; it does not recur and DB-backed pages serve 200.
  • Single item size exceeds maxSize — Next data-cache notice, item too big to cache. Present on staging too. Not a page failure.

To tell live from historical, don't read — measure. Snapshot wc -l on the log, curl the DB-backed pages, then print only the newly appended lines:

bash
L=/root/.pm2/logs/onchain-credit-error.log; B=$(wc -l < $L)
for p in / /portfolio /carries; do curl -s -o /dev/null -w "%{http_code} $p\n" "http://localhost:3001$p"; done
sleep 3; A=$(wc -l < $L); sed -n "$((B+1)),${A}p" $L

Anything not emitted by a fresh request is not the release's problem.

Scope is deliberate. The glob matches only the four creddit logs (onchain-credit-{out,error}.log, onchain-credit-staging-{out,error}.log). The event ingester's and the ledger worker's pm2 logs (creddit-event-ingester-{out,error}.log, creddit-ledger-worker-{out,error}.log) are ours, but the reference copy does not match them either, so nothing rotates them unless the box's copy was extended; adding both is carried in the ledger-first plan's §11 (docs/plans/portfolio-ledger-first-plan.md). The co-resident DEX HQ logs (dexhq, rindexer, creddit-indexer) are not rotated by this config and are not ours to touch — rindexer-out.log alone is 286 MB. For the same reason the config uses copytruncate rather than a pm2 reloadLogs postrotate, which would reopen every pm2 process's handles box-wide.

Access gate (NEXT_PUBLIC_ACCESS_CODE, per environment) ​

src/components/AccessGate.tsx is an opaque full-screen overlay that covers the app until the visitor types an access code (then remembered per browser in localStorage). It is off unless the environment sets NEXT_PUBLIC_ACCESS_CODE, so each environment decides independently. It replaces the old hardcoded EarlyAccessGate (code letsgetonchain), which was removed in v0.3.0.

  • Build-time, not runtime. Next inlines NEXT_PUBLIC_* into the client bundle and every page here is prerendered, so the value is baked by npm run build. Setting the var and only running pm2 restart does not arm the gate; a deploy (or a manual npm run build + restart) does.
  • Soft gate, zero access control. NEXT_PUBLIC_ACCESS_CODE is inlined into the client bundle at build time and compared in client React state, so the code ships in the bundle, the localStorage flag is trivially set by hand, and page content stays server-rendered underneath (deliberately, so SEO/structured data still indexes). It is cosmetic friction and beta signalling, never a security control. Everything private is protected server-side by the session cookie on /api/portfolio/* and /api/chat* regardless of the gate.
  • Rotating the code: bump STORAGE_KEY in the component (currently creddit_access_v2) to re-prompt browsers that already unlocked with the old one.
  • Staging does not need it. Staging is protected at the edge by nginx basic auth (real auth, no bundle leak); the overlay would only add a second prompt.
bash
# gate creddit.xyz
echo 'NEXT_PUBLIC_ACCESS_CODE=<code>' >> /opt/onchain-credit/.env.local
cd /opt/onchain-credit && npm run build && pm2 restart onchain-credit --update-env \
  && bash scripts/ops/restart-ingester.sh   # every rebuild restarts the ingester, or the freshness alarm pages
# ungate: remove the line, rebuild, restart both

Security response headers ​

Set centrally in next.config.ts (an async headers() block applied to /:path*, so every route including API routes and static assets gets them). There is no middleware.ts; keeping this in one place means the app ships secure headers with no per-route wiring. The same file sets poweredByHeader: false, which removes the default x-powered-by: Next.js response header so we do not advertise the framework/version.

HeaderValueWhy
Strict-Transport-Securitymax-age=63072000; includeSubDomains; preloadForce HTTPS for two years, including subdomains; preload-eligible. Only sent over HTTPS, so it is inert on local http://localhost.
X-Frame-OptionsDENYBlock framing (clickjacking). Reinforced by the CSP frame-ancestors 'none'.
X-Content-Type-OptionsnosniffStop MIME sniffing.
Referrer-Policystrict-origin-when-cross-originSend only the origin (not the path/query) on cross-origin navigations.
Permissions-Policycamera=(), microphone=(), geolocation=(), browsing-topics=()Deny powerful features the app never uses, and opt out of the Topics API.
Content-Security-Policy-Report-Onlysee belowObserve-and-report CSP. Does not block anything yet (see the rollout note).

CSP ships Report-Only first. The policy is emitted as Content-Security-Policy-Report-Only, not the enforcing Content-Security-Policy. In Report-Only mode the browser evaluates the policy and logs every violation but never blocks the resource, so we can watch real traffic and tune the policy before it can break a page. The current policy (assembled from the per-directive list in next.config.ts) is:

default-src 'self'; base-uri 'self'; object-src 'none'; frame-ancestors 'none'; form-action 'self'; script-src 'self' 'unsafe-inline'; style-src 'self' 'unsafe-inline'; img-src 'self' data: https:; font-src 'self' data:; connect-src 'self' https: wss:; manifest-src 'self'

The looser-than-'self' directives are deliberate and documented inline in next.config.ts: script-src 'unsafe-inline' (Next.js bootstrap scripts + inline JSON-LD from src/components/seo/JsonLd.tsx), style-src 'unsafe-inline' (Next.js inline styles + inline style attributes from React components and Recharts' chart wrappers), img-src data: https: (inline icons + remote logos), font-src data: (self-hosted + data-encoded glyphs), and connect-src https: wss: (same-origin fetch/stream + the injected window.ethereum wallet, which opens its own RPC/WebSocket connections client-side). Third-party price/aggregator/RPC APIs (KyberSwap, Pendle, Ethereum RPC) are called from the server, so they are not connect-src entries.

Reading violation reports. With no report-uri/report-to endpoint configured, reports surface in the browser only: open DevTools and look for Content-Security-Policy-Report-Only warnings in the Console (each names the blocked directive and resource), or the Network panel. Exercise the real pages while watching (home, an asset profile with charts, /portfolio, /carries, and connect a wallet) so injected wallet/chart/JSON-LD behaviour is covered. If a first-party resource trips the policy, widen the specific directive; do not blanket-add hosts.

Flipping to enforcing. Once the console is clean across the surfaces above, change the header key in next.config.ts from Content-Security-Policy-Report-Only to Content-Security-Policy (value unchanged), deploy, and re-verify. Do this as its own change so a regression is easy to bisect and revert.

Verify after a deploy:

bash
# prod edge: expect the six headers present and NO x-powered-by
curl -sI https://creddit.xyz/ | grep -iE 'strict-transport|x-frame|x-content-type|referrer-policy|permissions-policy|content-security|x-powered-by'

x-powered-by must be absent; content-security-policy-report-only (not content-security-policy) must be present until the flip. Note the edge (Cloudflare) may add or normalise some headers; the origin values above are what the app emits.

5. Verifying a deploy ​

After a push to main (or a manual deploy):

  1. CI is green. Watch the run on GitHub Actions ("Deploy to production"); the final log line is Deployed <short-sha> at <iso-ts>. A red run that ends in Rolled back to <sha> means the build failed and the previous commit is live — fix and re-push.
  2. Process is up and on the new code.
    bash
    pm2 list                                 # onchain-credit "online", restart count bumped
    ssh root@dexhq.io 'cd /opt/onchain-credit && git rev-parse --short HEAD'
    The short SHA must match the commit you deployed.
  3. App responds locally + DB reachable.
    bash
    ssh root@dexhq.io 'curl -sI http://localhost:3001/ | head -1'          # expect HTTP 200
    ssh root@dexhq.io 'curl -s http://localhost:3001/api/health'           # expect {"ok":true,"db":"up",...}
    /api/health is a 200 + SELECT 1 probe; the deploy workflows curl it (staging today, prod once trusted) and fail the deploy if it is not healthy.
  4. Edge serves it. Load https://creddit.xyz and the changed page. Remember ISR: if you also ran a refresher, the page may show the pre-refresh snapshot until the revalidate window elapses — re-deploy to force a re-prerender.
  5. No errors in logs. pm2 logs onchain-credit --lines 100.
  6. Data freshness (if relevant). Spot-check max(snapshot_ts) on the affected table via psql to confirm a manual refresher actually wrote.
  7. End-to-end smoke. The Playwright suite already ran against staging on this exact code (§3.1). It reports, it does not gate: read its result before opening the release PR, but nothing blocks the merge on it. Nothing runs automatically against prod. To point the suite at https://creddit.xyz by hand (read-only page loads, but still prod traffic), get explicit permission first; the target is E2E_BASE_URL (set it and playwright.config.ts boots no local server). Prod is served on the public hostname with no basic auth in front of it, so no credentials are involved; to point the suite at a firewalled app port on either box, tunnel to it as the staging smoke does (§3.1). Full procedure: Processes → Verifying a UI change.

External dependencies ​

The deploy itself only needs GitHub + the box. The running app and its refreshers depend on (configured in .env.local, preserved across deploys):

DependencyUsed forConfig
Ethereum RPC (current state)live on-chain readsETHEREUM_RPC_URL (fallback publicnode https://ethereum-rpc.publicnode.com)
Ethereum RPC (archive)historical/backfill block readsETHEREUM_ARCHIVE_RPC_URL (fallback dRPC https://eth.drpc.org)
DefiLlama Coins APIUSD prices; basis market pricesrc/lib/data/llama-prices.ts
NY FedSOFR ratesrefresh-sofr.ts
Dune APIthe token_price_bars price mirror, the weekly DEX market measurement + ad-hoc analytics (scarce credits — inspect before executing)DUNE_API_KEY, DUNE_QUERY_ID_HOURLY (8033816), DUNE_QUERY_ID_SUSDE (8283549), DUNE_QUERY_ID_DEX_MARKET (8808795 — the 30-day trading-day and volume measurement, ~0.04 credits an execution, vendored at scripts/dune/dex-market-measure.sql); optional DUNE_DEX_RATIO_MIN_RESYNC_SECONDS (default 12h — the routed feed's freshness/spend dial). Any query id unset ⇒ that path degrades to a logged skip + DB-only reads, never an error — EXCEPT DUNE_QUERY_ID_DEX_MARKET, which prints a [fail] line instead: without it the trading-day count never arrives and every pricing verdict stays null for ever, silently.
CoinGeckolive prices, hourly history for a hole, and the weekly REPORTED daily volume the pricing-category test's volume bar reads (market_chart/range total_volumes, exchanges and DEXes together; a 404 means the vendor lists no such coin and that row's bar falls back to the on-chain median)COINGECKO_API_KEY (Demo; keyless works at a much lower rate limit)
GeckoTerminal / CoinGecko /onchainpool reserves for the pricing categories' market limb (reserve_in_usd per pool), never counting a side that is the pool's own share token. One dataset, two doors: with COINGECKO_API_KEY set the job reads api.coingecko.com/api/v3/onchain/… at 30 requests a minute; unset it falls back to the keyless api.geckoterminal.com endpoint.COINGECKO_API_KEY (optional here, required by the live-price path). The keyless door publishes 30 a minute and delivers about five: measured 2026-09-22 against the live vendor, five 200s then a hard 429 on every request for ~40 seconds, so the fallback is paced at four a minute. TIME THE RUN BY PAGES, NOT BY TOKENS: the walk asks for a second page only where the first both failed the $1M bar and filled up, so 23 rows is 23 requests at best (about six minutes), 34 on the 2026-09-22 shape (about nine) and 69 at the three-page ceiling (about seventeen). Keyed it is about a minute. A key that is set and REFUSED (expired, mistyped, wrong tier: a 401 or a 403) falls back to the same door mid-run, with a [fail] line naming the variable, so a bad key costs a slow week rather than a lost one.
Morpho Blue GraphQL + Euler Goldsky subgraphcurator / market data—
Telegram bot (cron alerting)WS8 failure + unknown-asset alertsALERT_TG_BOT_TOKEN, ALERT_TG_CHAT_ID (both unset = alerting off); optional ALERT_TG_API_BASE (default https://api.telegram.org), ALERT_ENV (env label override)

Optional knobs live in the same .env.local and have nothing behind them to be down.PORTFOLIO_LIVE_COOLDOWN_MS is how long a signed-in wallet's live read is refused after one has just been taken, in milliseconds; unset it is 300000 (five minutes), and that window is what the portfolio's Synchronize button shows as its wait. Lowering it buys fresher readings for more RPC, and 0 removes the bound altogether, which makes one open tab per wallet enough to read the chain as fast as the pipeline runs. It is read per call rather than at import, so pm2 restart --update-env really does move it instead of no-opping on a module the app has already loaded; a value that is not a finite, non-negative number falls back to the default rather than being taken as zero. The full table of variables is External dependencies → Environment variables.

RPC access goes through src/lib/data/rpc.ts. Token-basis redemption rate comes from the token_yield_apy table; market price for basis from DefiLlama.

6. AI assistant (Creddit Agent) ​

The AI research assistant (Creddit Agent) is backed by /api/chat* routes that stream a tool-using Claude response over creddit's own data (carries, repo rates, capacity, oracles, basis, assets). It is code + schema + env: the code and migrations ship through the normal pipeline; the secrets and the nginx streaming block are box-local manual steps.

Its release is deferred and it has no UI. No route, nav entry or dock renders it; the API, the schema and the retention cron are all still in place so it can be picked up later. With no surface calling it, CHAT_ENABLED is the only thing between a hand-rolled POST and real token spend: leave it unset in every environment until the assistant is released. With the flag unset POST /api/chat answers 503 before any model call; the two GET /api/chat/conversations* routes do not read the flag, stay session-gated (401 signed out) and return only the caller's own stored history, so they spend nothing. Nothing else about this section changes when the UI comes back.

AI assistant env ​

Set in each environment's .env.local (preserved across deploys; next start re-reads it on pm2 restart). The feature stays dark (route returns 503) until CHAT_ENABLED is truthy and ANTHROPIC_API_KEY is set. While the release is deferred, that is the intended resting state everywhere.

VarPurpose
ANTHROPIC_API_KEYModel access (secret). Never in the repo.
CHAT_ENABLEDKill switch (1 to enable).
CHAT_MODELModel id (default claude-sonnet-5).
SESSION_SECRETHMAC key for the account session cookie (creddit_session). Generate 32+ random bytes; required for sign-in. Falls back to CHAT_SESSION_SECRET if unset, so an existing server keeps working.
CHAT_SESSION_SECRETLegacy name for the session-cookie HMAC key; still honored as the SESSION_SECRET fallback.
CHAT_DAILY_TOKEN_BUDGETPer-user daily token cap (default 250000).
CHAT_GLOBAL_DAILY_TOKEN_BUDGETWhole-app daily token cap (default 5000000).
SIWE_DOMAINRequired in production. Pins the SIWE domain (creddit.xyz in prod, staging.creddit.xyz in staging) so the domain binding never derives from the client-controlled Host header, which closes a phishing-relay / Host-spoofing account-takeover path (a signature captured on an attacker origin cannot be replayed against this server). Falls back to CHAT_SIWE_DOMAIN. If unset, the request Host is trusted only when it is in a built-in allowlist of known creddit hosts (a safety net, not a substitute for pinning); any Host outside that allowlist fails verification. See Origin & SIWE hardening.
CHAT_SIWE_DOMAINLegacy name for the SIWE domain pin; still honored as the SIWE_DOMAIN fallback.

nginx (streaming) ​

The chat response is a long-lived stream, so its location must disable proxy buffering. Add to each vhost (onchain-credit and onchain-credit-staging), above the catch-all location /:

nginx
location /api/chat {
    proxy_pass http://127.0.0.1:3001;   # 3002 on staging
    proxy_http_version 1.1;
    proxy_buffering off;
    proxy_cache off;
    proxy_read_timeout 300s;
    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;
}

Then nginx -t && nginx -s reload. Verify tokens arrive incrementally (not one flush) through Cloudflare.

Access ​

Sending a message requires wallet sign-in (Sign-In with Ethereum; injected wallets only in stage 1). Signing costs nothing (a message signature). The chat page and history are viewable without signing in; the per-address + global daily token budgets are the real spend guard.

Private documentation. creddit.xyz