Skip to content

built Built. This is a decision record, not documentation.

What is still current: All of it is the live cost discipline: the rolling prompt cache, the shaped tool payloads and the weighted per-user budget. Three modules cite this page by name (src/app/api/chat/route.ts, src/lib/agent/session.ts, src/lib/agent/message-shaping.ts).

Landed: v0.3.0

Header updated 2026-09-14. The body below is frozen history. All plans.

Plan: reduce chat assistant input-token cost ​

The assistant's cost is almost entirely input tokens. Measured baseline on staging (2026-07-05, claude-sonnet-5, first real session): 3 requests = 224,835 input / 12,012 output tokens, i.e. ~75k input per request against ~4k output. At Sonnet 5 pricing that is roughly $0.19-0.28 per request. The target of this plan is <15k full-price-equivalent input tokens per request (~5-10x cheaper) with no loss of answer grounding.

Ordered by impact/effort. Items 1-3 are the core; 4-6 are refinements. Each item names the exact files. Model choice (CHAT_MODEL) is deliberately NOT part of this plan — it is an env switch, orthogonal to these fixes, and they multiply with it.


0. Why input blows up (read first) ​

Three compounding mechanics, all in src/app/api/chat/route.ts + src/lib/agent/tools.ts:

  1. The tool loop re-sends everything per step. One user turn runs up to 8 model calls (stopWhen: stepCountIs(8)); every call re-sends the system prompt (~10k tokens: persona + docs/metrics.md + 13 tool schemas), the conversation history, and every tool result produced so far in the turn. A turn with 4 tool calls pays for its own tool outputs ~3x over.
  2. Tool outputs are fat. list_carries returns up to 25 CarrySummary objects (~25 fields each incl. a 10-field capacity subobject), every float at full double precision (0.04812345678901234 = ~18 chars where 4-5 significant digits carry the information). get_carry_history returns up to 200 points x 4 labeled fields. get_asset_profile returns multi- paragraph memo narratives verbatim.
  3. Past turns re-send their tool parts forever. The route sends the last 30 UIMessages through convertToModelMessages as-is, so turn 2 re-pays for turn 1's entire tool output volume, turn 3 for both, etc. This is why the average hit 75k within a 3-message session.

Prompt caching should discount the static prefix to ~0.1x, and a cacheControl breakpoint IS set on the system prompt — but we currently cannot see whether it hits: recordUsage stores totalUsage.inputTokens, which is the sum of cached + uncached. Instrument before optimizing.


1. Instrument cache + spend visibility (do this first) — S effort ​

We must separate "tokens sent" from "tokens billed at full price".

  • The AI SDK already exposes the split: totalUsage.inputTokenDetails (noCacheTokens, cacheReadTokens, cacheWriteTokens) in the onFinish of streamText (verified against installed ai@6.0.219 types).
  • Migration 040-chat-usage-cache-split.sql (additive, house style): ALTER TABLE onchain_credit.chat_usage ADD COLUMN IF NOT EXISTS no_cache_tokens bigint NOT NULL DEFAULT 0, ADD COLUMN IF NOT EXISTS cache_read_tokens bigint NOT NULL DEFAULT 0, ADD COLUMN IF NOT EXISTS cache_write_tokens bigint NOT NULL DEFAULT 0; (+ the GRANT is already table-level; no new grant needed).
  • src/lib/agent/budget.ts recordUsage: accept and upsert the three new columns alongside the existing totals.
  • src/app/api/chat/route.ts onFinish: pass totalUsage.inputTokenDetails through. Also console.log one structured line per request ([chat-usage] uid=… in=… noCache=… cacheRead=… cacheWrite=… out=… steps=…) so pm2 logs show it live.

Acceptance: after one staging conversation, a psql query shows the cache-read share per request. If cacheReadTokens is ~0 on the second and third requests of a session, caching is broken — fixing that (item 4 / checking byte-stability of the tool schemas + system prompt) becomes the top priority, because a working cache alone discounts the ~10k static prefix to ~1k-equivalent on every call after the first.

2. Stop re-sending historical tool outputs — M effort, biggest structural win ​

Past turns' tool results add nothing to grounding (the assistant's prose already summarized them) but dominate multi-turn input.

  • In src/app/api/chat/route.ts, before convertToModelMessages: map over messages.slice(-HISTORY_LIMIT) and, for every message EXCEPT those of the current turn (i.e. everything before the last user message), drop all non-text parts (m.parts.filter(p => p.type === "text")). Keep assistant text so conversational context survives; keep the DB/UI persistence untouched (full parts still stored and rendered — this changes only what the model sees).
  • Guard: a historical assistant message whose parts become empty (tool-only turns) should be dropped entirely, not sent as an empty message.
  • Note convertToModelMessages may reject tool parts without results in edge cases; filtering to text-only for history sidesteps that class too.

Expected effect: turn N's input stops growing with the sum of all prior tool volume; on the measured session this is the difference between 75k average and roughly half that, before the other items.

Acceptance: in a 3-turn session where turn 1 called list_carries, the [chat-usage] log line for turn 3 shows input well below turn 1 + turn 2's combined tool volume; answers still reference earlier turns correctly.

3. Slim the tool outputs — M effort ​

All in src/lib/agent/tools.ts (+ one helper). Grounding rule stays: keep asOf/source on every tool; the assistant must still cite them.

  • Rounding helper applied to every numeric field the tools return: fractional APYs/ratios to 5 decimal places (0.04812), USD to whole dollars, utilization to 4. Full-precision doubles are pure token waste.
  • list_carries: return compact screener rows, not full CarrySummary: { key, pair, protocol, category, isTermCarry, carry7d, volAdjCarry30d, maxLevCarry7d, leverage, suppliedUsd, headroomTokens+borrowSymbol, snapshotTs }. Everything else (LT/LTV, vol, trend, worst-week, full capacity object) already comes from get_carry_detail when the model drills in. Default limit 12 -> 10.
  • get_carry_history: cut ~4x. Default maxPoints 120 -> 60 (cap 200 -> 120); drop the derivable carry field (= target - funding); return columnar arrays instead of labeled objects: { cols: ["ts","targetApy","fundingApy"], rows: [["2026-07-01T06:00Z", 0.048,0.031], …] } — removes three JSON keys per point. Update the tool description so the model knows the shape.
  • get_asset_profile: cap the narrative. Return each section's heading + first paragraph only, plus note: "ask for the full <heading> section for more"; add an optional section input to fetch one full section. The memos are multiple paragraphs x 3 sections and are the single fattest non-carry output.
  • get_oracle_report: trim legs to { role, asset, provider, liveValue } (drop detail + onchainDescription from the default; they are verbose prose duplicated by mechanism/collateralPricing).
  • get_money_market_rates / get_curator_funds: apply rounding; drop allocationCount and availableLiquidity unless asked (keep deposited + utilization).
  • While here: ChatMessage.tsx's GenericToolCard slices its JSON preview — no UI change needed for the columnar history, but list_carries' CarryListCard reads carries[] fields — update it to the compact row shape (it uses collateral/funding -> now pair, and capacity.totalSuppliedUsd -> suppliedUsd).

Acceptance: log the serialized size of each tool result (one debug line in the tool wrapper): list_carries (10 rows) < ~2.5k chars, history (60 points) < ~3k chars, asset profile < ~2.5k chars.

4. Cache the growing prefix inside the loop — M effort ​

Item 1 tells us whether the static prefix caches. This item extends caching to the conversation + tool results, so even what must be re-sent bills at ~0.1x:

  • streamText's prepareStep hook (verified: may return a modified messages array) — on each step, attach providerOptions: { anthropic: { cacheControl: { type: "ephemeral" } } } to the last message of the array being sent. Each step then reads the previous step's prefix from cache instead of re-paying it. Mind Anthropic's 4-breakpoint limit: strip any breakpoint set by a previous step before adding the new one (keep: 1 on system, 1 rolling on the last message).
  • Keep the existing breakpoint on the static system prompt. If item 1 shows frequent cold starts (staging traffic gaps > the 5-min TTL), switch it to { type: "ephemeral", ttl: "1h" } — 2x write cost, worth it if >1 request/hour lands on average.
  • Verify the tool-schema prefix is byte-stable across requests (same tool order — buildTools returns a literal object, so order is stable; do not make tool descriptions dynamic).

Acceptance: cacheReadTokens share > 60% of input on the 2nd+ step of multi-tool turns and on follow-up turns within 5 minutes.

5. Tighten the loop + history caps — S effort ​

  • stopWhen: stepCountIs(8) -> stepCountIs(5) in route.ts. The system prompt already says "prefer one broad call over many narrow ones"; 5 is enough for screen -> detail -> history -> answer with margin.
  • Replace the message-count history cap (HISTORY_LIMIT = 30) with a character budget after item 2's text-only mapping: walk history newest -> oldest, keep messages until ~24k chars (~6k tokens), drop the rest. A 30-message cap can still be huge; a char cap bounds it.

6. Budget accounting follows real cost — S effort (after item 1) ​

chat_usage budgets currently count every input token at full weight, so cached reads (0.1x price) burn the user's daily budget 10x faster than they burn money. After item 1's columns exist, change checkBudget to compare against a cost-weighted total: no_cache_tokens + cache_write_tokens*1.25 + cache_read_tokens*0.1 + output_tokens*5 (output priced at ~5x input for Sonnet-class). Keep the env var semantics ("budget" = weighted tokens); adjust defaults if needed.


Rollout + measurement ​

  1. Ship items 1+5 together (tiny, immediate visibility + cap). Read the numbers from one staging session.
  2. Ship 2+3 (the volume cuts). Re-run the same session script; compare [chat-usage] lines.
  3. Ship 4 (+6). Re-measure.
  4. Re-run the eval questions in the plan (docs/plans/ai-assistant-plan.md §1.9) to confirm answer quality/grounding is unchanged — numbers must still match the UI and carry as-of citations. 5-decimal rounding keeps APYs exact to 0.001pp, well inside display precision (UI shows 2-3 decimals of a percent).

Targets: avg full-price-equivalent input (noCache + 1.25*cacheWrite + 0.1*cacheRead) < 15k/request; cache-read share > 60% on multi-turn sessions; no eval regression.

Out of scope here: switching CHAT_MODEL (env-only, stacks with all of the above); get_swap_cost; compaction of very long conversations (revisit if sessions regularly exceed the char cap).

Private documentation. creddit.xyz