Files
pepa-pi-bot/dev/v0.3.1/PRD.md
T
mayatnikovandClaude Opus 4.7 451c0343cb docs(v0.3.1): PRD — LLM prompt cost optimization
Design-only commit; no runtime changes. Spec for the next patch iteration.

Goal: cut per-advise() input tokens from ~800 to ≤300, preserving the
LLM's ability to produce valid registered skill ids and useful rationale.

Five proposed changes ranked by impact:

  P1  Compact registry format (saves ~350t/call) — group by namespace,
      comma-list ids, drop human titles. Default mode for advisor;
      verbose mode kept for postmortem/reflect.
  P2  Need-scoped registry (~50t additional) — show LLM only skills
      relevant to the active Maslow need + always-available safety
      skills (survive.flee, pillar-up, recovery.tunnel-out, explore.*).
  P3  Snapshot pruning (~50t) — drop weather/experience/dimension/biome/
      players from the user prompt; the LLM doesn't consult them.
  P4  Prompt caching probe — check if TimeWeb passes through
      prompt_tokens_details.cached_tokens. If yes, restructure prefix
      to maximize cache hits (cached input is ~10x cheaper at OpenAI).
  P5  Per-trigger cost telemetry in scripts/list-improvements.js --stats:
      avg_in / avg_out / cost_₽ / share% per trigger_reason, using
      TIMEWEB_PRICE_IN_RUB_PER_M and TIMEWEB_PRICE_OUT_RUB_PER_M env.

Trigger: TimeWeb admin panel after first day of v0.3.0 live showed
34K tokens / day at low activity. At cap budget that projects to
~480₽/month (101₽/M in, 608₽/M out for gpt-5.4-mini). Manageable
but the savings are mostly free — repeated infra tokens, not signal.

All changes are additive; runtime behaviour stays the same. If the
LLM produces worse advice with the compact registry, flip back via a
single constant in fast-advisor.js.

Acceptance: re-run scripts/check-timeweb.js probe 3 — expect
tokens_in ≤ 300 (was ~800). Live for 1h, check --stats: avg_in ≤ 300
per trigger group. Existing 360 tests still green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 19:52:05 +03:00

193 lines
8.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# pepa v0.3.1 — PRD: LLM prompt cost optimization
**Status**: Design draft. No code in this version yet — this PRD is the
spec future commits implement against. Owner: operator.
**Trigger**: TimeWeb admin panel after first day of v0.3.0 live:
~34K tokens used in a half-day session (mostly bot + some smoke).
At the 6-calls/hour cap that projects to **~480 ₽/month** (101 ₽/M in,
608 ₽/M out for gpt-5.4-mini). Manageable but worth shrinking — most
of the per-call cost is repeated infrastructure tokens, not the
situational signal the model actually uses.
## Goals
1. Cut per-advise() input tokens from ~800 → ≤300 (target 250).
2. Preserve correctness: the LLM must still see enough context to
produce a valid `skill_id` from the registry and a useful rationale.
3. Keep all changes transparent to the rest of the runtime — the
public `advise()` / `complete()` surface area doesn't change.
Non-goals:
- Switching providers. TimeWeb stays.
- Caching the LLM's *responses* (cache key would be situational, too
many misses to be worth the bookkeeping).
- Touching the analytical loops (postmortem / reflect). They're called
less often and need fuller context; cost there is acceptable.
## Cost breakdown — what we're optimizing
Measured on live advise() calls (TimeWeb gpt-5.4-mini, single advisor
trigger):
| Block | tokens (avg) | % of call |
|--------------------------------------|--------------|-----------|
| `skillRegistryPrompt({limit:1800})` | ~450 | 56% |
| System instructions (rules + JSON) | ~200 | 25% |
| User snapshot + threats + need + recent | ~150 | 19% |
| **Total input** | **~800** | **100%** |
| Output (JSON answer) | ~40-50 | — |
The registry block dominates. It currently lists all 30+ registered
skills with their human titles. The model rarely needs the full list —
most decisions are within 5-8 plausible skills per trigger.
## Proposed changes
### 1. Compact registry format (P1, biggest win)
Drop human titles and the per-skill descriptions; switch to
namespace-grouped, comma-separated id lists.
**Before** (~450 tokens):
```
Valid skill ids (USE ONLY THESE for avoid_skill / prefer_skill):
craft:
- craft.bed — Craft bed
- craft.chest — Craft chest
- craft.furnace — Craft furnace
...
survive:
- survive.acquire-food — Acquire food
- survive.eat — Eat
...
```
**After** (~100 tokens):
```
Valid skill ids (USE EXACTLY one of these or null):
craft: bed, chest, furnace, planks, sticks, torch, wooden-axe,
wooden-pickaxe, wooden-sword, stone-axe, stone-pickaxe, stone-sword
survive: acquire-food, eat, flee, pillar-up, sleep
gather: logs, stone, wool
recovery: tunnel-out
explore: far, wander
village: build-shelter, choose-base, deposit-surplus, place-chest
farm: wheat
diag: physics, scan, match
```
Saving: **~350 tokens/call**.
Implementation: add `skillRegistryPrompt({ mode: "compact" })` mode in
`runtime/skill-registry.js`. Default mode stays for slow analytical
loops (postmortem / reflect) which can afford the verbose form.
### 2. Need-scoped registry (P2, additional ~50 token saving)
When `activeNeed` is set, filter the registry to skills plausibly
relevant to that level + always-available safety skills.
Relevance table (manually curated, lives in `runtime/manifesto/needs.js`):
| Need | Relevant skills (in addition to ALWAYS set) |
|-------------------|------------------------------------------------------------|
| alive | survive.flee, survive.eat, recovery.tunnel-out |
| food | survive.acquire-food, survive.eat, farm.wheat |
| tools_wood | gather.logs, craft.planks, craft.sticks, craft.wooden-* |
| shelter_basic | gather.wool, craft.bed, village.build-shelter, village.choose-base |
| tools_stone | gather.stone, craft.sticks, craft.stone-* |
| armor_basic | gather.wool (placeholder) |
| food_security | farm.wheat, survive.acquire-food |
| tools_iron | gather.stone |
| armor_iron | (none — no skill yet) |
| village_seed | craft.chest, village.deposit-surplus, village.build-shelter |
| village_full | (full registry) |
| ALWAYS | survive.flee, survive.pillar-up, recovery.tunnel-out, |
| | explore.far, explore.wander |
Compact + scoped = **~50 tokens** for the registry block (down from 450).
Add a `prompt-builder.test.js` checking that:
- `survive.flee` is always present (emergency safety)
- The recommended skill from the previous call would still be in the
scoped registry (regression protection)
### 3. Snapshot pruning (P3, ~50 tokens)
The user-prompt snapshot includes fields the LLM rarely consults:
`weather`, `experience`, `dimension`, `biome`, `players[]`. Drop them
from the advise() user-prompt builder. Keep `position`, `hp`, `food`,
`isDay`, `closestHostile`, `activeNeed`, `recent dispatches`,
`hazards.footBlock` (lava detection), top inventory keys.
### 4. Prompt caching — investigation (P4)
OpenAI and Anthropic both support implicit prompt caching: when ≥1024
prefix tokens are identical across consecutive requests, the prefix
is billed once. TimeWeb's docs are silent on this.
Task: probe whether TimeWeb passes through OpenAI's `prompt_tokens_details.cached_tokens`
field. If yes, *increase* the system prefix length (keep verbose registry)
because cached input is ~10x cheaper than fresh. If no, full optimization
1+2+3 still wins.
Add a one-off check in `scripts/check-timeweb.js`: print
`payload?.usage?.prompt_tokens_details?.cached_tokens` if present.
### 5. Telemetry — per-trigger token attribution (P5)
Today `advisor_recommendations` records `tokens_in` per row but the
operator has no easy view of *which trigger types* are most expensive.
Extend `scripts/list-improvements.js --stats` to also print per-trigger:
```
trigger_reason total applied ok fail avg_in avg_out cost_₽ share%
wedged_* 20 18 3 15 280 45 2.1 45%
emergency_* 3 3 2 1 240 50 0.3 6%
repeat_* 8 7 0 7 260 42 0.8 17%
preempt_retry_* 14 12 2 10 290 44 1.5 32%
```
`cost_₽` = avg_in × calls × IN_PRICE + avg_out × calls × OUT_PRICE,
with prices read from env (`TIMEWEB_PRICE_IN_RUB_PER_M`,
`TIMEWEB_PRICE_OUT_RUB_PER_M`).
## Acceptance
After v0.3.1 lands:
- Re-run `node scripts/check-timeweb.js` probe 3 (`advise()`):
expect `tokens_in` ≤ 300 (was ~800).
- Re-run probe 4 (auto-trigger flow): rationale still references the
registered skill correctly.
- Run live for 1 hour, check `node scripts/list-improvements.js --stats`:
per-trigger `avg_in` ≤ 300.
- Existing 360 tests still green; new prompt-builder tests cover the
scoped registry behaviour.
## Out of scope (later versions)
- Tool/function-calling instead of free-form JSON (TimeWeb support unclear).
- Custom model selection per trigger (gpt-5.4-nano for routine repeats,
gpt-5.4-mini for emergencies). Defer until cost/quality data points
exist.
- Embedding-based prior-recommendation similarity check ("we already
told the bot to tunnel-out at this exact wedge 10 minutes ago, skip").
## Implementation order
When this version is greenlit, work in this order on a single branch
`v0.3.1` (one PR per session, per recently-updated workflow memory):
1. P1 compact registry mode + prompt-builder test
2. P2 need-scoped registry (extend `runtime/manifesto/needs.js` with
`relevantSkills`)
3. P3 snapshot pruning in `fast-advisor.js#buildUserPrompt`
4. P4 caching probe (one-off)
5. P5 cost telemetry in CLI viewer
6. STATUS.md + smoke retest + PR
All changes are additive; no behaviour regression expected. If
real-world after v0.3.1 shows the LLM giving worse advice with the
compact registry, fall back to default mode by flipping a single
constant in `fast-advisor.js`.