main
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
15b6c11002 |
v0.3.1: survival behaviour overhaul — storyline, biome-aware scout, wedge-relocate, food/perf fixes, monitor TUI (#28)
* docs(v0.3.1): PRD — LLM prompt cost optimization
Design-only commit; no runtime changes. Spec for the next patch iteration.
Goal: cut per-advise() input tokens from ~800 to ≤300, preserving the
LLM's ability to produce valid registered skill ids and useful rationale.
Five proposed changes ranked by impact:
P1 Compact registry format (saves ~350t/call) — group by namespace,
comma-list ids, drop human titles. Default mode for advisor;
verbose mode kept for postmortem/reflect.
P2 Need-scoped registry (~50t additional) — show LLM only skills
relevant to the active Maslow need + always-available safety
skills (survive.flee, pillar-up, recovery.tunnel-out, explore.*).
P3 Snapshot pruning (~50t) — drop weather/experience/dimension/biome/
players from the user prompt; the LLM doesn't consult them.
P4 Prompt caching probe — check if TimeWeb passes through
prompt_tokens_details.cached_tokens. If yes, restructure prefix
to maximize cache hits (cached input is ~10x cheaper at OpenAI).
P5 Per-trigger cost telemetry in scripts/list-improvements.js --stats:
avg_in / avg_out / cost_₽ / share% per trigger_reason, using
TIMEWEB_PRICE_IN_RUB_PER_M and TIMEWEB_PRICE_OUT_RUB_PER_M env.
Trigger: TimeWeb admin panel after first day of v0.3.0 live showed
34K tokens / day at low activity. At cap budget that projects to
~480₽/month (101₽/M in, 608₽/M out for gpt-5.4-mini). Manageable
but the savings are mostly free — repeated infra tokens, not signal.
All changes are additive; runtime behaviour stays the same. If the
LLM produces worse advice with the compact registry, flip back via a
single constant in fast-advisor.js.
Acceptance: re-run scripts/check-timeweb.js probe 3 — expect
tokens_in ≤ 300 (was ~800). Live for 1h, check --stats: avg_in ≤ 300
per trigger group. Existing 360 tests still green.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.1): storyline — canonical Minecraft survival quest
The bot has been stuck in a loop for two days:
acquire-food (fail: no nearby food) → explore.far → pillar-up (fail) → repeat
Diagnosis: manifesto + LLM advisor both correctly identify "you need
food" but neither expresses *what concretely to do next*. Manifesto is
a priority ladder (need-detection), not a narrative arc.
This commit adds the missing narrative layer — an ordered list of
operational steps that mirror the vanilla Minecraft survival path:
1. orient_self — Понять где я
2. first_wood — Собрать 8 поленьев
3. crafting_basics — Сделать верстак и палки
4. first_tools — Деревянные орудия
5. first_food — Найти первую еду
6. shelter_minimal — Простой шелтер с кроватью
7. stone_tier — Каменные орудия
8. food_security — Запас еды на 16+
9. iron_age — Железо и печь
10. settle_base — Постоянная база
11. village_grow — Развивать деревню (ongoing)
Each step has:
- completed(snapshot) → bool — detects achievement from snapshot
- suggestSkill(snapshot) → { skillId, args? } — concrete next dispatch
- emergencyPause(snapshot) → bool — defers to manifesto L0 alive
emergencies (low HP near hostile, lava under foot, food = 0)
- narration_ru — chat-friendly Russian one-liner spoken on entry
Components:
- runtime/goal/storyline.js — 11-step canonical quest catalogue
- runtime/goal/state.js — pickCurrentStep(snapshot) walks the list,
returns first non-completed step + its suggestion. 3s cache.
Validates suggestSkill's skillId against the live registry.
- runtime/reflex.js — curriculumReflex dispatch priority is now:
1. manifesto (L0 alive emergencies always win)
2. storyline (concrete operational subgoal)
3. curriculum plan (legacy fallback)
Tests pass ctx.disableStoryline=true for isolation.
- runtime/bot.js — snapshot.storyStep populated each tick so
chatter/advisor/reflect observers see the same view.
- runtime/coach/fast-advisor.js — buildUserPrompt now embeds the
current step + its suggested skill, so LLM advice is anchored
("step 5 first_food, storyline wants survive.acquire-food, but
recent dispatches show it's failing — try explore.far + scout").
- runtime/coach/advisor-trigger.js — forwards ctx.storyStep into
advise() and logs step id at trigger time.
- runtime/coach/reflect.js — reflection prompt includes storyline
progress so 30-min self-assessment is anchored.
- runtime/persona/chatter.js — narrates step.narration_ru on
transition. Rate-limited via existing maybeNarrateRaw().
New operator CLI:
- scripts/show-story.js — fetches the live snapshot via IPC sock and
prints step progress with ✓/→/ markers, current skill, inventory.
Falls back to --plain catalogue view when bot offline.
Token cost impact: ~+30 input tokens per advise() call (one extra
line in user prompt). Trivial vs the value of grounding LLM advice
in a concrete narrative.
Operator usage:
node scripts/show-story.js # live progress + which step + why
node scripts/show-story.js --plain # static catalogue of all 11 steps
Tests: 376 green (was 360, +16 storyline tests).
Also in this branch (already committed): dev/v0.3.1/PRD.md —
LLM prompt cost optimization design doc.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(v0.3.1): storyline beats manifesto L1+ (only L0 alive emergencies override)
Found in live logs after the previous commit deployed:
storyline: step 1/11: orient_self → explore.wander
advisor-trigger: firing because wedged (planned=survive.acquire-food, ...)
Manifesto was still picking survive.acquire-food (L1 food) over the
storyline's orient_self → explore.wander. That's the wrong precedence —
storyline expresses a *concrete operational subgoal* and L1+ manifesto
needs are just "you'd benefit from food" priorities, not emergencies.
New dispatch precedence in curriculumReflex:
1. manifesto L0 (alive emergencies: lava, low-HP+hostile, food=0)
2. storyline (concrete narrative subgoal — beats L1+ manifesto)
3. manifesto L1+ (fallback when storyline has no concrete suggestion)
4. curriculum plan (legacy fallback)
This way the bot starts following the narrative arc even while
manifesto's L1 food is technically unsatisfied — orient_self runs to
completion before pursuing food explicitly. Storyline already handles
food as step 5 (first_food), so we're not skipping it.
Tests: 378 green (+2 priority-ordering tests):
- L0 manifesto emergency: upstream reflex (defend/modes) catches before
curriculum dispatch
- storyline beats manifesto when both have suggestions: well-fed bot
with logs → craft.planks (storyline crafting_basics), not gather.logs
(manifesto L2)
- updated "manifesto fallback" test to require disableStoryline=true
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(tui): fullscreen monitor-only TUI (opencode-style)
Replaces the old tui/tui.tsx hotkey-heavy dashboard with a read-only
observability screen. Operator actions live in scripts/* now —
TUI is for watching, not driving.
Layout (top to bottom, all auto-resizing to terminal):
1. Header — MC/IPC status, pos, HP, food, day/night, hostiles
2. Storyline — current step + 11-step quest map (✓/→/○)
3. Activity — last N skill dispatches (colour by outcome)
4. MC Chat — last N chat lines (cyan for bot, yellow for players)
5. Advisor — last N LLM recommendations (trigger + outcome + tokens)
6. Improvements — open requests from knowledge.improvement_requests
7. Footer — 24h token usage + cost in ₽ + q-to-quit
Data sources:
- IPC sock: snapshot frames, log frames, chat frames (push)
- SQLite knowledge.db: advisor_recommendations + improvement_requests
polled every 5s (pull)
Token cost displayed live using TIMEWEB_PRICE_IN_RUB_PER_M /
TIMEWEB_PRICE_OUT_RUB_PER_M env vars (defaults: 101 / 608 for
gpt-5.4-mini).
Switches:
- npm run tui → new monitor (this file)
- npm run tui:legacy → old action-driven tui/tui.tsx (kept for now)
Implementation notes:
- Uses ink + alternate-screen-buffer ANSI for proper "opencode-feel"
fullscreen behaviour; restores prior terminal contents on quit.
- Skips alt-screen and useInput when stdin/stdout isn't a TTY
(smoke tests, piped output) — both gracefully degrade.
- Stable React keys via per-event uid counter, avoids reconciler
duplicate-key warnings as logs/chat/dispatches stream in.
- Resize handled via 1s stdout-dimension poll, NOT direct
'resize' listener (which conflicts with ink's own listener and
triggers MaxListenersExceededWarning).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* ui(tui): compact 4-section monitor (was 6) — fits 1080p without zoom
Operator reported the TUI overflowed the screen unless terminal was
zoomed way out. The 11-step storyline list alone was eating ~13
rows, and each advisor/improvement entry took 2-3 rows. Now:
- Header + storyline collapsed into one panel (2 lines):
line 1: pepa · ●MC ●IPC · 1m50s · pepa_bot · (697,61,702) · HP 20 · food 5 · ☀ · ⚔60(creeper@58b)
line 2: story ▓▒░░░░░░░░░ 1/11 orient_self · Понять где я → explore.wander
The 11-step ladder is now a unicode progress bar (▓ done, ▒ current,
░ pending) — same info, fits in one row.
- Advisor entries: one line each instead of two.
✓ wedged_60s → survive.flee 802t 1900ms
(outcome mark / trigger / target skill / tokens / latency)
- Improvements entries: one line each instead of two.
#1 P2 ×3 Add craft.iron-pickaxe skill
Description dropped from the row — use `node scripts/list-improvements.js`
for full text.
- Sections: 4 (was 6).
[header+story] · [activity | chat] · [advisor | improvements] · [footer]
Tested on a typical 1080p terminal — fits comfortably without zoom.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.1): real survival patterns — biome-aware scout, wedge-relocate, escape-pit-safe
Operator reported the bot wandered the same 50×50 patch for 2 hours
without making any progress toward food. Diagnosis showed three root
causes; this commit addresses all five open improvement_requests
the LLM (postmortem + tuner) flagged automatically.
Research basis (`Voyager`, `Plan4MC`, `GITM`, `Mindcraft`):
- Coverage / commit-to-cardinal exploration when local scan fails
- Biome-aware strategy switching using a static affordance table
- Wedge detector above the skill layer that triggers RELOCATE not
RETRY (per-skill stuck checks reset on re-entry — useless)
- Time-in-region bbox heuristic + need-duration AND skill-cycle gate
Concrete changes:
1. `runtime/goal/storyline.js`
- orient_self.completed: added timeout fallback (HP=full + session
>120s → done) so barren biomes don't block the bot on step 1.
Closes improvement #2 'Нет навыка оценки когда сменить район'.
- first_food.suggestSkill: now picks survive.scout-food (new) when
no passive mob is nearby; falls back to survive.acquire-food only
when something is in immediate range.
2. `runtime/biome-affordances.js` (new)
- Static table: 40+ biomes → {has_passive_mobs, has_trees,
has_water, has_crops, livable}.
- Unknown biomes return optimistic defaults to avoid regressions.
- Closes improvement #1 'Нет навыка целевого поиска еды по биому'.
3. `runtime/skills/scout-food.js` (new — survive.scout-food)
- Tiered strategy: biome check → scan 32 → scan 64 → commit a
cardinal for 200 blocks rescanning every 16. On cardinal
exhaustion, returns code:"exhausted" so the curriculum can
escalate to village.relocate.
- In barren biomes (desert/ocean/snowy_plains) the scan is
SKIPPED — bot walks straight toward the nearest neighbour
biome that affords passive mobs (8-direction biome probe at
radius 64).
4. `runtime/awareness/wedge-detector.js` (new)
- Rolling 10-min position bbox tracker. observe() called every
tick; isWedged() returns true when bbox<50 AND active need
unmet >5min AND skill cycles ≥3.
- markRelocationStarted() suppresses further wedge firings
until the bot has displaced ≥200b — prevents stack overflow
of relocate calls.
- Lives ABOVE the skill layer (in runtime/reflex.js), because
any per-skill stuck check resets on re-entry.
5. `runtime/skills/relocate.js` (new — village.relocate)
- 300-block walk in least-recently-used cardinal (per-incident
memory in ctx.recentRelocations).
- Re-paths every 32 blocks, soft-tolerates pathfinder failures
(3 consecutive throws → exit with code:"stuck_in_place").
- Closes improvement #2 + #4 ('low success rate trigger').
6. `runtime/skills/escape-pit-safe.js` (new — recovery.escape-pit-safe)
- Surveys 4 cardinals AND ceiling height before committing.
Picks the direction with most open blocks (≥3, no lava).
Falls through to pillar-up only if ceiling clear ≥4b. Returns
code:"no_strategy" if both blocked so curriculum can escalate
to relocate.
- Closes improvement #3 'Нет навыка для безопасного выхода'.
7. `runtime/reflex.js`
- Wedge detector wired before manifesto/storyline. If wedge.wedged
is true, dispatches village.relocate directly and returns —
bypasses every other branch.
- ctx.disableWedge flag for tests.
8. `runtime/coach/advisor-trigger.js`
- LLM provider outage backoff: 3 consecutive http_400 / timeout /
network_error → suppress advisor for 10 min. Today's TimeWeb
gpt-5.4-mini was 400'ing for an hour straight; we were spending
trigger budget on dead calls. Closes improvement implicit gap
in #5.
9. `runtime/bot.js`
- Tracks botSpawnedAt; snapshot._sessionMs exposed for storyline
orient_self timeout fallback.
Tests: 396 green (was 378, +18):
- runtime/biome-affordances.test.js — 8 tests
- runtime/awareness/wedge-detector.test.js — 9 tests
- runtime/goal/storyline.test.js — 1 new test (orient_self timeout)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(coach/trigger-tuner): crash after ~1h — runOnce is sync, not a Promise
The live bot died overnight with:
TypeError: runOnce(...).catch is not a function
at trigger-tuner.js:42 → [supervisor] child exited code=1
attach() wrapped the timer body as `runOnce().catch(...)` but
runOnce() returns a plain {ok, flagged, ...} object (pure SQL, no
await). The first tuner tick (60min after spawn) threw → killed the
whole bot process. Never surfaced before because the bot rarely ran
uninterrupted for a full hour during development.
Fix: guard the synchronous call with try/catch, matching how
persona/chatter.js already does its sync tick. (postmortem.drainOnce
and reflect.runOnce ARE async, so their .catch is correct — audited.)
Regression test added: captures the setInterval callback and invokes
it synchronously, asserting it does not throw.
Tests: 397 green.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(v0.3.1): mechanical food/stuck fixes — bot reaches the chicken now
The wedge wasn't only in the manifesto layer; several mechanical bugs
kept the bot in a dead random-walk:
- storyline / manifesto / curriculum: "local food" now means an edible
passive mob within <=32 blocks. A distant chicken or a cod no longer
fools the bot into dispatching acquire-food (which then fails on
no_path). Long-range food goes through scout-food instead.
- scout-food: partial approach to a target now counts as progress
(approached_target, e.g. moved:14); a blocked heading is NOT counted
as movement; added blind/tunnel fallback so it doesn't die when the
pathfinder can't route cleanly.
- acquire-food: on no_path it now also tries a blind/tunnel approach to
the animal; no_drop routes back into food scouting instead of giving
up.
- explore.far / relocate / flee: fewer false "done" results (micro-steps
no longer counted as success), more genuine escapes from stuck.
- scripts/show-story.js: live IPC now actually renders the current
storyline step.
Verification: scripts/lint-patch.js clean; npm test 404/404 green; bot
relaunched in tmux `pepa`. Live logs show real progress — bot switched
to survive.scout-food, approached the chicken (approached_target
moved:14), then reached survive.acquire-food: hunting chicken. Food
isn't fully closed yet but the remaining issue is concrete pickup/drop,
not dead random-walk.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(v0.3.1): sated bot stops chasing food + perf-leak + fuzzy improvement dedup
Third day of "bot just walks back and forth burning tokens". Root
causes were mechanical, not the manifesto:
1. SATED BOT CHASING FOOD (the big one)
Bot had food=17 (nearly full) but storyline first_food + manifesto
L1 required 2+ food ITEMS in inventory, so it looped scout-food /
acquire-food for hours instead of working. Now both treat a hunger
bar >= 14 (SATED_FOOD) as satisfied even with empty food inventory —
a full bot chops wood / makes tools and grabs food opportunistically,
only hard-pursuing food when actually hungry (< 14).
manifesto/needs.js foodDetect + goal/storyline.js first_food.completed.
2. perf_hooks MEMORY LEAK (overnight OOM suspect)
"MaxPerformanceEntryBufferExceededWarning: 1,000,001 measure entries".
mineflayer/pathfinder emit perf marks we never consume. Added a
60s reaper in bot.js (performance.clearMeasures/clearMarks). unref'd.
3. IMPROVEMENT QUEUE SELF-DUPLICATING
The LLM re-filed closed gaps with reworded titles (#5/#8/#9 were
dupes of implemented #1/#2/#3). Exact-title dedup missed them.
Replaced with token-set fuzzy match (isDuplicateTitle): jaccard>=0.75
OR >=3 shared meaningful tokens with jaccard>=0.5. Also: a re-filed
gap that's already implemented/rejected is NOT resurrected as a new
open row. Cleared all 5 open requests (now genuinely implemented).
Also confirmed (no change needed):
- canDig=true is a DELIBERATE codebase-wide choice ("without it the bot
gets permanently stuck", actions.js). The stale memory recommending
canDig=false is updated. ViaBackwards dig works partially (dug:1
moved:1.8 observed); false would trap the bot in every pit.
- scout-food already has blind/tunnel fallback + 12s step timeout
(operator's earlier edits) so trapped-pathfinder degrades instead of
hanging 30s.
Tests: 407 green (was 404). Updated needs/state/storyline tests for the
SATED_FOOD threshold; added fuzzy-dedup + tokenize/jaccard tests.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
7d96e44804 |
v0.3.0: Maslow + Awareness — self-learning bot with needs ladder, event-driven reflex, and TimeWeb fast advisor (#27)
* v0.3.0-rc.1: live skill registry + fast advisor scaffold
Roots out the v0.2.x failure mode: Pi-extracted lessons routinely named
hallucinated skill ids (relocate.surface, choose.safe.surface,
survive.shelter, gather.visible_log, …). All 47 Pi-lessons in the live DB
had applied_count=0 because normalisePreferSkill couldn't find them.
Fix:
1. runtime/skill-registry.js — single source of truth derived from
skills/index.js. Exports listSkillIds, isRegistered, and a
prompt-ready block (skillRegistryPrompt) grouped by namespace.
2. Pi prompts (coach/postmortem, coach/reflect) embed the live registry
with a "USE ONLY THESE, never invent" instruction. Lessons are
filtered at write-time too — anything not in the registry and not a
known mode name gets dropped.
3. coach/advice.js — normalisePreferSkill now returns null for unknown
ids, hardening consult() against any hallucinations that slip
through. Warn-logged for visibility.
Also lays the LLM substrate for the rest of v0.3.0:
- runtime/llm/provider.js — OpenAI-compatible chat client. Configured
via PEPA_FAST_LLM_{BASE_URL,API_KEY,MODEL,TIMEOUT_MS}. Safe no-op
unless API_KEY is set. Supports JSON-mode.
- runtime/coach/fast-advisor.js — tactical advisor tier (scaffold).
Exposes advise() that asks the fast LLM what to do RIGHT NOW when
the reflex is wedged/stuck. Rejects hallucinated skill ids using the
registry. Rate-limited 6/h, 30s cooldown. Not auto-triggered yet —
wired into reflex in rc.3 (awareness layer).
Tests: 279 green (+24 vs rc.3): 5 registry, 9 provider, 10 advisor.
See dev/v0.3.0/PLAN.md for the full iteration design (manifesto needs
ladder, event-driven awareness, skill pre-emption) and STATUS.md for
shipped/pending tracking.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* v0.3.0-rc.2: manifesto / needs ladder L0-L10
Adds an explicit hierarchical needs catalogue that the reflex consults
on every tick. The bot now pursues tangible intermediate goals (food,
wood tools, shelter, stone tools, ...) instead of inheriting whatever
the curriculum thought was "next".
Ladder:
L0 alive HP>5, food>0, not in lava, not panic-near hostile
L1 food ≥6 food items in inventory (or sated + any food)
L2 tools_wood wooden_pickaxe + wooden_axe + wooden_sword
L3 shelter_basic bed placed nearby or in inventory
L4 tools_stone stone-tier triplet
L5 armor_basic any chestplate (pursue=null until craft.leather-*
lands; ladder gracefully skips)
L6 food_security ≥16 food items
L7 tools_iron iron-tier triplet (pursue=gather.stone for now)
L8 armor_iron iron chestplate (pursue=null for now)
L9 village_seed bed + chest in nearby blocks
L10 village_full never detected, falls through to curriculum
Each need has detect(snapshot) → bool and pursue(snapshot) →
{skillId, args} | null. The ladder picks the LOWEST unsatisfied
pursuable need. Needs whose pursue is null get recorded as
blockedNeeds and the walk continues — no stalling on missing skills.
Wired into curriculumReflex: manifesto takes precedence over
curriculum.plan when it has a concrete suggestion. Tests can pass
ctx.disableManifesto=true to exercise the curriculum branch
in isolation (existing reflex tests keep passing this way).
Pi self-reflection prompt now includes
"activeNeed (Maslow ladder L0-L10): L2 tools_wood → gather.logs"
so Pi advises at the right level instead of giving generic guidance.
skillId returned by pursue() is validated against the live registry
(rc.1 plumbing) — manifesto cannot accidentally dispatch a
hallucinated skill name.
Tests: 315 green (was 279 on rc.1, +36 new):
- runtime/manifesto/needs.test.js — 24 tests (per-need detect/pursue,
helper sums)
- runtime/manifesto/state.test.js — 10 tests (ladder walk, hostile
takeover at L0, armor skipping, caching)
- runtime/reflex.test.js — 2 integration tests (manifesto overrides
curriculum plan; well-fed bot pursues tools_stone)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* v0.3.0-rc.3: event-driven awareness + skill pre-emption
Adds a reactive layer on top of the polling reflex. The bot now
notices environmental shocks (forced moves, HP plunges, hostile
spawns) within ~100ms instead of waiting for the next DISPATCH tick,
and the in-flight skill is preempted so the next reflex cycle can
re-plan against the current world state.
This is the rc that wires the "rc.1 plumbing + rc.2 manifesto" into
a feedback loop:
- awareness fires preempt → dispatch aborts
- reflex tick re-evaluates → manifesto walks the ladder
- new dispatch picks the right skill for the new world state
Pieces:
- runtime/awareness/events.js (new) — bot.on listeners:
- move: single-tick Δposition ≥ 5 blocks → forced_move flag + preempt
- health: HP drop ≥ 2 → health_plunge flag + preempt
- entitySpawn: hostile mob within 12 blocks → hostile_added + preempt
- blockUpdate: nearby block change → env_changed flag (no preempt,
throttled 800ms; otherwise gather skills would self-preempt
every dig)
- runtime/skills/index.js — RUNNER_CODES.PREEMPTED + raceWithAbort()
wraps every execute() against ctx.abortSignal. Existing skills get
preemption for free; they don't have to check the signal manually.
- runtime/bot.js:
- dispatchAction creates a fresh AbortController per dispatch and
stores it on reflexCtx.currentAbort
- attachAwareness fires controller.abort() when something disrupts
the active skill; runSkill returns code: "preempted" and the
reflex moves on
- reflexCtx.lastPreempt records the most recent shock
Tests: 332 green (was 315 on rc.2, +17 new):
- runtime/awareness/events.test.js — 12 tests (each event type +
thresholds + throttling + passive-mob filter)
- runtime/skills/contract.test.js — 3 abortSignal tests
(mid-flight, pre-armed, clean signal)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(.env): add PEPA_FAST_LLM_* placeholders for v0.3.0 fast advisor
Empty values keep the fast-advisor tier disabled (safe no-op). Fill
in BASE_URL + API_KEY + MODEL to enable. TimeWeb-style endpoint
example included.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(v0.3.0): rename fast-LLM env vars to TIMEWEB_* (match other projects)
Aligns with the user's other repos (proso) which use TIMEWEB_API_GROK /
TIMEWEB_URL_GROK. Single naming convention across projects avoids the
'which env var was it for this repo' mental tax.
PEPA_FAST_LLM_BASE_URL → TIMEWEB_BASE_URL
PEPA_FAST_LLM_API_KEY → TIMEWEB_API_KEY
PEPA_FAST_LLM_MODEL → TIMEWEB_MODEL
PEPA_FAST_LLM_TIMEOUT_MS → TIMEWEB_TIMEOUT_MS
Provider still works with any OpenAI-compatible endpoint — TimeWeb is
the default but the variable name doesn't lock us in. Tests + docs +
.env / .env.example updated.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(scripts): TimeWeb smoke test + bump default LLM timeout to 20s
scripts/check-timeweb.js — three probes: plain text, JSON mode, full
fast-advisor stack (registry injection + skill validation). Loads .env,
prints {ok, latency, reply preview} for each. Doesn't touch bot state.
Bumped DEFAULT_TIMEOUT_MS 8s → 20s in runtime/llm/provider.js. TimeWeb's
hosted agent endpoint takes 5-15s for the fast-advisor prompt
(registry block + snapshot context), so 8s was producing spurious
timeouts. OpenAI direct returns much faster; env var TIMEWEB_TIMEOUT_MS
overrides if needed.
Smoke verified live (PR #27 branch):
probe 1: 6.3s, plain prompt → "pepa hears you"
probe 2: 5.4s, JSON mode → {"alive":true,"name":"pepa"}
probe 3: 14.9s, advise() → action=switch_skill, skill=recovery.tunnel-out
(correct registered skill, sensible rationale — registry
injection successfully prevents hallucination)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.0): auto-trigger fast-advisor + token usage tracking
Closes the awareness → LLM → action loop that the rc.1/2/3 sequence
left as a followup. When the bot is wedged, looping, or just suffered
a preempt-then-retry, the reflex fires advise() in the background;
when the recommendation lands it overrides the next dispatch.
Async by design: advise() takes 5-15s on TimeWeb's hosted endpoint —
too slow for a synchronous reflex tick. tickAdvisor() is fire-and-
forget, the result lands on ctx.advisorRecommendation, and the *next*
tick reads and consumes it. Recommendations age out after 60s.
Components:
- runtime/coach/advisor-trigger.js — policy + async fire path
- tickAdvisor(ctx, {plannedSkillId}) checks three triggers:
1. wedged > 60s (no significant move)
2. last 4+ dispatches are the same skill AND it's planned again
3. preempt within last 30s + same skill being retried
- 90s trigger cooldown, single-in-flight guard
- consumeFreshRecommendation(ctx) reads/clears the cache
- runtime/reflex.js — curriculumReflex calls tickAdvisor() every tick
and consumes a fresh recommendation BEFORE dispatching. ctx flag
disableAdvisor=true for tests.
- runtime/bot.js — dispatchAction maintains a rolling 8-slot
reflexCtx.recentSkillIds for the loop-detection trigger.
Token usage:
- runtime/llm/provider.js — normaliseUsage() reads OpenAI/TimeWeb-
style {prompt_tokens, completion_tokens, total_tokens} from the
response. Returned on every complete() result and logged at info
level as "in=Nt/out=Mt".
- runtime/coach/fast-advisor.js — getUsageSnapshot() aggregates
total tokens across all calls in the session.
Measured on live TimeWeb endpoint (gpt-5.4-mini agent):
per call: ~705 input + 45 output = ~750 tokens
rate limit: 6 calls/hour
worst case at full budget: ~108K tokens/day
estimated cost (OpenAI gpt-5-mini reference price): ~$0.60/month
Well within any reasonable budget — model can run hot 24/7.
Smoke verified: scripts/check-timeweb.js probe 4 produces
trigger fired: true (wedged_90s)
recommendation: recovery.tunnel-out
rationale: "Stuck wedged for 90s; exploration is failing."
latency: 5302ms
Tests: 345 green (was 332, +13 advisor-trigger).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.0): paradigm shift — TimeWeb-only LLM + persistent advisor trail + improvement queue
This is the rc.4 batch the user requested:
1. Emergency triggers (low HP + close hostile, lava-under-foot)
bypass the long cooldown so the LLM is consulted BEFORE the bot
dies, not after.
2. Active manifesto need is now included in the advisor user prompt
— the LLM picks suggestions that satisfy the bot's current
concrete need (L2 tools_wood → "gather logs nearby" not
"explore further").
3. Every advisor recommendation is persisted to SQLite
(advisor_recommendations table) with full token usage. The
reflex marks 'applied=1' when it dispatches and updates
outcome_ok/code when the dispatch completes. Ground truth for
"is the LLM actually helping" lives in the DB, not in logs.
4. Pi CLI is OUT of every background loop. coach/postmortem and
coach/reflect now go through the same TimeWeb endpoint
fast-advisor uses, via the shared coach/llm-call.js helper.
Pi is reserved for manual operator commands.
5. The LLM (postmortem, reflect, advisor) can flag "structural
gaps" — missing skills/features the operator should implement.
These land in the new improvement_requests table. Dedup by
title bumps `votes` instead of inserting duplicates so the
queue doesn't bloat. Operator views via
`node scripts/list-improvements.js`.
6. A deterministic trigger-tuner runs hourly: reads 24h of
recommendation stats, flags triggers whose success rate is
below 25% (sample ≥ 5) or whose prompts are expensive (>1000
input tokens) with mediocre payoff. Improvements get
source="tuner", category="tuning". No LLM call.
New files:
runtime/coach/llm-call.js — askAnalytical() helper
runtime/coach/trigger-tuner.js — stats → improvements
runtime/coach/trigger-tuner.test.js
scripts/list-improvements.js — operator CLI
Schema additions:
advisor_recommendations: id, ts, trigger_reason, planned_skill,
recommended_skill, action, rationale, active_need, tokens_in,
tokens_out, latency_ms, applied, outcome_ok, outcome_code, outcome_at
improvement_requests: id, ts, source, category, title, description,
context, priority, status, duplicate_of, votes, implemented_at, notes
Renamed env-var consumers:
Pi-coach drainOnce({ askPi }) → drainOnce({ askAnalyticalFn? })
Pi-reflect runOnce({ askPi }) → runOnce({ askAnalyticalFn? })
bot.js attachCoach/attachReflect no longer pass askPi
attachTuner() added to bot.js spawn handler
lessons.source 'pi-coach' → 'timeweb-coach'
lessons.source 'pi-reflect' → 'timeweb-reflect'
Token cost measured live:
~705 input + 45 output = ~750 total per advisor call
worst case @ 6 calls/hour rate cap = ~108K tokens/day
OpenAI gpt-5-mini reference price: ~$0.60/month
Operator usage:
node scripts/list-improvements.js # open queue
node scripts/list-improvements.js --stats # advisor performance
node scripts/list-improvements.js --done 17 "shipped in 0.3.1"
node scripts/list-improvements.js --reject 18 "duplicate"
Tests: 360 green (was 332, +28 new).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
042793a53f |
docs+chore: v0.2.0 README refresh + auto-patch opens PR (not direct merge)
- README: update lede, architecture block, self-improvement section, status. Mentions knowledge.db, coach/advice loop, persona narration, and the new PR-based auto-patch flow. - scripts/auto-patch.js: replace cherry-pick-to-main with `git push` + `gh pr create`. The operator is now the only one who can merge into main (enforced by branch protection rules on the remote). Legacy direct-merge path remains behind PEPA_AUTO_PATCH_MERGE=cherry-pick for emergencies. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
4ae63dabe1 |
feat(runtime): v0.1.0 — adopt Voyager critic + Mindcraft modes/library/lint
Five concrete patterns from Voyager and Mindcraft, applied in our shape
without abandoning the git-as-evolution-substrate that makes pepa
distinct. Plus a first multi-agent surface so two bots from the same
repo can share intent.
1. runtime/critic.js (Voyager critic.txt)
- Spawns `pi -p` with a JSON-only critic prompt before a proposal is
written. {reasoning, success, critique}.
- success=true short-circuits the proposal (bot recovered between
detector tripping and now), saving Pi tokens on false positives.
- critique is spliced into the proposal body via attachCritique() so
the downstream auto-patcher has a sharp spec.
- Graceful: pi missing / timeout / unparseable JSON → proposal still
filed without the critic block.
2. scripts/lint-patch.js (Mindcraft coder._lintCode)
- Pre-flight gate between Pi commit and npm test: node --check, dynamic
import (catches missing named exports), regex extraction of
runSkill("id") calls cross-checked against the live registry.
- Cheaper than npm test, fails fast with a clear reason.
3. runtime/stuck-incident.renderActionTemplate (Voyager action_template.txt)
- All proposal bodies now follow the same fixed-section layout: Task /
Last result / Execution error / State / Metrics / Journal /
Scenarios / Critique / Fix / Edit scope / Forbidden.
4. runtime/skill-library.js (Mindcraft skill_library.getRelevantSkillDocs)
- Word-overlap ranking (Mindcraft's offline fallback) — zero deps,
deterministic. auto-patch.js injects top-3 similar skills into the
Pi prompt as "look at these patterns".
5. runtime/modes.js (Mindcraft modes.js)
- Declarative {name, interrupts, on, active, update(ctx)} chain that
runs BEFORE the curriculum each tick.
- Ships self_preservation (low HP → eat/flee), hunger (food<14 → eat),
night_shelter (night + bed in hand → sleep). Cleaner than ad-hoc
lastFleeAttempt cooldowns in reflex.js.
6. runtime/social/conversation.js + cmd:conv-say/conv-recent/conv-list
- File-JSONL topic channel so two bots from the same repo (different
usernames, different host dirs under state/) can append turns and
read peers. Skeleton — multi-agent collaboration on top later.
Differentiator preserved: every Pi-written skill still lands on main via
auto-patch.js (real git branch + smoke gate + cherry-pick). Voyager
keeps skills in a Chroma JSON, Mindcraft keeps them in RAM — pepa keeps
them as versioned source code reviewable in `git log`.
package.json: 0.0.1 → 0.1.0. 174/174 tests pass. README + AGENTS updated.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
69d1298fbd |
fix(auto-patch): lockfile-coordinated supervisor restarts
When Pi writes a multi-file runtime patch, the supervisor's file watcher can fire between two consecutive writes, kill the bot mid-edit, and load a half-saved file with a SyntaxError. Loop until the operator stops it. scripts/auto-patch.js now creates state/auto-patch.lock with its PID right after the branch checkout (before spawning pi -p), and removes it on every exit path. runtime/supervisor.js defers any watch-triggered restart while the lock holder is alive, polling every 2 s; once the lock drops it waits 1.5 s for the final write to settle, then runs \`node --check\` on the changed file and only restarts if it parses. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
82a8250a12 |
feat(auto-patch): enforce proposal editScope + run npm test before cherry-pick (#19)
Two safety rails on the unattended self-improvement loop:
1. scripts/edit-scope.js + .test.js: pure helpers that parse
`editScope: [...]` out of a proposal frontmatter (the field Phase 6
started writing) and validate a list of changed files against it.
13 tests covering null/missing/malformed frontmatter, directory
prefix matching, exact-file matching, default-scope fallback.
2. scripts/auto-patch.js:
- reads editScope from the proposal (falls back to ["runtime/"])
- injects the allowed paths into the Pi prompt so Pi knows the
boundaries up front
- validates the diff against scope + auto-allows any
runtime/**/*.test.js files Pi added
- runs `npm test` on the patched branch BEFORE cherry-picking;
refuses to land a patch that breaks the suite
Closes the "Smoke checks run before applying patch" item from
plans/autonomous-survival-bot-prd.md §7 Phase 6. npm test now 92/92.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
7797dd3d5a |
feat(runtime): fully autonomous self-healing — no operator approval (#10)
Operator feedback: "бот должен быть полностью автономным — сам себя улучшать и чинить, в этом и есть смысл; все что я вижу пока что он стоит на месте и кидает proposals на каждый чих — это кардинально не то что я хочу". Acted on: 1. Trigger filter — proposals only on real bugs. runtime/bot.js classifies failure detail into bug / timeout / feature-gap / other. The 5-in-a-row trigger fires only when the run contains a bug (TypeError / Cannot read / is not defined …) OR is entirely timeouts on the same operation. Feature gaps like "no reachable log within 32 blocks", "no food in inventory", "no bed in range", "no target in reach" are SKIPPED — the reflex layer routes around them (noTreesUntil → wander, etc). The LLM has no business patching code for missing inventory. Threshold raised 3 → 5 in a row. Cooldown unchanged (30 min). 2. Auto-apply, no operator-in-the-loop. New runtime/auto-improve.js polls proposals/ every 2s. When it sees a new .md and 10s have passed since first sighting (debounce), spawns scripts/auto-patch.js detached. New scripts/auto-patch.js: refuses on dirty tree, moves proposal pending → approved/, branches `auto/<slug>` off main, runs `pi -p` with 10-min timeout. If Pi committed AND every changed file is under runtime/ → cherry-picks onto main. Otherwise discards the branch. No push, no PR. Audit trail in state/<host>/proposals/approved/. Rate limit: 15-min cooldown between finished runs + 4/hour hard cap. 3. Auto-rollback on bad patches. runtime/supervisor.js: when MAX_RESTARTS_PER_MINUTE is exceeded AND `git log -1 HEAD` is younger than 15 min AND HEAD touched runtime/, runs `git reset --hard HEAD~1`. Up to MAX_ROLLBACKS=3 lifetime, then exits 1 for manual investigation. Restart counters are reset after a successful rollback so the next attempt isn't immediately killed. 4. current-task.json slim. No longer stores the full perception snapshot (was ~3 KB per write × every action). Position only — sufficient as a resume anchor. Slim snapshot still goes into the proposal markdown for context. docs/runtime.md — rewrote the self-improvement section: full flow diagram, classification rules, all rate-limit knobs, manual escape hatches kept but documented as rarely-needed. Also cleared 5 stale proposals from previous smoke tests so the first production run isn't burning Pi tokens on stale bugs that have since been fixed. Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2ecadd3bb2 |
fix(supervisor): pidfile lock prevents two supervisors racing on the nickname (#9)
User report 2026-05-25: launched 'npm run bot' fresh, MC server kicked
every login with "Игрок с данным никнеймом уже играет на сервере" and
the bot fell into a perpetual reconnect-then-kicked loop. Root cause:
a smoke-test supervisor from an earlier shell was still running in the
background, holding the pepa_bot session open. Two supervisors racing
on the same nickname is undefined behaviour from the server's side and
results in this exact failure mode.
Changes:
runtime/supervisor.js — acquires state/<host>/supervisor.pid before
spawning the child. If another supervisor is alive (kill -0 check), the
new one exits with a clear message telling the operator how to recover.
On SIGINT/SIGTERM/exit the lock is released; stale pidfiles are detected
when the recorded PID is no longer alive.
scripts/stop.sh — emergency cleanup helper:
- kills any supervisor or bot.js processes matching this repo
- removes pidfile + bot.sock
- reminds the operator to wait ~30s for the MC server to drop the old
session before re-launching
package.json — new `npm run stop` script.
Smoke-tested:
- first 'node runtime/supervisor.js' acquires lock, writes pid
- second call refuses with diagnostic message
- first SIGTERM ⇒ pidfile removed automatically
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
cd14bbf89a |
feat(runtime): state persistence + proposals + supervisor hot-restart (#6)
Closes the self-improvement loop end-to-end:
reflex fails 3× → proposal file → operator approves in TUI →
`npm run propose:apply <file>` spawns Pi on a feature branch →
Pi commits the patch → supervisor watches runtime/*.js and
restarts the child on change.
runtime/state-store.js — atomic current-task.json writes, daily diary
append, proposals/ + proposals/approved/ helpers.
runtime/bot.js:
- ctx.dispatch writes current-task.json on start and updates it on
completion / failure / throw.
- failure tracker: 3 consecutive same-label failures → writeProposal()
with the snapshot, labels, and a suggested-next-step section.
30-min cooldown prevents proposal spam.
- on startup, surfaces resume info (previous task + pending proposal
count); on death, clears current-task.json + writes diary line.
- new IPC commands: PROPOSAL_LATEST returns the newest pending
proposal body; PROPOSAL_APPROVE moves it to proposals/approved/.
tui/tui.tsx — status bar shows `[proposals N, press y]` badge when
bot.pendingProposals > 0. Hotkey 'y' opens the proposal panel; 'y'
approves, 'n'/Esc closes.
scripts/propose-apply.js — given an approved proposal filename, creates
a `feat/proposal-<slug>` branch and spawns `pi -p` with the proposal
+ repo-conventions prompt. Refuses on dirty tree. No auto-push, no
auto-merge — operator reviews the diff and decides.
runtime/supervisor.js — forks bot.js as a child, watches runtime/*.js,
restarts on file change or on child exit code 42. Rate-limited at 5
restarts/minute. SIGINT/SIGTERM forward cleanly. `npm run bot` now
goes through the supervisor; `npm run bot:bare` skips it.
Smoke-tested: supervisor spawned, bot connected to MC, spawned at
expected coords, diary line written, state cleanup on SIGTERM correct.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|