87 Commits
Author SHA1 Message Date
c910457817 v0.4.0 vNext — closed-loop world model + settlement contract (#29)
* feat(v0.4.0): vNext — closed-loop world model + settlement contract

Implements the vNext architecture from the research doc: demote the noisy
multi-rail planner in favour of a closed loop (world truth → invariant check)
plus a single utility-driven goal authority.

L1 services (fix no_drop / silent pathfinder hang first):
- InventoryLedger: diff-based "did I actually get it" verifier; acquire-food
  now confirms via ledger.gainedSince instead of the unreliable count/event.
- MotionService.gotoSafe: wall-clock timeout + progress watchdog +
  path_update(noPath/timeout) → structured {reached|stuck|timeout|nopath}.

L3 plan — unify the three competing rails (curriculum/manifesto/storyline):
- Settlement Contract: ordered M0–M9 milestones, each invariant-checked
  against an authoritative world view (early steps delegate to the proven
  curriculum; late game adds farming).
- InvariantChecker + predicate library; GoalManager selects the lowest unmet
  milestone via utility argmax (food-urgency preempts, DEPS-style).
- Wired into the scheduler: bot.js precomputes snapshot.contract; reflex.js
  consumes it in place of the storyline rail. Manifesto L0 still preempts.

Eval + robustness:
- Village Score (single 0..1 metric) on the snapshot + TUI "build" line.
- survive.dig-in skill + dusk_dig_in mode (exposed at night, no bed → cover).
- approach_block helper (GoalNear + lookAt, avoids GoalLookAtBlock #341).

+28 new tests (450 total green). LLM remains entirely off the tick path.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.4.0): finish vNext plan — anti-loop, skill-graph, worldDelta diff, flee→motion

Completes the remaining v0.4.0 plan items and one fix motivated by a live
in-game observation (flee hanging 30s against a persistent zombie).

- flee → MotionService.gotoSafe: structured {stuck|timeout|nopath} in ~4s with
  a blind-retreat fallback, instead of the observed 30s pathfinder hang + 3
  watchdog replans. Movements setup guarded so it is unit-testable.
- QW5 anti-loop (runtime/anti-loop.js): same skill failing >=3x in 5min →
  30min blacklist (reflex shouldSkip) + one-shot improvement_request
  (bot.js drainFired -> writeProposal).
- 4.1 closed-loop worldDelta: runSkill snapshots inventory before execute and
  attaches the real delta (_invObserved) to every successful result; opt-in
  skill.expectGain asserts the claimed gain or returns world_unchanged.
- 3.6 skill-graph (Plan4MC): declarative requires/produces for ~20 skills;
  prerequisitesMet/canRun/runnableFrontier; GoalManager annotates suggestions
  with blockedBy when prereqs are unmet.

+22 tests (472 total green). Live smoke confirmed dig-in works and no new
errors; flee loop is what this commit's flee migration addresses.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 11:29:39 +03:00
15b6c11002 v0.3.1: survival behaviour overhaul — storyline, biome-aware scout, wedge-relocate, food/perf fixes, monitor TUI (#28)
* docs(v0.3.1): PRD — LLM prompt cost optimization

Design-only commit; no runtime changes. Spec for the next patch iteration.

Goal: cut per-advise() input tokens from ~800 to ≤300, preserving the
LLM's ability to produce valid registered skill ids and useful rationale.

Five proposed changes ranked by impact:

  P1  Compact registry format (saves ~350t/call) — group by namespace,
      comma-list ids, drop human titles. Default mode for advisor;
      verbose mode kept for postmortem/reflect.
  P2  Need-scoped registry (~50t additional) — show LLM only skills
      relevant to the active Maslow need + always-available safety
      skills (survive.flee, pillar-up, recovery.tunnel-out, explore.*).
  P3  Snapshot pruning (~50t) — drop weather/experience/dimension/biome/
      players from the user prompt; the LLM doesn't consult them.
  P4  Prompt caching probe — check if TimeWeb passes through
      prompt_tokens_details.cached_tokens. If yes, restructure prefix
      to maximize cache hits (cached input is ~10x cheaper at OpenAI).
  P5  Per-trigger cost telemetry in scripts/list-improvements.js --stats:
      avg_in / avg_out / cost_₽ / share% per trigger_reason, using
      TIMEWEB_PRICE_IN_RUB_PER_M and TIMEWEB_PRICE_OUT_RUB_PER_M env.

Trigger: TimeWeb admin panel after first day of v0.3.0 live showed
34K tokens / day at low activity. At cap budget that projects to
~480₽/month (101₽/M in, 608₽/M out for gpt-5.4-mini). Manageable
but the savings are mostly free — repeated infra tokens, not signal.

All changes are additive; runtime behaviour stays the same. If the
LLM produces worse advice with the compact registry, flip back via a
single constant in fast-advisor.js.

Acceptance: re-run scripts/check-timeweb.js probe 3 — expect
tokens_in ≤ 300 (was ~800). Live for 1h, check --stats: avg_in ≤ 300
per trigger group. Existing 360 tests still green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.3.1): storyline — canonical Minecraft survival quest

The bot has been stuck in a loop for two days:
  acquire-food (fail: no nearby food) → explore.far → pillar-up (fail) → repeat

Diagnosis: manifesto + LLM advisor both correctly identify "you need
food" but neither expresses *what concretely to do next*. Manifesto is
a priority ladder (need-detection), not a narrative arc.

This commit adds the missing narrative layer — an ordered list of
operational steps that mirror the vanilla Minecraft survival path:

  1. orient_self      — Понять где я
  2. first_wood       — Собрать 8 поленьев
  3. crafting_basics  — Сделать верстак и палки
  4. first_tools      — Деревянные орудия
  5. first_food       — Найти первую еду
  6. shelter_minimal  — Простой шелтер с кроватью
  7. stone_tier       — Каменные орудия
  8. food_security    — Запас еды на 16+
  9. iron_age         — Железо и печь
  10. settle_base     — Постоянная база
  11. village_grow    — Развивать деревню (ongoing)

Each step has:
  - completed(snapshot) → bool — detects achievement from snapshot
  - suggestSkill(snapshot) → { skillId, args? } — concrete next dispatch
  - emergencyPause(snapshot) → bool — defers to manifesto L0 alive
    emergencies (low HP near hostile, lava under foot, food = 0)
  - narration_ru — chat-friendly Russian one-liner spoken on entry

Components:

- runtime/goal/storyline.js — 11-step canonical quest catalogue
- runtime/goal/state.js — pickCurrentStep(snapshot) walks the list,
  returns first non-completed step + its suggestion. 3s cache.
  Validates suggestSkill's skillId against the live registry.
- runtime/reflex.js — curriculumReflex dispatch priority is now:
    1. manifesto (L0 alive emergencies always win)
    2. storyline (concrete operational subgoal)
    3. curriculum plan (legacy fallback)
  Tests pass ctx.disableStoryline=true for isolation.
- runtime/bot.js — snapshot.storyStep populated each tick so
  chatter/advisor/reflect observers see the same view.
- runtime/coach/fast-advisor.js — buildUserPrompt now embeds the
  current step + its suggested skill, so LLM advice is anchored
  ("step 5 first_food, storyline wants survive.acquire-food, but
  recent dispatches show it's failing — try explore.far + scout").
- runtime/coach/advisor-trigger.js — forwards ctx.storyStep into
  advise() and logs step id at trigger time.
- runtime/coach/reflect.js — reflection prompt includes storyline
  progress so 30-min self-assessment is anchored.
- runtime/persona/chatter.js — narrates step.narration_ru on
  transition. Rate-limited via existing maybeNarrateRaw().

New operator CLI:

- scripts/show-story.js — fetches the live snapshot via IPC sock and
  prints step progress with ✓/→/ markers, current skill, inventory.
  Falls back to --plain catalogue view when bot offline.

Token cost impact: ~+30 input tokens per advise() call (one extra
line in user prompt). Trivial vs the value of grounding LLM advice
in a concrete narrative.

Operator usage:

  node scripts/show-story.js          # live progress + which step + why
  node scripts/show-story.js --plain  # static catalogue of all 11 steps

Tests: 376 green (was 360, +16 storyline tests).

Also in this branch (already committed): dev/v0.3.1/PRD.md —
LLM prompt cost optimization design doc.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(v0.3.1): storyline beats manifesto L1+ (only L0 alive emergencies override)

Found in live logs after the previous commit deployed:
  storyline: step 1/11: orient_self → explore.wander
  advisor-trigger: firing because wedged (planned=survive.acquire-food, ...)

Manifesto was still picking survive.acquire-food (L1 food) over the
storyline's orient_self → explore.wander. That's the wrong precedence —
storyline expresses a *concrete operational subgoal* and L1+ manifesto
needs are just "you'd benefit from food" priorities, not emergencies.

New dispatch precedence in curriculumReflex:
  1. manifesto L0 (alive emergencies: lava, low-HP+hostile, food=0)
  2. storyline (concrete narrative subgoal — beats L1+ manifesto)
  3. manifesto L1+ (fallback when storyline has no concrete suggestion)
  4. curriculum plan (legacy fallback)

This way the bot starts following the narrative arc even while
manifesto's L1 food is technically unsatisfied — orient_self runs to
completion before pursuing food explicitly. Storyline already handles
food as step 5 (first_food), so we're not skipping it.

Tests: 378 green (+2 priority-ordering tests):
- L0 manifesto emergency: upstream reflex (defend/modes) catches before
  curriculum dispatch
- storyline beats manifesto when both have suggestions: well-fed bot
  with logs → craft.planks (storyline crafting_basics), not gather.logs
  (manifesto L2)
- updated "manifesto fallback" test to require disableStoryline=true

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(tui): fullscreen monitor-only TUI (opencode-style)

Replaces the old tui/tui.tsx hotkey-heavy dashboard with a read-only
observability screen. Operator actions live in scripts/* now —
TUI is for watching, not driving.

Layout (top to bottom, all auto-resizing to terminal):
  1. Header     — MC/IPC status, pos, HP, food, day/night, hostiles
  2. Storyline  — current step + 11-step quest map (✓/→/○)
  3. Activity   — last N skill dispatches (colour by outcome)
  4. MC Chat    — last N chat lines (cyan for bot, yellow for players)
  5. Advisor    — last N LLM recommendations (trigger + outcome + tokens)
  6. Improvements — open requests from knowledge.improvement_requests
  7. Footer     — 24h token usage + cost in ₽ + q-to-quit

Data sources:
  - IPC sock: snapshot frames, log frames, chat frames (push)
  - SQLite knowledge.db: advisor_recommendations + improvement_requests
    polled every 5s (pull)

Token cost displayed live using TIMEWEB_PRICE_IN_RUB_PER_M /
TIMEWEB_PRICE_OUT_RUB_PER_M env vars (defaults: 101 / 608 for
gpt-5.4-mini).

Switches:
  - npm run tui          → new monitor (this file)
  - npm run tui:legacy   → old action-driven tui/tui.tsx (kept for now)

Implementation notes:
  - Uses ink + alternate-screen-buffer ANSI for proper "opencode-feel"
    fullscreen behaviour; restores prior terminal contents on quit.
  - Skips alt-screen and useInput when stdin/stdout isn't a TTY
    (smoke tests, piped output) — both gracefully degrade.
  - Stable React keys via per-event uid counter, avoids reconciler
    duplicate-key warnings as logs/chat/dispatches stream in.
  - Resize handled via 1s stdout-dimension poll, NOT direct
    'resize' listener (which conflicts with ink's own listener and
    triggers MaxListenersExceededWarning).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* ui(tui): compact 4-section monitor (was 6) — fits 1080p without zoom

Operator reported the TUI overflowed the screen unless terminal was
zoomed way out. The 11-step storyline list alone was eating ~13
rows, and each advisor/improvement entry took 2-3 rows. Now:

- Header + storyline collapsed into one panel (2 lines):
    line 1: pepa · ●MC ●IPC · 1m50s · pepa_bot · (697,61,702) · HP 20 · food 5 · ☀ · ⚔60(creeper@58b)
    line 2: story ▓▒░░░░░░░░░ 1/11 orient_self · Понять где я → explore.wander
  The 11-step ladder is now a unicode progress bar (▓ done, ▒ current,
  ░ pending) — same info, fits in one row.

- Advisor entries: one line each instead of two.
    ✓ wedged_60s → survive.flee  802t 1900ms
    (outcome mark / trigger / target skill / tokens / latency)

- Improvements entries: one line each instead of two.
    #1 P2 ×3  Add craft.iron-pickaxe skill
  Description dropped from the row — use `node scripts/list-improvements.js`
  for full text.

- Sections: 4 (was 6).
    [header+story] · [activity | chat] · [advisor | improvements] · [footer]

Tested on a typical 1080p terminal — fits comfortably without zoom.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.3.1): real survival patterns — biome-aware scout, wedge-relocate, escape-pit-safe

Operator reported the bot wandered the same 50×50 patch for 2 hours
without making any progress toward food. Diagnosis showed three root
causes; this commit addresses all five open improvement_requests
the LLM (postmortem + tuner) flagged automatically.

Research basis (`Voyager`, `Plan4MC`, `GITM`, `Mindcraft`):
  - Coverage / commit-to-cardinal exploration when local scan fails
  - Biome-aware strategy switching using a static affordance table
  - Wedge detector above the skill layer that triggers RELOCATE not
    RETRY (per-skill stuck checks reset on re-entry — useless)
  - Time-in-region bbox heuristic + need-duration AND skill-cycle gate

Concrete changes:

1. `runtime/goal/storyline.js`
   - orient_self.completed: added timeout fallback (HP=full + session
     >120s → done) so barren biomes don't block the bot on step 1.
     Closes improvement #2 'Нет навыка оценки когда сменить район'.
   - first_food.suggestSkill: now picks survive.scout-food (new) when
     no passive mob is nearby; falls back to survive.acquire-food only
     when something is in immediate range.

2. `runtime/biome-affordances.js` (new)
   - Static table: 40+ biomes → {has_passive_mobs, has_trees,
     has_water, has_crops, livable}.
   - Unknown biomes return optimistic defaults to avoid regressions.
   - Closes improvement #1 'Нет навыка целевого поиска еды по биому'.

3. `runtime/skills/scout-food.js` (new — survive.scout-food)
   - Tiered strategy: biome check → scan 32 → scan 64 → commit a
     cardinal for 200 blocks rescanning every 16. On cardinal
     exhaustion, returns code:"exhausted" so the curriculum can
     escalate to village.relocate.
   - In barren biomes (desert/ocean/snowy_plains) the scan is
     SKIPPED — bot walks straight toward the nearest neighbour
     biome that affords passive mobs (8-direction biome probe at
     radius 64).

4. `runtime/awareness/wedge-detector.js` (new)
   - Rolling 10-min position bbox tracker. observe() called every
     tick; isWedged() returns true when bbox<50 AND active need
     unmet >5min AND skill cycles ≥3.
   - markRelocationStarted() suppresses further wedge firings
     until the bot has displaced ≥200b — prevents stack overflow
     of relocate calls.
   - Lives ABOVE the skill layer (in runtime/reflex.js), because
     any per-skill stuck check resets on re-entry.

5. `runtime/skills/relocate.js` (new — village.relocate)
   - 300-block walk in least-recently-used cardinal (per-incident
     memory in ctx.recentRelocations).
   - Re-paths every 32 blocks, soft-tolerates pathfinder failures
     (3 consecutive throws → exit with code:"stuck_in_place").
   - Closes improvement #2 + #4 ('low success rate trigger').

6. `runtime/skills/escape-pit-safe.js` (new — recovery.escape-pit-safe)
   - Surveys 4 cardinals AND ceiling height before committing.
     Picks the direction with most open blocks (≥3, no lava).
     Falls through to pillar-up only if ceiling clear ≥4b. Returns
     code:"no_strategy" if both blocked so curriculum can escalate
     to relocate.
   - Closes improvement #3 'Нет навыка для безопасного выхода'.

7. `runtime/reflex.js`
   - Wedge detector wired before manifesto/storyline. If wedge.wedged
     is true, dispatches village.relocate directly and returns —
     bypasses every other branch.
   - ctx.disableWedge flag for tests.

8. `runtime/coach/advisor-trigger.js`
   - LLM provider outage backoff: 3 consecutive http_400 / timeout /
     network_error → suppress advisor for 10 min. Today's TimeWeb
     gpt-5.4-mini was 400'ing for an hour straight; we were spending
     trigger budget on dead calls. Closes improvement implicit gap
     in #5.

9. `runtime/bot.js`
   - Tracks botSpawnedAt; snapshot._sessionMs exposed for storyline
     orient_self timeout fallback.

Tests: 396 green (was 378, +18):
  - runtime/biome-affordances.test.js — 8 tests
  - runtime/awareness/wedge-detector.test.js — 9 tests
  - runtime/goal/storyline.test.js — 1 new test (orient_self timeout)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(coach/trigger-tuner): crash after ~1h — runOnce is sync, not a Promise

The live bot died overnight with:
  TypeError: runOnce(...).catch is not a function
  at trigger-tuner.js:42  →  [supervisor] child exited code=1

attach() wrapped the timer body as `runOnce().catch(...)` but
runOnce() returns a plain {ok, flagged, ...} object (pure SQL, no
await). The first tuner tick (60min after spawn) threw → killed the
whole bot process. Never surfaced before because the bot rarely ran
uninterrupted for a full hour during development.

Fix: guard the synchronous call with try/catch, matching how
persona/chatter.js already does its sync tick. (postmortem.drainOnce
and reflect.runOnce ARE async, so their .catch is correct — audited.)

Regression test added: captures the setInterval callback and invokes
it synchronously, asserting it does not throw.

Tests: 397 green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(v0.3.1): mechanical food/stuck fixes — bot reaches the chicken now

The wedge wasn't only in the manifesto layer; several mechanical bugs
kept the bot in a dead random-walk:

- storyline / manifesto / curriculum: "local food" now means an edible
  passive mob within <=32 blocks. A distant chicken or a cod no longer
  fools the bot into dispatching acquire-food (which then fails on
  no_path). Long-range food goes through scout-food instead.

- scout-food: partial approach to a target now counts as progress
  (approached_target, e.g. moved:14); a blocked heading is NOT counted
  as movement; added blind/tunnel fallback so it doesn't die when the
  pathfinder can't route cleanly.

- acquire-food: on no_path it now also tries a blind/tunnel approach to
  the animal; no_drop routes back into food scouting instead of giving
  up.

- explore.far / relocate / flee: fewer false "done" results (micro-steps
  no longer counted as success), more genuine escapes from stuck.

- scripts/show-story.js: live IPC now actually renders the current
  storyline step.

Verification: scripts/lint-patch.js clean; npm test 404/404 green; bot
relaunched in tmux `pepa`. Live logs show real progress — bot switched
to survive.scout-food, approached the chicken (approached_target
moved:14), then reached survive.acquire-food: hunting chicken. Food
isn't fully closed yet but the remaining issue is concrete pickup/drop,
not dead random-walk.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(v0.3.1): sated bot stops chasing food + perf-leak + fuzzy improvement dedup

Third day of "bot just walks back and forth burning tokens". Root
causes were mechanical, not the manifesto:

1. SATED BOT CHASING FOOD (the big one)
   Bot had food=17 (nearly full) but storyline first_food + manifesto
   L1 required 2+ food ITEMS in inventory, so it looped scout-food /
   acquire-food for hours instead of working. Now both treat a hunger
   bar >= 14 (SATED_FOOD) as satisfied even with empty food inventory —
   a full bot chops wood / makes tools and grabs food opportunistically,
   only hard-pursuing food when actually hungry (< 14).
   manifesto/needs.js foodDetect + goal/storyline.js first_food.completed.

2. perf_hooks MEMORY LEAK (overnight OOM suspect)
   "MaxPerformanceEntryBufferExceededWarning: 1,000,001 measure entries".
   mineflayer/pathfinder emit perf marks we never consume. Added a
   60s reaper in bot.js (performance.clearMeasures/clearMarks). unref'd.

3. IMPROVEMENT QUEUE SELF-DUPLICATING
   The LLM re-filed closed gaps with reworded titles (#5/#8/#9 were
   dupes of implemented #1/#2/#3). Exact-title dedup missed them.
   Replaced with token-set fuzzy match (isDuplicateTitle): jaccard>=0.75
   OR >=3 shared meaningful tokens with jaccard>=0.5. Also: a re-filed
   gap that's already implemented/rejected is NOT resurrected as a new
   open row. Cleared all 5 open requests (now genuinely implemented).

Also confirmed (no change needed):
- canDig=true is a DELIBERATE codebase-wide choice ("without it the bot
  gets permanently stuck", actions.js). The stale memory recommending
  canDig=false is updated. ViaBackwards dig works partially (dug:1
  moved:1.8 observed); false would trap the bot in every pit.
- scout-food already has blind/tunnel fallback + 12s step timeout
  (operator's earlier edits) so trapped-pathfinder degrades instead of
  hanging 30s.

Tests: 407 green (was 404). Updated needs/state/storyline tests for the
SATED_FOOD threshold; added fuzzy-dedup + tokenize/jaccard tests.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 09:43:57 +03:00
7d96e44804 v0.3.0: Maslow + Awareness — self-learning bot with needs ladder, event-driven reflex, and TimeWeb fast advisor (#27)
* v0.3.0-rc.1: live skill registry + fast advisor scaffold

Roots out the v0.2.x failure mode: Pi-extracted lessons routinely named
hallucinated skill ids (relocate.surface, choose.safe.surface,
survive.shelter, gather.visible_log, …). All 47 Pi-lessons in the live DB
had applied_count=0 because normalisePreferSkill couldn't find them.

Fix:
1. runtime/skill-registry.js — single source of truth derived from
   skills/index.js. Exports listSkillIds, isRegistered, and a
   prompt-ready block (skillRegistryPrompt) grouped by namespace.
2. Pi prompts (coach/postmortem, coach/reflect) embed the live registry
   with a "USE ONLY THESE, never invent" instruction. Lessons are
   filtered at write-time too — anything not in the registry and not a
   known mode name gets dropped.
3. coach/advice.js — normalisePreferSkill now returns null for unknown
   ids, hardening consult() against any hallucinations that slip
   through. Warn-logged for visibility.

Also lays the LLM substrate for the rest of v0.3.0:

- runtime/llm/provider.js — OpenAI-compatible chat client. Configured
  via PEPA_FAST_LLM_{BASE_URL,API_KEY,MODEL,TIMEOUT_MS}. Safe no-op
  unless API_KEY is set. Supports JSON-mode.
- runtime/coach/fast-advisor.js — tactical advisor tier (scaffold).
  Exposes advise() that asks the fast LLM what to do RIGHT NOW when
  the reflex is wedged/stuck. Rejects hallucinated skill ids using the
  registry. Rate-limited 6/h, 30s cooldown. Not auto-triggered yet —
  wired into reflex in rc.3 (awareness layer).

Tests: 279 green (+24 vs rc.3): 5 registry, 9 provider, 10 advisor.

See dev/v0.3.0/PLAN.md for the full iteration design (manifesto needs
ladder, event-driven awareness, skill pre-emption) and STATUS.md for
shipped/pending tracking.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* v0.3.0-rc.2: manifesto / needs ladder L0-L10

Adds an explicit hierarchical needs catalogue that the reflex consults
on every tick. The bot now pursues tangible intermediate goals (food,
wood tools, shelter, stone tools, ...) instead of inheriting whatever
the curriculum thought was "next".

Ladder:
  L0  alive          HP>5, food>0, not in lava, not panic-near hostile
  L1  food           ≥6 food items in inventory (or sated + any food)
  L2  tools_wood     wooden_pickaxe + wooden_axe + wooden_sword
  L3  shelter_basic  bed placed nearby or in inventory
  L4  tools_stone    stone-tier triplet
  L5  armor_basic    any chestplate (pursue=null until craft.leather-*
                     lands; ladder gracefully skips)
  L6  food_security  ≥16 food items
  L7  tools_iron     iron-tier triplet (pursue=gather.stone for now)
  L8  armor_iron     iron chestplate (pursue=null for now)
  L9  village_seed   bed + chest in nearby blocks
  L10 village_full   never detected, falls through to curriculum

Each need has detect(snapshot) → bool and pursue(snapshot) →
{skillId, args} | null. The ladder picks the LOWEST unsatisfied
pursuable need. Needs whose pursue is null get recorded as
blockedNeeds and the walk continues — no stalling on missing skills.

Wired into curriculumReflex: manifesto takes precedence over
curriculum.plan when it has a concrete suggestion. Tests can pass
ctx.disableManifesto=true to exercise the curriculum branch
in isolation (existing reflex tests keep passing this way).

Pi self-reflection prompt now includes
"activeNeed (Maslow ladder L0-L10): L2 tools_wood → gather.logs"
so Pi advises at the right level instead of giving generic guidance.

skillId returned by pursue() is validated against the live registry
(rc.1 plumbing) — manifesto cannot accidentally dispatch a
hallucinated skill name.

Tests: 315 green (was 279 on rc.1, +36 new):
- runtime/manifesto/needs.test.js — 24 tests (per-need detect/pursue,
  helper sums)
- runtime/manifesto/state.test.js — 10 tests (ladder walk, hostile
  takeover at L0, armor skipping, caching)
- runtime/reflex.test.js — 2 integration tests (manifesto overrides
  curriculum plan; well-fed bot pursues tools_stone)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* v0.3.0-rc.3: event-driven awareness + skill pre-emption

Adds a reactive layer on top of the polling reflex. The bot now
notices environmental shocks (forced moves, HP plunges, hostile
spawns) within ~100ms instead of waiting for the next DISPATCH tick,
and the in-flight skill is preempted so the next reflex cycle can
re-plan against the current world state.

This is the rc that wires the "rc.1 plumbing + rc.2 manifesto" into
a feedback loop:
  - awareness fires preempt → dispatch aborts
  - reflex tick re-evaluates → manifesto walks the ladder
  - new dispatch picks the right skill for the new world state

Pieces:

- runtime/awareness/events.js (new) — bot.on listeners:
  - move: single-tick Δposition ≥ 5 blocks → forced_move flag + preempt
  - health: HP drop ≥ 2 → health_plunge flag + preempt
  - entitySpawn: hostile mob within 12 blocks → hostile_added + preempt
  - blockUpdate: nearby block change → env_changed flag (no preempt,
    throttled 800ms; otherwise gather skills would self-preempt
    every dig)

- runtime/skills/index.js — RUNNER_CODES.PREEMPTED + raceWithAbort()
  wraps every execute() against ctx.abortSignal. Existing skills get
  preemption for free; they don't have to check the signal manually.

- runtime/bot.js:
  - dispatchAction creates a fresh AbortController per dispatch and
    stores it on reflexCtx.currentAbort
  - attachAwareness fires controller.abort() when something disrupts
    the active skill; runSkill returns code: "preempted" and the
    reflex moves on
  - reflexCtx.lastPreempt records the most recent shock

Tests: 332 green (was 315 on rc.2, +17 new):
- runtime/awareness/events.test.js — 12 tests (each event type +
  thresholds + throttling + passive-mob filter)
- runtime/skills/contract.test.js — 3 abortSignal tests
  (mid-flight, pre-armed, clean signal)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(.env): add PEPA_FAST_LLM_* placeholders for v0.3.0 fast advisor

Empty values keep the fast-advisor tier disabled (safe no-op). Fill
in BASE_URL + API_KEY + MODEL to enable. TimeWeb-style endpoint
example included.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(v0.3.0): rename fast-LLM env vars to TIMEWEB_* (match other projects)

Aligns with the user's other repos (proso) which use TIMEWEB_API_GROK /
TIMEWEB_URL_GROK. Single naming convention across projects avoids the
'which env var was it for this repo' mental tax.

  PEPA_FAST_LLM_BASE_URL → TIMEWEB_BASE_URL
  PEPA_FAST_LLM_API_KEY  → TIMEWEB_API_KEY
  PEPA_FAST_LLM_MODEL    → TIMEWEB_MODEL
  PEPA_FAST_LLM_TIMEOUT_MS → TIMEWEB_TIMEOUT_MS

Provider still works with any OpenAI-compatible endpoint — TimeWeb is
the default but the variable name doesn't lock us in. Tests + docs +
.env / .env.example updated.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(scripts): TimeWeb smoke test + bump default LLM timeout to 20s

scripts/check-timeweb.js — three probes: plain text, JSON mode, full
fast-advisor stack (registry injection + skill validation). Loads .env,
prints {ok, latency, reply preview} for each. Doesn't touch bot state.

Bumped DEFAULT_TIMEOUT_MS 8s → 20s in runtime/llm/provider.js. TimeWeb's
hosted agent endpoint takes 5-15s for the fast-advisor prompt
(registry block + snapshot context), so 8s was producing spurious
timeouts. OpenAI direct returns much faster; env var TIMEWEB_TIMEOUT_MS
overrides if needed.

Smoke verified live (PR #27 branch):
  probe 1: 6.3s, plain prompt → "pepa hears you"
  probe 2: 5.4s, JSON mode → {"alive":true,"name":"pepa"}
  probe 3: 14.9s, advise() → action=switch_skill, skill=recovery.tunnel-out
           (correct registered skill, sensible rationale — registry
           injection successfully prevents hallucination)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.3.0): auto-trigger fast-advisor + token usage tracking

Closes the awareness → LLM → action loop that the rc.1/2/3 sequence
left as a followup. When the bot is wedged, looping, or just suffered
a preempt-then-retry, the reflex fires advise() in the background;
when the recommendation lands it overrides the next dispatch.

Async by design: advise() takes 5-15s on TimeWeb's hosted endpoint —
too slow for a synchronous reflex tick. tickAdvisor() is fire-and-
forget, the result lands on ctx.advisorRecommendation, and the *next*
tick reads and consumes it. Recommendations age out after 60s.

Components:

- runtime/coach/advisor-trigger.js — policy + async fire path
  - tickAdvisor(ctx, {plannedSkillId}) checks three triggers:
    1. wedged > 60s (no significant move)
    2. last 4+ dispatches are the same skill AND it's planned again
    3. preempt within last 30s + same skill being retried
  - 90s trigger cooldown, single-in-flight guard
  - consumeFreshRecommendation(ctx) reads/clears the cache
- runtime/reflex.js — curriculumReflex calls tickAdvisor() every tick
  and consumes a fresh recommendation BEFORE dispatching. ctx flag
  disableAdvisor=true for tests.
- runtime/bot.js — dispatchAction maintains a rolling 8-slot
  reflexCtx.recentSkillIds for the loop-detection trigger.

Token usage:

- runtime/llm/provider.js — normaliseUsage() reads OpenAI/TimeWeb-
  style {prompt_tokens, completion_tokens, total_tokens} from the
  response. Returned on every complete() result and logged at info
  level as "in=Nt/out=Mt".
- runtime/coach/fast-advisor.js — getUsageSnapshot() aggregates
  total tokens across all calls in the session.

Measured on live TimeWeb endpoint (gpt-5.4-mini agent):
  per call: ~705 input + 45 output = ~750 tokens
  rate limit: 6 calls/hour
  worst case at full budget: ~108K tokens/day
  estimated cost (OpenAI gpt-5-mini reference price): ~$0.60/month

Well within any reasonable budget — model can run hot 24/7.

Smoke verified: scripts/check-timeweb.js probe 4 produces
  trigger fired: true (wedged_90s)
  recommendation: recovery.tunnel-out
  rationale: "Stuck wedged for 90s; exploration is failing."
  latency: 5302ms

Tests: 345 green (was 332, +13 advisor-trigger).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.3.0): paradigm shift — TimeWeb-only LLM + persistent advisor trail + improvement queue

This is the rc.4 batch the user requested:

  1. Emergency triggers (low HP + close hostile, lava-under-foot)
     bypass the long cooldown so the LLM is consulted BEFORE the bot
     dies, not after.
  2. Active manifesto need is now included in the advisor user prompt
     — the LLM picks suggestions that satisfy the bot's current
     concrete need (L2 tools_wood → "gather logs nearby" not
     "explore further").
  3. Every advisor recommendation is persisted to SQLite
     (advisor_recommendations table) with full token usage. The
     reflex marks 'applied=1' when it dispatches and updates
     outcome_ok/code when the dispatch completes. Ground truth for
     "is the LLM actually helping" lives in the DB, not in logs.
  4. Pi CLI is OUT of every background loop. coach/postmortem and
     coach/reflect now go through the same TimeWeb endpoint
     fast-advisor uses, via the shared coach/llm-call.js helper.
     Pi is reserved for manual operator commands.
  5. The LLM (postmortem, reflect, advisor) can flag "structural
     gaps" — missing skills/features the operator should implement.
     These land in the new improvement_requests table. Dedup by
     title bumps `votes` instead of inserting duplicates so the
     queue doesn't bloat. Operator views via
     `node scripts/list-improvements.js`.
  6. A deterministic trigger-tuner runs hourly: reads 24h of
     recommendation stats, flags triggers whose success rate is
     below 25% (sample ≥ 5) or whose prompts are expensive (>1000
     input tokens) with mediocre payoff. Improvements get
     source="tuner", category="tuning". No LLM call.

New files:
  runtime/coach/llm-call.js        — askAnalytical() helper
  runtime/coach/trigger-tuner.js   — stats → improvements
  runtime/coach/trigger-tuner.test.js
  scripts/list-improvements.js     — operator CLI

Schema additions:
  advisor_recommendations: id, ts, trigger_reason, planned_skill,
    recommended_skill, action, rationale, active_need, tokens_in,
    tokens_out, latency_ms, applied, outcome_ok, outcome_code, outcome_at
  improvement_requests: id, ts, source, category, title, description,
    context, priority, status, duplicate_of, votes, implemented_at, notes

Renamed env-var consumers:
  Pi-coach drainOnce({ askPi })   → drainOnce({ askAnalyticalFn? })
  Pi-reflect runOnce({ askPi })   → runOnce({ askAnalyticalFn? })
  bot.js attachCoach/attachReflect no longer pass askPi
  attachTuner() added to bot.js spawn handler
  lessons.source 'pi-coach'   → 'timeweb-coach'
  lessons.source 'pi-reflect' → 'timeweb-reflect'

Token cost measured live:
  ~705 input + 45 output = ~750 total per advisor call
  worst case @ 6 calls/hour rate cap = ~108K tokens/day
  OpenAI gpt-5-mini reference price: ~$0.60/month

Operator usage:
  node scripts/list-improvements.js                # open queue
  node scripts/list-improvements.js --stats        # advisor performance
  node scripts/list-improvements.js --done 17 "shipped in 0.3.1"
  node scripts/list-improvements.js --reject 18 "duplicate"

Tests: 360 green (was 332, +28 new).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 19:37:56 +03:00
85f7ad90c8 docs: dev/v0.2.0/STATUS.md as session-continuation context (#23)
Single source of truth for the v0.2.0 iteration:
- what shipped per rc (rc.1/rc.2/rc.3), with links to PRs/commits
- live DB snapshot from 2026-05-27 ~15:30
- 7 known issues / followups ranked by impact (rc.4 candidates)
- quick-start checklist for the next session

Also:
- docs/v0.2.0-self-learning.md: phase checklist updated to reflect
  shipped state and points at dev/v0.2.0/STATUS.md for live data
- README status bullet refreshed with rc.1/2/3 summary

Folder convention: dev/v<version>/ for per-version development notes.
Anything still in flight or candidate for the next iteration lives
here; design docs that pre-date the iteration stay in docs/.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 19:37:33 +03:00
865aae1213 v0.2.0-rc.3: pillar-up escape + advice in fallback + danger POI (#22)
* v0.2.0-rc.3: pillar-up escape + advice everywhere + danger POI

Closes the gap rc.2 left open. Live observation showed:
- Pi-coach extracted 5 high-quality lessons (do not explore.far at
  night near zombies, etc.) but none of them fired (applied_count=0
  across the board). Root cause: dispatcher consulted advice only on
  the main curriculum path; the bot was falling into the wander/
  explore.far FALLBACK after each gather attempt bailed, which
  bypassed consult().
- Bot was wedged in a pit on (608, 90) with stone walls. recovery.
  tunnel-out kept failing ("Digging aborted") because mining stone
  with fists takes ~10s/block; pathfinder watchdog kills it.

This patch:

1. survive.pillar-up (runtime/skills/pillar-up.js) — new escape skill.
   Places a placeable block under the bot and jumps onto it; repeats
   up to 8 steps. No pickaxe required. Works in dirt/cobble/planks/
   sand/gravel/wool/etc. The bot's vertical exit from any pit it can
   stand in.

2. Wedged-emergency reflex (runtime/reflex.js). At the top of
   curriculumReflex, if noProgressReason is wedged-like AND position
   hasn't shifted ≥16 blocks in 60s AND no hostile in 6m AND pillar
   block in inventory → dispatch survive.pillar-up. 2-min cooldown
   between attempts.

3. consult() now also runs on the WANDER/explore.far fallback path
   (runtime/reflex.js curriculumReflex). Pi-coach lessons can finally
   take effect. If the fallback skill is overridden to a non-eligible
   skill but the bot has a placeable block, falls back to pillar-up.
   Outcomes feed reportAdviceOutcome so confidence stays grounded.

4. recordPOI("danger") on death (runtime/coach/postmortem.js). Spatial
   memory now flags where the bot died, expires after 6h. POI table
   was empty in rc.2.

5. SAFE_OVERRIDES extended (runtime/coach/advice.js): adds
   survive.pillar-up and village.choose-base so coach lessons can
   route there.

Tests: 255/255 green (+9 pillar-up).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* v0.2.0-rc.3 fixup: relax wedged-escape trigger

Drop the WEDGED_REASONS check — noProgressReason is a string that
may or may not be set when the bot is stuck. Fire pillar-up purely on
"no horizontal progress ≥ 60s, no hostile in 6m, placeable block in
inv". Pillar-up is a constructive no-op when it's not needed (places
one dirt under self) so the false-positive cost is small.

Live observation: rc.3 was deployed and bot was wedged with tunnel-out
repeatedly aborted on stone, but wedged-escape never fired because
the runtime's noProgressReason wasn't in my whitelist. Removing the
gate lets the trigger actually engage.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 15:31:21 +03:00
e84148d189 v0.2.0-rc.2: P0 hardening — Pi headless, test state isolation, advice fixes (#21)
P0 (correctness):

1. PEPA_HEADLESS=1 guard in extensions/mineflayer-bridge.ts. When `pi -p`
   spawns a subprocess (banter, coach, planner, reflect, auto-patch), the
   bridge no longer attempts a second MC connect — the hybrid runtime
   already owns the nickname. runtime/pi-bridge.js sets the env var on
   every spawn. Root cause of the "two pepa_bot's racing for the slot"
   bug seen in reply-pi stderr.

2. Test state isolation in runtime/config.js. When running under the node
   test runner (detected via execArgv/argv) — or when PEPA_STATE_DIR is
   set — stateDir redirects to /tmp/pepa-test-state-<pid>/. log.js,
   scenario-memory, world-journal, and knowledge.db all follow.
   `npm test` no longer pollutes live scenarios.jsonl, world-journal.jsonl,
   or daily log files. Verified empirically: post-fix run added 0 test
   rows to the live scenarios file. Cleaned ~550 historical test rows
   from live state in the same change.

3. defendReflex outcome reporting (runtime/reflex.js). Previously a
   creeper-rule override marked the lesson succeeded=false BEFORE the
   flee skill returned. Now dispatchDefendFlee accepts {lessonId} and
   the onComplete fires reportAdviceOutcome with the actual flee result.

4. Mode-name → skill-id translation in runtime/coach/advice.js. Pi-coach
   occasionally returns prefer_skill values that are mode names
   ("night_shelter", "self_preservation", "hunger"). normalisePreferSkill
   maps these to SAFE_OVERRIDES entries before dispatch. Also handles
   "tunnel-out", "survive_flee", "survive flee" shapes.

New behavior:

5. Self-reflection loop (runtime/coach/reflect.js). Every 30 min, the
   bot asks Pi: "Are you making progress, or stuck in a loop? What
   should you do differently?" Pi answers with a verdict
   (progress/loop/recovering/idle/emergency), summary, next-action, and
   0-N new lessons. The reflection is written to
   state/<host>/reflections/<ts>.md and lessons land in the DB with
   source="pi-reflect". Rate-limited to 2 calls/hour. Wired through
   bot.js with the existing askPi + lastSnapshot accessor.

Tests: 246/246 green (+9 new: 6 advice mode-name + 4 reflect).

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 14:10:28 +03:00
mayatnikovandClaude Opus 4.7 042793a53f docs+chore: v0.2.0 README refresh + auto-patch opens PR (not direct merge)
- README: update lede, architecture block, self-improvement section, status.
  Mentions knowledge.db, coach/advice loop, persona narration, and the new
  PR-based auto-patch flow.
- scripts/auto-patch.js: replace cherry-pick-to-main with `git push` +
  `gh pr create`. The operator is now the only one who can merge into main
  (enforced by branch protection rules on the remote). Legacy direct-merge
  path remains behind PEPA_AUTO_PATCH_MERGE=cherry-pick for emergencies.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 13:31:27 +03:00
mayatnikov 9c2c56f30e v0.2.0-rc.1: self-learning agent — knowledge DB + post-mortem coach + persona
Major iteration. The bot now:

1. Maintains a per-server SQLite knowledge.db with structured memory:
   recipes, mob_intel, block_intel, lessons, deaths, postmortems, poi,
   wiki_pages, chat_log, code_changes. Seeded from docs/.

2. Captures every death event into the `deaths` table with full context
   (last skill, hostile, recent dispatches, snapshot). A periodic coach
   loop (≤3 Pi calls/hour) extracts generalised lessons from batches of
   unanalysed deaths and writes them to the `lessons` table.

3. Honours learned lessons at dispatch time. curriculumReflex and
   defendReflex consult knowledge.topAdvice before action: if a
   high-confidence lesson says "avoid <skill>" or "prefer <alternative>",
   the dispatcher swaps or backs off. Closes the learning loop.

4. Narrates its actions in Russian MC chat (≤8 lines/hour, ≥75s gap):
   skill starts, threat sightings, dusk/dawn, respawn, milestones.

5. Includes pre-v0.2.0 WIP in the baseline commit: pathfinder watchdog
   refinements that avoid interrupting collectBlock, metric-driven
   reflex recovery, persistent skill metrics, new skills (flee,
   acquire-food, place-chest, sleep), gather.logs hostile-proximity
   bail-out (auto-patch).

See docs/v0.2.0-self-learning.md for the rc.2/final roadmap.
237/237 tests pass. better-sqlite3 dep added; the subsystem
gracefully no-ops when the driver is missing.
2026-05-27 13:23:50 +03:00
mayatnikovandClaude Opus 4.7 e183aaec4e feat(runtime/coach,reflex): retrieval-augmented dispatch via learned lessons
This closes the learning loop. Lessons in knowledge.db now actually
influence behaviour:

- runtime/coach/advice.js: consult({plannedSkillId, snapshot}) reads
  knowledge.topAdvice() and returns 'override' / 'avoid' / 'proceed'.
  When a lesson says "avoid <skill>" with prefer="survive.flee" (etc.),
  the dispatcher swaps in the alternative.

- runtime/reflex.js:
  * curriculumReflex now consults advice before dispatch; on 'avoid'
    backs off the planned skill + sets wander hint; on 'override'
    dispatches the lesson's preferred alternative.
  * defendReflex (dist≤4 melee branch) consults advice too — so a
    creeper at 4m honours the starter rule "attack creeper → flee".
    Failure outcomes feed back via markApplied so confidence stays
    grounded.

SAFE_OVERRIDES whitelist contains only known runSkill targets
(survive.flee, survive.sleep, survive.eat, recovery.tunnel-out,
explore.far/wander, village.build-shelter); unknown prefers fall back
to plain 'avoid'.

7 advice tests; total suite 237 green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 13:23:20 +03:00
mayatnikovandClaude Opus 4.7 5449ad2e0d feat(runtime/coach,persona): death post-mortem coach + Russian chat narration
Coach (runtime/coach/postmortem.js):
- bot.on('death') → captures context (last skill, hostile, recent
  scenarios, journal nearby, snapshot) → inserts row into knowledge.deaths
- Periodic drain (every 5min, ≤3 Pi calls/hour, 12min cooldown):
  batches up to 8 unanalysed deaths, asks Pi to extract 1-3 generalised
  lessons in JSON, persists to knowledge.lessons + knowledge.postmortems
- Pi prompt asks for structured advice (trigger_skill, trigger_hostile,
  avoid_skill, prefer_skill, confidence) so future dispatch can act on it

Persona (runtime/persona/chatter.js):
- Polls snapshot every 5s, narrates Russian lines on transitions:
  skill start, threat spotted, dusk/dawn, respawn, stuck, milestone done
- Rate-limited: min 75s gap, max 8/hour, duplicate suppression
- ~14 template buckets covering gather/craft/build/travel/combat/weather

Wire-up in bot.js (5 lines):
- initKnowledge({stateDir}) fire-and-forget at module load
- attachCoach + attachChatter inside bot.once("spawn")

Tests: 14 new (7 coach, 7 persona). Total suite 230 green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 13:19:56 +03:00
mayatnikovandClaude Opus 4.7 3da10ce58c feat(runtime/knowledge): SQLite knowledge base with starter intel + lessons
Per-server state/<host>/knowledge.db. Schema:
- recipes (seeded from docs/minecraft-recipes.json, 38 rows)
- mob_intel (15 mobs incl. creeper/zombie/skeleton with verdict_no_weapon)
- block_intel (30 blocks with required_tool / drops)
- lessons (12 starter rules — "don't attack creepers with fists", etc.)
- deaths / postmortems / poi / wiki_pages / chat_log / code_changes

Public API in runtime/knowledge/index.js: initKnowledge, recall, record,
markApplied, topAdvice, lookupRecipe/Mob/Block, insertDeath, insertPostmortem,
recordPOI, poiNearby, logChat. All operations gracefully no-op when
better-sqlite3 is unavailable.

10 tests; covers init, seeding, recall filters, lesson lifecycle, death/PM
round-trip, spatial POI queries, chat log.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 13:14:59 +03:00
mayatnikov b5954f193d fix(runtime/skills/chop-logs): skip unsafe log gathering near hostiles 2026-05-27 13:07:40 +03:00
mayatnikovandClaude Opus 4.7 fe6961e5f7 v0.2.0-rc.1: design doc + version bump + better-sqlite3 dep
Major iteration: self-learning agent with knowledge base, post-mortem
coach, and persona narration. See docs/v0.2.0-self-learning.md for the
full design. This commit only adds the scaffolding (design + deps);
subsequent commits add the modules.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 13:06:54 +03:00
mayatnikovandClaude Opus 4.7 86e5294bb8 chore: snapshot pre-v0.2.0 WIP (pathfinder/reflex/metrics/skills improvements)
Baseline for the v0.2.0 self-learning iteration. All 205 tests pass on this
state. Subsequent commits in this branch layer the knowledge base,
post-mortem coach, and persona narration on top.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-27 13:05:12 +03:00
mayatnikov 3ee2ab886e chore(runtime/pathfinder): explicit info log when watchdog arms (visibility) 2026-05-26 20:00:54 +03:00
mayatnikovandClaude Opus 4.7 01531c54c8 feat(runtime/pathfinder): stuck-replan watchdog — react to mid-path obstacles
Problem (live 2026-05-26 screenshot): player drops a block in front of
the bot mid-path. mineflayer-pathfinder computes the path once when
goto() is called and never recomputes for world changes. Bot pushes
forward against the new block until the 30–60 s goto timeout fires,
visible as the bot just standing there pressing W.

Fix — runtime/pathfinder-watchdog.js: per-bot poll loop (2 s tick) that
runs while bot.pathfinder.goal is non-null. Tracks horizontal position.
If movement < 0.5 blocks for > 6 s after an initial 1.5 s grace,
forces a replan: setGoal(null) + setGoal(<same goal>) on a 250 ms
delay. That makes the planner rebuild the path against the current
world, so it routes around the new block — or, with canDig=true in our
profiles, digs through it. Capped at 3 replans per goal so a genuinely
unreachable target still bubbles up to the caller's timeout.

Side benefit: catches mineflayer-pathfinder issue #222 ("path hangs
on unreachable goal") much earlier than our 45 s goto wrappers.

Wired into bot.js on the "spawn" event and stopped on gracefulExit.
7 new unit tests cover the polling math + replan cap. 197/197 green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 20:00:01 +03:00
mayatnikovandClaude Opus 4.7 d9f11b7d99 feat(social): Pi-driven Russian chat replies with persistent per-player history
Replaces the canned-template "yo" path for greetings / status / addressed
banter with a Pi roundtrip that takes a real persona, the bot's current
in-game context, the operator's diary tail, AND the last 8 chat turns
with THIS specific player. Per-player cooldown (8 s) + per-message
chat-rate cooldown survive the existing throttle so a chatty player
can't drain Pi tokens.

runtime/social/chat-history.js — append/recent per player into
state/<host>/chat/<player>.jsonl, 1000-line rolling cap. Each entry
stores { ts, dir, text, snap? } where snap is a compact position +
activeSkill + milestone at the time of the turn, so Pi can later say
"помнишь когда мы тогда у воды лес рубили". Survives restarts and
auto-patch cherry-picks.

runtime/social/reply-pi.js — Russian-first system prompt locking the
bot as "pepa_bot, автономный игрок-фермер" on play.xmatic.team, one-
line answers, no emojis, no AI/bot self-mentions, no sycophancy. Spawns
`pi -p`, sanitises the response (strips pepa: prefix, code fences,
quotes, multi-line), caps at 200 chars before sending into MC chat.
Graceful: timeout/parse-fail → returns null, caller falls through to
the existing template path so the bot never goes mute.

bot.js handleChat:
- Skip messages from our own username (defensive — never reply to self).
- Record every inbound line into chat-history.
- For non-COMMAND_LIKE / non-UNSAFE intents, try piReply first; on
  success, send + record outbound; on null/throw, fall through to the
  templated generateReply.

14 new unit tests (sanitiseReply edge cases, history rotation, snap
compaction, prompt content). 190/190 green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 19:12:58 +03:00
mayatnikov 49ce0f9bb5 fix(runtime/actions): clear stale collectblock queue after chop timeout 2026-05-26 19:00:09 +03:00
mayatnikov f13799de28 fix(runtime/reflex): retreat after repeated melee clears 2026-05-26 16:13:11 +03:00
mayatnikov b23aad2128 fix(runtime/reflex): verify melee clears hostile 2026-05-26 16:09:21 +03:00
mayatnikov d87b83ea35 chore: gitignore .claude/ (subagent worktrees) so auto-patch doesn't refuse on dirty tree 2026-05-26 16:04:13 +03:00
mayatnikovandClaude Opus 4.7 8bc9b41c01 feat(runtime): cmd:screenshot + cmd:force-incident + prismarine-viewer dep
Three live-verification surfaces on top of v0.1.0:

1. cmd:screenshot { reason?, frames? } → runtime/viewer.takeScreenshot
   - Headless POV render via prismarine-viewer.headless to
     state/<host>/screenshots/<ISO>-<reason>.mp4 (1 frame ≈ ~10 KB).
   - Lazy-loads the heavy GL stack on first call so the bot doesn't pay
     the cost on startup or in TUI-only sessions.
   - Returns { ok, path, error } over IPC LOG event.
   - Known limitation: needs node-canvas. node-canvas v3 (current npm
     default) is incompatible with prismarine-viewer's API; v2 doesn't
     build under Node 24 (node-pre-gyp fail). So today the feature is
     wired and the IPC contract is stable, but the underlying render
     fails fast with "createCanvas is not a function". A future cleanup
     can either fork the renderer or pin a Node 20 toolchain.

2. cmd:force-incident { kind?, reason? } → filePostCritique path
   - Operator-triggered demo of the critic → proposal → auto-improve →
     auto-patch chain. Was previously only observable when the bot
     genuinely got stuck. Now a single IPC call exercises the full
     loop on demand.
   - Verified live 2026-05-26: critic call returned a real, useful
     critique ("attack zombie returns done while target is alive →
     blocks gather.logs"), proposal landed with all sections including
     the Critic block, auto-improve picked it up within 10 s.

3. prismarine-viewer + canvas added to dependencies so npm install
   builds the deps once and the IPC surface is always available.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 16:01:43 +03:00
mayatnikovandClaude Opus 4.7 4ae63dabe1 feat(runtime): v0.1.0 — adopt Voyager critic + Mindcraft modes/library/lint
Five concrete patterns from Voyager and Mindcraft, applied in our shape
without abandoning the git-as-evolution-substrate that makes pepa
distinct. Plus a first multi-agent surface so two bots from the same
repo can share intent.

1. runtime/critic.js (Voyager critic.txt)
   - Spawns `pi -p` with a JSON-only critic prompt before a proposal is
     written. {reasoning, success, critique}.
   - success=true short-circuits the proposal (bot recovered between
     detector tripping and now), saving Pi tokens on false positives.
   - critique is spliced into the proposal body via attachCritique() so
     the downstream auto-patcher has a sharp spec.
   - Graceful: pi missing / timeout / unparseable JSON → proposal still
     filed without the critic block.

2. scripts/lint-patch.js (Mindcraft coder._lintCode)
   - Pre-flight gate between Pi commit and npm test: node --check, dynamic
     import (catches missing named exports), regex extraction of
     runSkill("id") calls cross-checked against the live registry.
   - Cheaper than npm test, fails fast with a clear reason.

3. runtime/stuck-incident.renderActionTemplate (Voyager action_template.txt)
   - All proposal bodies now follow the same fixed-section layout: Task /
     Last result / Execution error / State / Metrics / Journal /
     Scenarios / Critique / Fix / Edit scope / Forbidden.

4. runtime/skill-library.js (Mindcraft skill_library.getRelevantSkillDocs)
   - Word-overlap ranking (Mindcraft's offline fallback) — zero deps,
     deterministic. auto-patch.js injects top-3 similar skills into the
     Pi prompt as "look at these patterns".

5. runtime/modes.js (Mindcraft modes.js)
   - Declarative {name, interrupts, on, active, update(ctx)} chain that
     runs BEFORE the curriculum each tick.
   - Ships self_preservation (low HP → eat/flee), hunger (food<14 → eat),
     night_shelter (night + bed in hand → sleep). Cleaner than ad-hoc
     lastFleeAttempt cooldowns in reflex.js.

6. runtime/social/conversation.js + cmd:conv-say/conv-recent/conv-list
   - File-JSONL topic channel so two bots from the same repo (different
     usernames, different host dirs under state/) can append turns and
     read peers. Skeleton — multi-agent collaboration on top later.

Differentiator preserved: every Pi-written skill still lands on main via
auto-patch.js (real git branch + smoke gate + cherry-pick). Voyager
keeps skills in a Chroma JSON, Mindcraft keeps them in RAM — pepa keeps
them as versioned source code reviewable in `git log`.

package.json: 0.0.1 → 0.1.0. 174/174 tests pass. README + AGENTS updated.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 15:16:42 +03:00
mayatnikov 25e39c7244 fix(runtime/actions): place crafting table beside bot 2026-05-26 14:52:44 +03:00
mayatnikov ecbfa28f3c fix(runtime/actions): cancel timed-out log collection 2026-05-26 14:26:43 +03:00
mayatnikov b215d5b2c2 fix(runtime/skills/chop-logs): stop repeated gather log timeouts 2026-05-26 13:54:26 +03:00
mayatnikovandClaude Opus 4.7 28f5d9e483 fix(perception): use numeric block ids — callback matchers silently fail under ViaBackwards
Root cause of "bot just stands still": every gather.* skill was using
bot.findBlock({ matching: (b) => names.includes(b.name) }), and under
mineflayer 1.21.4 + ViaBackwards the Block objects fed into the
callback have a wrong .name field (Block.type / numeric id is still
correct — this is mineflayer issue #2347). Every search returned null,
every skill reported "no_target", reflex looped wander → tunnel-out
forever. The bot's logs said "dispatch ok" while the operator watched
it pace in circles.

Proven live with a new diag.match skill on play.xmatic.team:
  findBlocks({matching: numericIds})              → 50 hits
  findBlocks({matching: (b) => b.name === ...})   →  0 hits  ← the bug
  findBlock({matching: (b) => b.name === ...})    → null     ← the bug
  findBlock({matching: numericIds})               → dark_oak_log @ (606,62,110)

After this fix the same bot from the same spawn dispatches gather.logs
and reaches the chop loop ("chop: dark_oak_log at 606,62,110 (tool=fists)")
instead of returning "no reachable log within 64 blocks".

Changes:
- runtime/perception.js (new): findBlocksByName / findNearestBlockByName
  centralise the numeric-id workaround for any future skill.
- runtime/actions.js: chopNearestTree, sleepInBed, placeCraftingTable now
  use perception. Also load mineflayer-tool plugin alongside collectblock
  (collectblock 1.6 hard-requires bot.tool to dispatch a dig).
- gather-stone, gather-wool, deposit-surplus rewritten to numeric-id
  search. gather-wool also loads mineflayer-tool.
- diagnose-scan.js (new): two diagnostic skills — diag.scan reports
  findBlocks counts per radius for common blocks; diag.match cross-tests
  the four matcher styles so this regression can be re-proven on demand.
- runtime/skills/index.js: registers diag.scan + diag.match.

Memory: project_findblock_callback_broken_under_viabackwards.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 13:45:29 +03:00
mayatnikovandClaude Opus 4.7 69d1298fbd fix(auto-patch): lockfile-coordinated supervisor restarts
When Pi writes a multi-file runtime patch, the supervisor's file watcher
can fire between two consecutive writes, kill the bot mid-edit, and load
a half-saved file with a SyntaxError. Loop until the operator stops it.

scripts/auto-patch.js now creates state/auto-patch.lock with its PID
right after the branch checkout (before spawning pi -p), and removes
it on every exit path. runtime/supervisor.js defers any watch-triggered
restart while the lock holder is alive, polling every 2 s; once the
lock drops it waits 1.5 s for the final write to settle, then runs
\`node --check\` on the changed file and only restarts if it parses.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 13:24:54 +03:00
mayatnikov 6909715f75 fix(runtime/skills): guard stationary blind fallback 2026-05-26 13:22:30 +03:00
mayatnikov 36e547a896 fix(runtime/actions): require horizontal escape movement 2026-05-26 12:45:13 +03:00
mayatnikovandpepa_bot self-improvement loop 6ba6bcdbb7 test(recovery-tunnel-out): mock bot.dig in jump-in-place test
Pi-authored follow-up from a second auto-patch cycle — the original
"does not count jumping in place as escape" test forgot to provide a
dig() mock, so the in-place jump path crashed when the test exercised
escape-pit's "dig above" fallback. Adding the mock makes the test
actually verify the assertion it claims.

138/138 tests now.

Co-Authored-By: pepa_bot self-improvement loop <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:44:12 +03:00
mayatnikovandClaude Opus 4.7 6560c0765c fix(auto-improve): detach auto-patch + recovery-tunnel-out test in suite
Two bugs the live self-improvement run exposed:

1) Auto-patch was spawned with detached:false, so when supervisor
   restarted bot.js (file change after Pi's commit landed on the
   auto branch), the auto-patch child was killed mid-way — Pi's
   commit lived in the auto branch but never got cherry-picked.
   Recovered manually this round via reflog + cherry-pick. Now
   detached:true + child.unref() + a per-run log at
   state/_auto-patch-last.log so the operator can read Pi's full
   output later.

2) Pi's recovery-tunnel-out.test.js was created but not in npm test
   script; tests would have stayed unrun forever. Added.

Also commits the Pi-authored skill (eb29591 cherry-picked):
- runtime/skills/recovery-tunnel-out.js (+ test)
- improvements to runtime/actions.js + runtime/skills/explore-far.js
- wired into runtime/skills/index.js

npm test 137/137.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:43:30 +03:00
mayatnikov 564557450d fix(runtime/skills): recover from wedged pits with tunnel-out 2026-05-26 12:39:23 +03:00
mayatnikovandClaude Opus 4.7 a57625b541 feat(runtime): wedged-cant-escape proposal trigger (self-improvement v2)
Closes the self-improvement loop: when escape-pit + wedged-jump +
blind fallback all run 3× in a row without freeing the bot, fire a
dedicated proposal at the auto-improver. Pi gets the full context
(journal byKind, last 12 scenario-memory entries, current slim
snapshot) and is asked to either improve escapePit() (dig forward +
down + side, not only up) OR add a brand-new recovery.tunnel-out skill.

- runtime/stuck-incident.js: new checkWedged() path with separate
  cooldown (10 min) from the no-progress path. noteResult() ingests
  every dispatched action's detail.mode to count wedged completions.
- runtime/bot.js: dispatchAction calls stuckIncident.noteResult(res)
  after each result; tick() calls checkWedged() and files the proposal
  via writeProposal({editScope:[runtime/actions.js, runtime/skills/,
  runtime/reflex.js]}).

This is the architectural piece: a bot wedged in a 1×1 hole now
generates a proposal that Pi can act on (with edit-scope guard rails
+ npm test smoke gate from PR #19), instead of looping wedged-jump
forever.

Verified live: bot now also picks direction from journal —
"explore.far: journal says leanest quadrant=NE → prefer N" — first
time the bot uses persistent memory to choose where to go next.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:24:22 +03:00
mayatnikovandClaude Opus 4.7 d960db4819 feat(runtime): persistent memory — world-journal + scenario-memory
Closes a structural gap: the bot now actually REMEMBERS what it
discovered and what it tried. Two stores live under state/<host>/ and
are wired in automatically.

runtime/world-journal.js
- Append-only JSONL of discovered points (chopped, placed, base,
  shelter, farm, dead_end). Indexed by 16-block spatial grid; O(neighbors)
  nearest() lookups; 6 h age prune; 10k line ceiling with trim.
- leanestQuadrant({x,z}) reports the quadrant the bot has the FEWEST
  markers in — used by explore.far to circle rather than retread.
- summary() exposed for the stuck-incident proposal body.

runtime/scenario-memory.js
- Sliding window of (skillId, situationHash, code, ok, detail) tuples.
- situationHash() is a coarse fingerprint (16x8x16 cell + day/night +
  food/hp bucket + inv key set + closest hostile). So "same kind of
  place + same kind of state" matches.
- shouldSkip({skillId, situation}) → true after ≥3 failures within 30
  min UNLESS a more-recent success in the same situation un-locks it.
- recentTailFor() exposed for the stuck-incident body.

Wiring (runtime/bot.js):
- dispatchAction captures situationHash BEFORE the action runs and
  records (skillId, situation, code, ok) after — failures are attributed
  to the dispatch-time state, not the partial-effect state.
- worldDelta fields (choppedAt, minedAt, placedAt, baseAt, shelterAt,
  plantedAt, harvestedAt, tilledAt) auto-flow into the journal.
- no_target + silent_dig_failure also write dead_end markers.

Scheduler / skills now consume memory:
- reflex.js curriculum reflex calls memory.shouldSkip — if the same
  (skill, situation) failed 3+ times recently, auto-converts to a
  wander hint so the bot leaves and tries elsewhere.
- explore.far calls journal.leanestQuadrant when multiple cardinal
  directions are walkable and prefers the less-explored one.
- gather.logs walks to the nearest known "chopped" bucket within 96
  blocks before falling through to findBlock — chunks with confirmed
  trees are more likely to yield another.

stuck-incident body now includes journal byKind + last 12 scenario
entries so Pi can write a structural fix, not just a guard clause.

Architecturally: this is the foundation for "bot rewrites itself".
The proposals Pi now receives carry real signal about what was tried
and what's around, instead of a single snapshot in isolation.

10 new tests (world-journal × 5, scenario-memory × 5). npm test 134/134.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:20:03 +03:00
mayatnikovandClaude Opus 4.7 3e3ea3e597 fix(runtime): escape-pit fallback for wedged bot
When probe-cardinal shows all 4 directions blocked (the bot is in a
1×1 pit, surrounded by leaves, or in a corridor corner), don't just
hold forward+jump — actually dig the block above the bot's head,
jump into the new gap, repeat up to 3 times. Both wander and
explore.far now call escapePit() in this branch.

Observed live: bot fell into a pit at (623,71,106) after first
explore.far and looped wedged-jump→still-wedged→wedged-jump for 60s
before this fix. With escape-pit, the bot now actually breaks out.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:07:59 +03:00
mayatnikovandClaude Opus 4.7 0aae5e2e68 fix(runtime): wander/explore.far probe-then-go (bot actually moves)
Ground-truth finding (diag.physics):
  forward N:0.03 E:3.38 S:0 W:3.26  → forward WORKS in unobstructed dirs
  jump ΔY=1.25                       → jump WORKS (vanilla height)
  dig untested (no soft block within 6 of spawn)

So the bot CAN move and jump — the previous "stands still" symptom was
our wander/explore code picking blocked random angles and trusting a
pathfinder that times out on this server's terrain. Each retry just
picked another random direction, often the same blocked one.

- runtime/actions.js wander: probe 4 cardinal yaws for 800ms each,
  measure actual Δ, commit to the best one for the remaining budget.
  Falls back to "wedged-jump" (forward+jump 2.5s) only when ALL four
  cardinals are <0.5 blocks.
- runtime/skills/explore-far.js: same probe-then-go shape, scaled to a
  ~48-block long walk in the best direction. Replaces the static
  NE/SE/SW/NW quadrant rotation that ignored what was actually
  walkable.
- runtime/movement-profiles.js: canDig back to true on gather/travel/
  flee. The earlier "everything false" defensive default was based on
  a wrong hypothesis (silent dig failure) — diag.physics + server-side
  inspection (no anti-cheat plugin, spawn-protection=0) showed dig is
  fine.
- runtime/compat.test.js: assertions follow profile defaults.
- runtime/skills/diagnose-physics.js: forward probe now tries 4
  cardinals and returns trials + bestDir + bestDist so it can be used
  to debug "wedged" reports later.

Verified live: bot now actually walks 47 blocks north after
probe.cardinal showed N:2.4 free. First end-to-end real movement on
play.xmatic.team since this session started.

npm test 124/124.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:05:14 +03:00
mayatnikovandClaude Opus 4.7 86f0c1799e fix(runtime): pin MC_VERSION=1.21.4 + ground-truth probe + close-loop dig
Two-pronged response to user-confirmed "bot stands still, doesn't actually
chop" on play.xmatic.team:

1. Pin protocol — .env now sets MC_VERSION=1.21.4. minecraft-data has
   wrong packet ID mappings for protocol 775 (server 26.1.2 via
   ViaBackwards 5.9.1) — see mineflayer#3888 and #3717. 1.21.5 also has
   an enchants decoder bug that breaks bot.dig. 1.21.4 is the last
   protocol mineflayer 4.37.1 can speak cleanly through VIA.

2. Don't trust dig success — runtime/actions.js chopNearestTree and
   runtime/skills/gather-stone.js now lookAt(face center)+forceLook,
   await collectBlock, then re-read the target block. If the log/stone
   is STILL there, return ok:false code:"silent_dig_failure" and
   blacklist the position. Prevents the curriculum from reporting
   "wood.16 in progress" while the world hasn't actually changed.

3. Defensive default — runtime/movement-profiles.js: canDig=false on
   every profile until dig is confirmed working live. Otherwise
   pathfinder schedules paths through must-dig blocks and the bot loops.

4. Ground-truth probe — runtime/skills/diagnose-physics.js dispatches
   forward/jump/dig probes and writes the result to the diary. New
   IPC command cmd:run-skill lets the operator (or a future curriculum
   trigger) fire any skill on demand; it waits for the current action
   to finish before dispatching. /tmp/pepa-runskill.mjs is a one-shot
   client.

Live probe on play.xmatic.team confirmed: forward Δ=0.003 over 2s
(BROKEN — server rejects movement packets), jump ΔY=0.42 (likely
physics jitter, not a real jump). Strongly suggests an anti-cheat
plugin gating bot-style movements server-side — beyond protocol pin.

npm test 124/124.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 11:51:43 +03:00
mayatnikovandClaude Opus 4.7 19dc8e12c6 fix(runtime): unstick wander loop + chop radius + explore.far skill
Follow-up to the iteration-1 fixes. Live smoke on play.xmatic.team
revealed the bot was spawning into a tree-less plain (no log within
32 blocks of spawn), looping wander→gather→no_target→wander
forever inside a 16-block box.

- runtime/actions.js: chopNearestTree search radius 32 → 64 (still no
  trees on this spawn, but a normal biome will be served well by it).
  wander now has a blind-walk fallback when pathfinder times out
  (look+forward+jump for 3 s) so the bot at least unsticks from leaves
  or pillars. Pathfinder timeout reduced 30 s → 15 s.
- runtime/skills/explore-far.js: new explore.far skill — walks ~48
  blocks in a quadrant (NE/SE/SW/NW, rotating per call) so successive
  hints actually circle the spawn instead of bouncing in place. Blind
  walk fallback included.
- runtime/reflex.js: when the scheduler is told to wander twice in a
  row by gather.* recover hints, it now dispatches explore.far instead
  so the bot actually leaves the patch it's stuck in. Resets the
  consecutiveWanderHints counter on any success.
- runtime/reflex.js (sleep): no longer dispatches when the bot has
  neither a bed in inventory NOR a known shelter/base location —
  saved one dispatch + 5-min cooldown per restart at night.
- runtime/reflex.js (eat): inventory check + lastEatAt always updated
  fix the eat-spam loop observed live (every tick fired "eat" → "no
  food in inventory" → again).
- runtime/skills/chop-logs.js: recognise "no log within ..." as
  no_target so the recover hint switches the bot to wander/explore.

npm test 124/124.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 11:18:36 +03:00
mayatnikovandClaude Opus 4.7 29542f0559 fix(runtime): unstick scheduler + chop + sleep + bed/shelter/farm skills
Recovers the bot from the live-server symptoms reported 2026-05-26:
1) constant supervisor reconnects, 2) chop "clicks once and stops",
3) sleep does nothing without a bed and so blocks night-skipping for
other players, 4) curriculum reflex always fell through to wander.

Supervisor (#38):
- runtime/watch-filter.js: pure predicate excluding *.test.js + the
  supervisor itself; recursive:true so skills/ + social/ edits also
  restart. Burned a working main once when test files counted toward
  the rollback threshold.
- runtime/supervisor.js: watch-triggered restarts no longer count
  toward the crash-loop rollback path. Watcher is now recursive.

Chop / mine (#39):
- runtime/actions.js + runtime/skills/gather-stone.js: replaced raw
  pathfinder.goto + bot.dig with mineflayer-collectblock's
  bot.collectBlock.collect — handles approach, repositioning, LoS,
  dig and pickup as one primitive. Old version "swung once" because
  GoalGetToBlock often parked the bot in leaves above the log.

Sleep + bed (#40):
- runtime/actions.js: sleepInBed now ALSO places a carried bed on
  solid ground next to the bot and sleeps on it. Critical so the bot
  stops blocking player night-skipping the moment it owns a bed.

Bed pipeline (#41):
- runtime/skills/gather-wool.js: gather.wool skill — mines wool block
  if any nearby, otherwise shears or attacks the nearest sheep.
- runtime/skills/craft.js: craftBedSkill (any colour the bot has ≥3
  wool of, plus 3 planks, plus a table).
- runtime/curriculum.js: new milestone survive.bed sits between
  wood.tools and stone.32 so the bot gets a bed BEFORE everything else.
  Test fixture updated to include a red_bed in post-survive.bed stages.

Village / shelter / wheat (#42, #43):
- runtime/skills/build-shelter.js: village.build-shelter — real 3×3×3
  resumable hut blueprint around the recorded base, places one block
  per loop, idempotent so an interrupted build resumes correctly,
  marks each placed block in the owned-blocks ledger.
- runtime/skills/deposit-surplus.js: village.deposit-surplus opens
  the nearest chest and transfers surplus stacks while keeping a
  reserve of tools/food/bed.
- runtime/skills/farm-wheat.js: farm.wheat does one step per call
  (till adjacent-to-water grass, plant seeds, or harvest ripe wheat).
- runtime/curriculum.js: village.shelter milestone after base-site.

Scheduler glitch (root of "always wander"):
- runtime/bot.js: curriculum + locations are now computed BEFORE
  runTick. Previously they were stamped AFTER, so reflex.js saw
  snapshot.curriculum=undefined every tick and fell through to the
  wander fallback. Verified live: scheduler now dispatches
  gather.logs/gather.stone/craft.* by id via runSkill.

Eat-spam:
- runtime/reflex.js: eatReflex now checks inventory for actual food
  and updates lastEatAt on EVERY dispatch (not only successes), so a
  failed eat respects the 5 s cooldown instead of firing every tick.

npm test 123/123. Validated live on play.xmatic.team (curriculum
dispatched gather.logs via runSkill, recover hint switched to wander
when no log in range).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 11:12:21 +03:00
ea4f16a0da feat(runtime): scheduler-via-runSkill + Pi banter escalation + base-site (follow-ups) (#20)
Three closures of remaining PRD follow-ups, one merge:

1. Reflex scheduler now drives behaviour from the curriculum.
   - reflex.js: replaced ad-hoc techTreeReflex + autonomousReflex with
     curriculumReflex that dispatches the skill suggested by
     snapshot.curriculum.plan via runSkill. Per-skill backoff for
     missing_tool / missing_material / no_target / no_food_source /
     unsupported_version. recover() hint with `{hint:"wander"}` swaps
     the next tick to wander for 60 s.
   - Chain is now: defend > eat > sleep > curriculum > idle.
   - reflex.test.js: 11 new tests covering busy/disconnected,
     defend/eat preemption, dispatch by id, unknown-skill fallback,
     per-skill + wander-hint backoffs, onComplete updating backoff.

2. Pi escalation for ADDRESSED_BANTER with hard rate limit.
   - bot.js: when generateReply returns {escalate:true}, spawn askPi
     with bot state + last 5 lines from that speaker (redacted via
     chatMemory). Reply capped at 200 chars, sent as one chat line.
   - Rate cap: 6 calls/hour, 90 s min gap. Suppressed escalations
     log once and silently drop.

3. Phase 4 substrate.
   - runtime/locations.js: atomic JSON store
     (state/<host>/locations.json) with setLocation / getLocation /
     nearestLocation / removeLocation; 6 tests.
   - runtime/base-site.js: scoreCurrentPosition(bot) + pure scoreSite
     bundle (wood / stone / water / flatness / no-players /
     no-foreign-builds, owned-blocks excluded from claim penalty);
     6 tests.
   - runtime/skills/choose-base.js: village.choose-base skill — scores
     the current spot, writes locations.base if score ≥ 8, otherwise
     returns code:"too_weak" with a wander recover hint.
   - curriculum.js: new final milestone village.base-site fires
     village.choose-base until a base location exists.
   - bot.js: stamps snapshot.locations from listLocations() each tick
     so the curriculum can read it without coupling to disk.

docs/runtime.md updated with three new sections.
npm test now 116/116.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 10:46:14 +03:00
82a8250a12 feat(auto-patch): enforce proposal editScope + run npm test before cherry-pick (#19)
Two safety rails on the unattended self-improvement loop:

1. scripts/edit-scope.js + .test.js: pure helpers that parse
   `editScope: [...]` out of a proposal frontmatter (the field Phase 6
   started writing) and validate a list of changed files against it.
   13 tests covering null/missing/malformed frontmatter, directory
   prefix matching, exact-file matching, default-scope fallback.

2. scripts/auto-patch.js:
   - reads editScope from the proposal (falls back to ["runtime/"])
   - injects the allowed paths into the Pi prompt so Pi knows the
     boundaries up front
   - validates the diff against scope + auto-allows any
     runtime/**/*.test.js files Pi added
   - runs `npm test` on the patched branch BEFORE cherry-picking;
     refuses to land a patch that breaks the suite

Closes the "Smoke checks run before applying patch" item from
plans/autonomous-survival-bot-prd.md §7 Phase 6. npm test now 92/92.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 10:39:18 +03:00
d2e52a1b79 feat(runtime): compatibility hardening (Phase 7) (#18)
Phase 7 of plans/autonomous-survival-bot-prd.md. Five small modules
that close the recurring "shared state" and "version-pinned list"
failure modes the PRD flags in §7 and §5.4.

New:
- runtime/movement-profiles.js: named profiles (GATHER, TRAVEL, FLEE,
  BUILD, RETURN_TO_BASE) as pure descriptors via PROFILE_DEFAULTS,
  plus applyProfile(profile, bot) that hands a fresh Movements to
  pathfinder. Avoids the "flee left canDig=false on the shared
  Movements, next chop got stuck in canopy" regression.
- runtime/owned-blocks.js: JSONL ledger of blocks this bot placed/
  removed (state/<host>/owned-blocks.jsonl); isOwned({x,y,z}) for
  O(1) lookups; ensureDir() makes the parent dir lazily.
- runtime/claim-avoidance.js: classifyArea({blocks, isOwned}) returns
  player_build / natural_or_owned / insufficient_data based on
  man-made block density vs ownership ratio; shouldAvoid(area) helper.
  Designed for gather/place skills to call before touching contested
  area.
- runtime/skills/compat.test.js: runs runtime/skills/groups.js against
  real minecraft-data registries for 1.18.2, 1.20.4, 1.21.5; spot-
  checks that pale_oak_log only appears on 1.21+ etc.
- runtime/compat.test.js: 10 tests covering movement descriptors,
  isManMadeBlockName, classifyArea, owned-blocks markPlaced/dedup/
  isOwned/markRemoved.

npm test now 79/79.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:39:38 +03:00
c7eab06f22 feat(runtime): stuck-incident detector + skill metrics + edit scope (Phase 6) (#17)
Phase 6 of plans/autonomous-survival-bot-prd.md. Expand the
self-improvement loop so the bot can spot and report no-progress
stagnation, not just exception-class failures.

New:
- runtime/stuck-incident.js: detector fires a structured proposal when
  the same noProgressReason persists past 5 min (cooldown 30 min).
  Body includes runtimeState, milestone, suggested skill, slim
  snapshot, last action result, per-skill success/failure metrics
  and a forbidden-paths list. Pure module — caller (bot.js) writes
  the proposal.
- runtime/skill-metrics.js: in-memory per-skill ok/fail counters
  surfaced on snapshot.skillMetrics for the TUI and the incident
  body.
- runtime/stuck-incident.test.js: 6 tests covering null reason,
  threshold gating, cooldown, reason change resetting the timer,
  body composition and metrics snapshot.

Wiring:
- runtime/state-store.js: writeProposal accepts {editScope: string[]}
  and persists it in the frontmatter; readProposalEditScope() reads
  it back so future auto-patch.js can refuse cherry-picks that touch
  other areas.
- runtime/bot.js: tick() invokes the stuck detector each tick,
  records skill ok/fail via skillMetrics, stamps snapshot.skillMetrics
  and writes the stuck proposal via writeProposal({editScope}).
  dispatchAction now records into skillMetrics for both the
  resolved-result and the exception path.

npm test now 46/46.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:35:01 +03:00
fc62160524 feat(runtime): social layer — intent / templates / chat memory (Phase 5) (#16)
Phase 5 of plans/autonomous-survival-bot-prd.md. Make the bot feel
present in chat without ever becoming a command executor.

New: runtime/social/
- intent.js: classifyIntent({text, botName}) returns one of GREETING /
  STATUS_QUESTION / ADDRESSED_BANTER / COMMAND_LIKE / UNSAFE_REQUEST /
  AMBIENT. Unicode-aware word boundaries so cyrillic + latin both work
  ("Привет всем" → GREETING, "build me a tower" → AMBIENT unless
  addressed).
- reply.js: generateReply({intent, speaker, snapshot, diaryTail}) →
  short templated response, or {send: null, escalate: true} for the
  caller to decide whether to spend Pi tokens.
- memory.js: createChatMemory() — per-speaker LRU buffer of recent
  lines; redacts password / api_key / JWT-shaped tokens at append
  time, so the buffer can be safely fed back into any future prompt.
- social.test.js: 12 tests (intent edges, memory eviction, redaction,
  reply routing). npm test now 40/40.

state-store.js additions:
- readDiaryTail(n) — reads the last N lines of today's diary; used by
  status replies.
- writeEscalation({from, request, whyUnsure, wouldHave}) /
  listEscalations() — JSONL log under state/<host>/escalations.jsonl
  for UNSAFE_REQUEST classifications and future operator review.

bot.js: handleChat() now routes through social/intent + social/reply
(replacing the Phase-0 inline regexes), records every line into
chatMemory, and writes an escalation when classifyIntent returns
UNSAFE_REQUEST. Command-like notice + dialog-only behaviour from
Phase 0 are preserved.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:30:48 +03:00
ae7b4d89cb feat(runtime): early-game survival curriculum + stone/craft skills (Phase 3) (#15)
Phase 3 of plans/autonomous-survival-bot-prd.md. Gives the bot a
deterministic path from empty inventory through stone-tier tools and
basic storage, without an LLM call per tick.

New:
- runtime/curriculum.js: ordered milestone chooser
  (wood.16 → wood.planks-and-sticks → wood.tools → stone.32 →
   stone.tools → food.basic → storage.chest → shelter.torch). Each
  milestone exposes isDone(inventory, snapshot) and suggest() returning
  a { skillId } plan the scheduler can dispatch via runSkill. isDone
  uses "stage reached" escapes so progress is monotonic — crafting
  planks doesn't bounce the chooser back to "gather 16 logs".
- runtime/skills/gather-stone.js: gather.stone with pickaxe-required
  precondition, blacklist on failed paths, registry-aware matching
  (stone / cobblestone / deepslate / cobbled_deepslate / andesite /
  diorite / granite).
- runtime/skills/craft.js: factory + concrete skills for craft.planks,
  craft.sticks, craft.wooden-axe/-pickaxe/-sword, craft.stone-axe/
  -pickaxe/-sword, craft.furnace, craft.chest, craft.torch (torch
  requires coal or charcoal preflight).

Tests:
- runtime/curriculum.test.js: 14 tests covering chooser ordering,
  per-milestone skill suggestion, inventoryFull threshold, monotonic
  advancement across stage transitions.
- npm test now runs the full suite: 28/28 passing.

Wiring:
- runtime/bot.js: lastSnapshot.curriculum carries the next milestone
  + suggested skill on every tick; lastSnapshot.currentMilestone
  prefers the curriculum title over the planner.md line.
- tui/tui.tsx: milestone line shows the curriculum's suggested skill
  and an [inventory full] flag when isInventoryFull fires.

Reflex.js still calls actions.js directly; wiring the scheduler to
runSkill(plan.skillId, …) lands in Phase 4.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:23:33 +03:00
4b7541435d feat(runtime): skill substrate + dynamic groups + reference skills (Phase 2) (#14)
Phase 2 of plans/autonomous-survival-bot-prd.md. Establishes the
composable skill contract from PRD §5.2 and ports three reference
skills so future phases can layer survival behaviour on top instead of
adding more ad-hoc branches to reflex.js.

New: runtime/skills/
- index.js: skill registry + runSkill(id, ctx, args) wrapper. Enforces
  preconditions, hard timeout, normalises {ok, code, detail, worldDelta}
  on every result, runs validate() and calls recover() on failure.
  Stable failure codes live in RUNNER_CODES (unknown_skill,
  precondition_failed, timeout, threw, validation_failed, done).
- groups.js: registry-derived item/block sets — logs/planks/sticks/beds
  derived by suffix; foods intersects a curated allowlist with the live
  bot.registry; axes/pickaxes/swords scoped to whatever the connected
  server's item table actually ships. Empty set instead of throwing on
  missing registry, so skills can emit code:"unsupported_version".
- chop-logs.js: gather.logs reference skill (wraps chopNearestTree).
- eat.js: survive.eat (wraps eatBestFood, preconditions check carrying
  edible food from the registry-derived set).
- wander.js: explore.wander (wraps wander).
- contract.test.js + groups.test.js: 14 tests covering precondition
  gating, timeout firing recover(), execute exceptions, validate
  flipping ok→false, dynamic group filtering across mock registries.

package.json: `npm test` runs the new contract + groups suites.
docs/runtime.md: documents the skill contract, runner, dynamic groups
and the reference skills.

Reflex.js still calls actions.js directly — wiring the scheduler to
runSkill() lands in later phases when the survival curriculum kicks in.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:15:57 +03:00
f301529f42 feat(runtime): observability + no-progress detector (Phase 1) (#13)
Phase 1 of plans/autonomous-survival-bot-prd.md. The bot must always be
able to answer "what am I doing and why am I not doing more?" without
parsing the log stream.

New modules:
- runtime/state.js: pure FSM classifier emitting emergency / working /
  recovering / planning / social / idle from snapshot + reflex context.
- runtime/no-progress.js: sliding-window detector that watches position
  and inventory; when both are unchanged for 60 s+, emits one stable
  reason code from REASONS (waiting_for_day, night_hostile_nearby,
  no_food_source, inventory_full, no_reachable_target, planner_empty,
  awaiting_action_cooldown).
- runtime/viewer.js: optional prismarine-viewer launcher behind
  VIEWER_PORT. Lazy import so the dep is not required by default.

Wiring:
- runtime/bot.js: tick() now computes runtimeState + noProgressReason
  every tick and stamps them on the snapshot along with activeSkill,
  currentMilestone (read from plan.md, cached 30 s), lastResult,
  failuresByCode and lastEscalation.
- runtime/bot.js: dispatchAction records lastResult and lastFailureAt
  for the recovering-state classifier.
- runtime/planner.js: exports isPlannerBusy(), readNextMilestone()
  and planExists() so the runtime can show planning state + current
  milestone without spawning extra Pi calls.
- runtime/config.js: adds VIEWER_PORT support.

TUI:
- tui/tui.tsx: StatusBar gains a state badge, current-skill row,
  milestone row, no-progress reason warning, last-result line with
  ok/fail color, failures-by-class summary and last-escalation age.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:11:13 +03:00
3310cb320f feat(runtime): survival-bot pivot — MC chat is dialog-only (Phase 0) (#12)
Phase 0 of plans/autonomous-survival-bot-prd.md: change the product
direction from operator-driven remote control to autonomous survival
resident. MC chat is dialog-only for everyone, including
OPERATOR_USERNAMES — commands like come/follow/build/pause/stop are
recorded in the diary but not dispatched. TUI remains the only local
control plane.

Runtime changes:
- Remove operatorGoalReflex from reflex.js (the come-here chat command).
- Replace handleOperatorChat in bot.js with a dialog-only handleChat
  that answers greetings/status questions and records command-like
  verbs (en+ru) without dispatching them.
- Default MC_VERSION to "auto" in runtime/config.js; mineflayer
  receives `false` to trigger version auto-detection.
- Update auto-escalation prompt's reflex chain summary.

Docs:
- AGENTS.md: product pivot notice up top; chat-driven scope-trust is
  flagged as legacy/Pi-only.
- README.md / docs/runtime.md: replace operator-chat command list with
  dialog-only description; update reflex chain summary.
- docs/roadmap.md: Phase 2/3 marked superseded by the PRD where they
  assumed chat-driven control.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:01:26 +03:00
mayatnikovandClaude Opus 4.7 22de63d37d chore: ignore local plans/ scratch directory
Operator PRDs and scratch implementation notes live under plans/ and
should not be committed to the repo.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 21:47:50 +03:00