Closes a structural gap: the bot now actually REMEMBERS what it
discovered and what it tried. Two stores live under state/<host>/ and
are wired in automatically.
runtime/world-journal.js
- Append-only JSONL of discovered points (chopped, placed, base,
shelter, farm, dead_end). Indexed by 16-block spatial grid; O(neighbors)
nearest() lookups; 6 h age prune; 10k line ceiling with trim.
- leanestQuadrant({x,z}) reports the quadrant the bot has the FEWEST
markers in — used by explore.far to circle rather than retread.
- summary() exposed for the stuck-incident proposal body.
runtime/scenario-memory.js
- Sliding window of (skillId, situationHash, code, ok, detail) tuples.
- situationHash() is a coarse fingerprint (16x8x16 cell + day/night +
food/hp bucket + inv key set + closest hostile). So "same kind of
place + same kind of state" matches.
- shouldSkip({skillId, situation}) → true after ≥3 failures within 30
min UNLESS a more-recent success in the same situation un-locks it.
- recentTailFor() exposed for the stuck-incident body.
Wiring (runtime/bot.js):
- dispatchAction captures situationHash BEFORE the action runs and
records (skillId, situation, code, ok) after — failures are attributed
to the dispatch-time state, not the partial-effect state.
- worldDelta fields (choppedAt, minedAt, placedAt, baseAt, shelterAt,
plantedAt, harvestedAt, tilledAt) auto-flow into the journal.
- no_target + silent_dig_failure also write dead_end markers.
Scheduler / skills now consume memory:
- reflex.js curriculum reflex calls memory.shouldSkip — if the same
(skill, situation) failed 3+ times recently, auto-converts to a
wander hint so the bot leaves and tries elsewhere.
- explore.far calls journal.leanestQuadrant when multiple cardinal
directions are walkable and prefers the less-explored one.
- gather.logs walks to the nearest known "chopped" bucket within 96
blocks before falling through to findBlock — chunks with confirmed
trees are more likely to yield another.
stuck-incident body now includes journal byKind + last 12 scenario
entries so Pi can write a structural fix, not just a guard clause.
Architecturally: this is the foundation for "bot rewrites itself".
The proposals Pi now receives carry real signal about what was tried
and what's around, instead of a single snapshot in isolation.
10 new tests (world-journal × 5, scenario-memory × 5). npm test 134/134.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
When probe-cardinal shows all 4 directions blocked (the bot is in a
1×1 pit, surrounded by leaves, or in a corridor corner), don't just
hold forward+jump — actually dig the block above the bot's head,
jump into the new gap, repeat up to 3 times. Both wander and
explore.far now call escapePit() in this branch.
Observed live: bot fell into a pit at (623,71,106) after first
explore.far and looped wedged-jump→still-wedged→wedged-jump for 60s
before this fix. With escape-pit, the bot now actually breaks out.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Ground-truth finding (diag.physics):
forward N:0.03 E:3.38 S:0 W:3.26 → forward WORKS in unobstructed dirs
jump ΔY=1.25 → jump WORKS (vanilla height)
dig untested (no soft block within 6 of spawn)
So the bot CAN move and jump — the previous "stands still" symptom was
our wander/explore code picking blocked random angles and trusting a
pathfinder that times out on this server's terrain. Each retry just
picked another random direction, often the same blocked one.
- runtime/actions.js wander: probe 4 cardinal yaws for 800ms each,
measure actual Δ, commit to the best one for the remaining budget.
Falls back to "wedged-jump" (forward+jump 2.5s) only when ALL four
cardinals are <0.5 blocks.
- runtime/skills/explore-far.js: same probe-then-go shape, scaled to a
~48-block long walk in the best direction. Replaces the static
NE/SE/SW/NW quadrant rotation that ignored what was actually
walkable.
- runtime/movement-profiles.js: canDig back to true on gather/travel/
flee. The earlier "everything false" defensive default was based on
a wrong hypothesis (silent dig failure) — diag.physics + server-side
inspection (no anti-cheat plugin, spawn-protection=0) showed dig is
fine.
- runtime/compat.test.js: assertions follow profile defaults.
- runtime/skills/diagnose-physics.js: forward probe now tries 4
cardinals and returns trials + bestDir + bestDist so it can be used
to debug "wedged" reports later.
Verified live: bot now actually walks 47 blocks north after
probe.cardinal showed N:2.4 free. First end-to-end real movement on
play.xmatic.team since this session started.
npm test 124/124.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-pronged response to user-confirmed "bot stands still, doesn't actually
chop" on play.xmatic.team:
1. Pin protocol — .env now sets MC_VERSION=1.21.4. minecraft-data has
wrong packet ID mappings for protocol 775 (server 26.1.2 via
ViaBackwards 5.9.1) — see mineflayer#3888 and #3717. 1.21.5 also has
an enchants decoder bug that breaks bot.dig. 1.21.4 is the last
protocol mineflayer 4.37.1 can speak cleanly through VIA.
2. Don't trust dig success — runtime/actions.js chopNearestTree and
runtime/skills/gather-stone.js now lookAt(face center)+forceLook,
await collectBlock, then re-read the target block. If the log/stone
is STILL there, return ok:false code:"silent_dig_failure" and
blacklist the position. Prevents the curriculum from reporting
"wood.16 in progress" while the world hasn't actually changed.
3. Defensive default — runtime/movement-profiles.js: canDig=false on
every profile until dig is confirmed working live. Otherwise
pathfinder schedules paths through must-dig blocks and the bot loops.
4. Ground-truth probe — runtime/skills/diagnose-physics.js dispatches
forward/jump/dig probes and writes the result to the diary. New
IPC command cmd:run-skill lets the operator (or a future curriculum
trigger) fire any skill on demand; it waits for the current action
to finish before dispatching. /tmp/pepa-runskill.mjs is a one-shot
client.
Live probe on play.xmatic.team confirmed: forward Δ=0.003 over 2s
(BROKEN — server rejects movement packets), jump ΔY=0.42 (likely
physics jitter, not a real jump). Strongly suggests an anti-cheat
plugin gating bot-style movements server-side — beyond protocol pin.
npm test 124/124.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Follow-up to the iteration-1 fixes. Live smoke on play.xmatic.team
revealed the bot was spawning into a tree-less plain (no log within
32 blocks of spawn), looping wander→gather→no_target→wander
forever inside a 16-block box.
- runtime/actions.js: chopNearestTree search radius 32 → 64 (still no
trees on this spawn, but a normal biome will be served well by it).
wander now has a blind-walk fallback when pathfinder times out
(look+forward+jump for 3 s) so the bot at least unsticks from leaves
or pillars. Pathfinder timeout reduced 30 s → 15 s.
- runtime/skills/explore-far.js: new explore.far skill — walks ~48
blocks in a quadrant (NE/SE/SW/NW, rotating per call) so successive
hints actually circle the spawn instead of bouncing in place. Blind
walk fallback included.
- runtime/reflex.js: when the scheduler is told to wander twice in a
row by gather.* recover hints, it now dispatches explore.far instead
so the bot actually leaves the patch it's stuck in. Resets the
consecutiveWanderHints counter on any success.
- runtime/reflex.js (sleep): no longer dispatches when the bot has
neither a bed in inventory NOR a known shelter/base location —
saved one dispatch + 5-min cooldown per restart at night.
- runtime/reflex.js (eat): inventory check + lastEatAt always updated
fix the eat-spam loop observed live (every tick fired "eat" → "no
food in inventory" → again).
- runtime/skills/chop-logs.js: recognise "no log within ..." as
no_target so the recover hint switches the bot to wander/explore.
npm test 124/124.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Recovers the bot from the live-server symptoms reported 2026-05-26:
1) constant supervisor reconnects, 2) chop "clicks once and stops",
3) sleep does nothing without a bed and so blocks night-skipping for
other players, 4) curriculum reflex always fell through to wander.
Supervisor (#38):
- runtime/watch-filter.js: pure predicate excluding *.test.js + the
supervisor itself; recursive:true so skills/ + social/ edits also
restart. Burned a working main once when test files counted toward
the rollback threshold.
- runtime/supervisor.js: watch-triggered restarts no longer count
toward the crash-loop rollback path. Watcher is now recursive.
Chop / mine (#39):
- runtime/actions.js + runtime/skills/gather-stone.js: replaced raw
pathfinder.goto + bot.dig with mineflayer-collectblock's
bot.collectBlock.collect — handles approach, repositioning, LoS,
dig and pickup as one primitive. Old version "swung once" because
GoalGetToBlock often parked the bot in leaves above the log.
Sleep + bed (#40):
- runtime/actions.js: sleepInBed now ALSO places a carried bed on
solid ground next to the bot and sleeps on it. Critical so the bot
stops blocking player night-skipping the moment it owns a bed.
Bed pipeline (#41):
- runtime/skills/gather-wool.js: gather.wool skill — mines wool block
if any nearby, otherwise shears or attacks the nearest sheep.
- runtime/skills/craft.js: craftBedSkill (any colour the bot has ≥3
wool of, plus 3 planks, plus a table).
- runtime/curriculum.js: new milestone survive.bed sits between
wood.tools and stone.32 so the bot gets a bed BEFORE everything else.
Test fixture updated to include a red_bed in post-survive.bed stages.
Village / shelter / wheat (#42, #43):
- runtime/skills/build-shelter.js: village.build-shelter — real 3×3×3
resumable hut blueprint around the recorded base, places one block
per loop, idempotent so an interrupted build resumes correctly,
marks each placed block in the owned-blocks ledger.
- runtime/skills/deposit-surplus.js: village.deposit-surplus opens
the nearest chest and transfers surplus stacks while keeping a
reserve of tools/food/bed.
- runtime/skills/farm-wheat.js: farm.wheat does one step per call
(till adjacent-to-water grass, plant seeds, or harvest ripe wheat).
- runtime/curriculum.js: village.shelter milestone after base-site.
Scheduler glitch (root of "always wander"):
- runtime/bot.js: curriculum + locations are now computed BEFORE
runTick. Previously they were stamped AFTER, so reflex.js saw
snapshot.curriculum=undefined every tick and fell through to the
wander fallback. Verified live: scheduler now dispatches
gather.logs/gather.stone/craft.* by id via runSkill.
Eat-spam:
- runtime/reflex.js: eatReflex now checks inventory for actual food
and updates lastEatAt on EVERY dispatch (not only successes), so a
failed eat respects the 5 s cooldown instead of firing every tick.
npm test 123/123. Validated live on play.xmatic.team (curriculum
dispatched gather.logs via runSkill, recover hint switched to wander
when no log in range).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three closures of remaining PRD follow-ups, one merge:
1. Reflex scheduler now drives behaviour from the curriculum.
- reflex.js: replaced ad-hoc techTreeReflex + autonomousReflex with
curriculumReflex that dispatches the skill suggested by
snapshot.curriculum.plan via runSkill. Per-skill backoff for
missing_tool / missing_material / no_target / no_food_source /
unsupported_version. recover() hint with `{hint:"wander"}` swaps
the next tick to wander for 60 s.
- Chain is now: defend > eat > sleep > curriculum > idle.
- reflex.test.js: 11 new tests covering busy/disconnected,
defend/eat preemption, dispatch by id, unknown-skill fallback,
per-skill + wander-hint backoffs, onComplete updating backoff.
2. Pi escalation for ADDRESSED_BANTER with hard rate limit.
- bot.js: when generateReply returns {escalate:true}, spawn askPi
with bot state + last 5 lines from that speaker (redacted via
chatMemory). Reply capped at 200 chars, sent as one chat line.
- Rate cap: 6 calls/hour, 90 s min gap. Suppressed escalations
log once and silently drop.
3. Phase 4 substrate.
- runtime/locations.js: atomic JSON store
(state/<host>/locations.json) with setLocation / getLocation /
nearestLocation / removeLocation; 6 tests.
- runtime/base-site.js: scoreCurrentPosition(bot) + pure scoreSite
bundle (wood / stone / water / flatness / no-players /
no-foreign-builds, owned-blocks excluded from claim penalty);
6 tests.
- runtime/skills/choose-base.js: village.choose-base skill — scores
the current spot, writes locations.base if score ≥ 8, otherwise
returns code:"too_weak" with a wander recover hint.
- curriculum.js: new final milestone village.base-site fires
village.choose-base until a base location exists.
- bot.js: stamps snapshot.locations from listLocations() each tick
so the curriculum can read it without coupling to disk.
docs/runtime.md updated with three new sections.
npm test now 116/116.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two safety rails on the unattended self-improvement loop:
1. scripts/edit-scope.js + .test.js: pure helpers that parse
`editScope: [...]` out of a proposal frontmatter (the field Phase 6
started writing) and validate a list of changed files against it.
13 tests covering null/missing/malformed frontmatter, directory
prefix matching, exact-file matching, default-scope fallback.
2. scripts/auto-patch.js:
- reads editScope from the proposal (falls back to ["runtime/"])
- injects the allowed paths into the Pi prompt so Pi knows the
boundaries up front
- validates the diff against scope + auto-allows any
runtime/**/*.test.js files Pi added
- runs `npm test` on the patched branch BEFORE cherry-picking;
refuses to land a patch that breaks the suite
Closes the "Smoke checks run before applying patch" item from
plans/autonomous-survival-bot-prd.md §7 Phase 6. npm test now 92/92.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 7 of plans/autonomous-survival-bot-prd.md. Five small modules
that close the recurring "shared state" and "version-pinned list"
failure modes the PRD flags in §7 and §5.4.
New:
- runtime/movement-profiles.js: named profiles (GATHER, TRAVEL, FLEE,
BUILD, RETURN_TO_BASE) as pure descriptors via PROFILE_DEFAULTS,
plus applyProfile(profile, bot) that hands a fresh Movements to
pathfinder. Avoids the "flee left canDig=false on the shared
Movements, next chop got stuck in canopy" regression.
- runtime/owned-blocks.js: JSONL ledger of blocks this bot placed/
removed (state/<host>/owned-blocks.jsonl); isOwned({x,y,z}) for
O(1) lookups; ensureDir() makes the parent dir lazily.
- runtime/claim-avoidance.js: classifyArea({blocks, isOwned}) returns
player_build / natural_or_owned / insufficient_data based on
man-made block density vs ownership ratio; shouldAvoid(area) helper.
Designed for gather/place skills to call before touching contested
area.
- runtime/skills/compat.test.js: runs runtime/skills/groups.js against
real minecraft-data registries for 1.18.2, 1.20.4, 1.21.5; spot-
checks that pale_oak_log only appears on 1.21+ etc.
- runtime/compat.test.js: 10 tests covering movement descriptors,
isManMadeBlockName, classifyArea, owned-blocks markPlaced/dedup/
isOwned/markRemoved.
npm test now 79/79.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 6 of plans/autonomous-survival-bot-prd.md. Expand the
self-improvement loop so the bot can spot and report no-progress
stagnation, not just exception-class failures.
New:
- runtime/stuck-incident.js: detector fires a structured proposal when
the same noProgressReason persists past 5 min (cooldown 30 min).
Body includes runtimeState, milestone, suggested skill, slim
snapshot, last action result, per-skill success/failure metrics
and a forbidden-paths list. Pure module — caller (bot.js) writes
the proposal.
- runtime/skill-metrics.js: in-memory per-skill ok/fail counters
surfaced on snapshot.skillMetrics for the TUI and the incident
body.
- runtime/stuck-incident.test.js: 6 tests covering null reason,
threshold gating, cooldown, reason change resetting the timer,
body composition and metrics snapshot.
Wiring:
- runtime/state-store.js: writeProposal accepts {editScope: string[]}
and persists it in the frontmatter; readProposalEditScope() reads
it back so future auto-patch.js can refuse cherry-picks that touch
other areas.
- runtime/bot.js: tick() invokes the stuck detector each tick,
records skill ok/fail via skillMetrics, stamps snapshot.skillMetrics
and writes the stuck proposal via writeProposal({editScope}).
dispatchAction now records into skillMetrics for both the
resolved-result and the exception path.
npm test now 46/46.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 5 of plans/autonomous-survival-bot-prd.md. Make the bot feel
present in chat without ever becoming a command executor.
New: runtime/social/
- intent.js: classifyIntent({text, botName}) returns one of GREETING /
STATUS_QUESTION / ADDRESSED_BANTER / COMMAND_LIKE / UNSAFE_REQUEST /
AMBIENT. Unicode-aware word boundaries so cyrillic + latin both work
("Привет всем" → GREETING, "build me a tower" → AMBIENT unless
addressed).
- reply.js: generateReply({intent, speaker, snapshot, diaryTail}) →
short templated response, or {send: null, escalate: true} for the
caller to decide whether to spend Pi tokens.
- memory.js: createChatMemory() — per-speaker LRU buffer of recent
lines; redacts password / api_key / JWT-shaped tokens at append
time, so the buffer can be safely fed back into any future prompt.
- social.test.js: 12 tests (intent edges, memory eviction, redaction,
reply routing). npm test now 40/40.
state-store.js additions:
- readDiaryTail(n) — reads the last N lines of today's diary; used by
status replies.
- writeEscalation({from, request, whyUnsure, wouldHave}) /
listEscalations() — JSONL log under state/<host>/escalations.jsonl
for UNSAFE_REQUEST classifications and future operator review.
bot.js: handleChat() now routes through social/intent + social/reply
(replacing the Phase-0 inline regexes), records every line into
chatMemory, and writes an escalation when classifyIntent returns
UNSAFE_REQUEST. Command-like notice + dialog-only behaviour from
Phase 0 are preserved.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 3 of plans/autonomous-survival-bot-prd.md. Gives the bot a
deterministic path from empty inventory through stone-tier tools and
basic storage, without an LLM call per tick.
New:
- runtime/curriculum.js: ordered milestone chooser
(wood.16 → wood.planks-and-sticks → wood.tools → stone.32 →
stone.tools → food.basic → storage.chest → shelter.torch). Each
milestone exposes isDone(inventory, snapshot) and suggest() returning
a { skillId } plan the scheduler can dispatch via runSkill. isDone
uses "stage reached" escapes so progress is monotonic — crafting
planks doesn't bounce the chooser back to "gather 16 logs".
- runtime/skills/gather-stone.js: gather.stone with pickaxe-required
precondition, blacklist on failed paths, registry-aware matching
(stone / cobblestone / deepslate / cobbled_deepslate / andesite /
diorite / granite).
- runtime/skills/craft.js: factory + concrete skills for craft.planks,
craft.sticks, craft.wooden-axe/-pickaxe/-sword, craft.stone-axe/
-pickaxe/-sword, craft.furnace, craft.chest, craft.torch (torch
requires coal or charcoal preflight).
Tests:
- runtime/curriculum.test.js: 14 tests covering chooser ordering,
per-milestone skill suggestion, inventoryFull threshold, monotonic
advancement across stage transitions.
- npm test now runs the full suite: 28/28 passing.
Wiring:
- runtime/bot.js: lastSnapshot.curriculum carries the next milestone
+ suggested skill on every tick; lastSnapshot.currentMilestone
prefers the curriculum title over the planner.md line.
- tui/tui.tsx: milestone line shows the curriculum's suggested skill
and an [inventory full] flag when isInventoryFull fires.
Reflex.js still calls actions.js directly; wiring the scheduler to
runSkill(plan.skillId, …) lands in Phase 4.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 2 of plans/autonomous-survival-bot-prd.md. Establishes the
composable skill contract from PRD §5.2 and ports three reference
skills so future phases can layer survival behaviour on top instead of
adding more ad-hoc branches to reflex.js.
New: runtime/skills/
- index.js: skill registry + runSkill(id, ctx, args) wrapper. Enforces
preconditions, hard timeout, normalises {ok, code, detail, worldDelta}
on every result, runs validate() and calls recover() on failure.
Stable failure codes live in RUNNER_CODES (unknown_skill,
precondition_failed, timeout, threw, validation_failed, done).
- groups.js: registry-derived item/block sets — logs/planks/sticks/beds
derived by suffix; foods intersects a curated allowlist with the live
bot.registry; axes/pickaxes/swords scoped to whatever the connected
server's item table actually ships. Empty set instead of throwing on
missing registry, so skills can emit code:"unsupported_version".
- chop-logs.js: gather.logs reference skill (wraps chopNearestTree).
- eat.js: survive.eat (wraps eatBestFood, preconditions check carrying
edible food from the registry-derived set).
- wander.js: explore.wander (wraps wander).
- contract.test.js + groups.test.js: 14 tests covering precondition
gating, timeout firing recover(), execute exceptions, validate
flipping ok→false, dynamic group filtering across mock registries.
package.json: `npm test` runs the new contract + groups suites.
docs/runtime.md: documents the skill contract, runner, dynamic groups
and the reference skills.
Reflex.js still calls actions.js directly — wiring the scheduler to
runSkill() lands in later phases when the survival curriculum kicks in.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 1 of plans/autonomous-survival-bot-prd.md. The bot must always be
able to answer "what am I doing and why am I not doing more?" without
parsing the log stream.
New modules:
- runtime/state.js: pure FSM classifier emitting emergency / working /
recovering / planning / social / idle from snapshot + reflex context.
- runtime/no-progress.js: sliding-window detector that watches position
and inventory; when both are unchanged for 60 s+, emits one stable
reason code from REASONS (waiting_for_day, night_hostile_nearby,
no_food_source, inventory_full, no_reachable_target, planner_empty,
awaiting_action_cooldown).
- runtime/viewer.js: optional prismarine-viewer launcher behind
VIEWER_PORT. Lazy import so the dep is not required by default.
Wiring:
- runtime/bot.js: tick() now computes runtimeState + noProgressReason
every tick and stamps them on the snapshot along with activeSkill,
currentMilestone (read from plan.md, cached 30 s), lastResult,
failuresByCode and lastEscalation.
- runtime/bot.js: dispatchAction records lastResult and lastFailureAt
for the recovering-state classifier.
- runtime/planner.js: exports isPlannerBusy(), readNextMilestone()
and planExists() so the runtime can show planning state + current
milestone without spawning extra Pi calls.
- runtime/config.js: adds VIEWER_PORT support.
TUI:
- tui/tui.tsx: StatusBar gains a state badge, current-skill row,
milestone row, no-progress reason warning, last-result line with
ok/fail color, failures-by-class summary and last-escalation age.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 0 of plans/autonomous-survival-bot-prd.md: change the product
direction from operator-driven remote control to autonomous survival
resident. MC chat is dialog-only for everyone, including
OPERATOR_USERNAMES — commands like come/follow/build/pause/stop are
recorded in the diary but not dispatched. TUI remains the only local
control plane.
Runtime changes:
- Remove operatorGoalReflex from reflex.js (the come-here chat command).
- Replace handleOperatorChat in bot.js with a dialog-only handleChat
that answers greetings/status questions and records command-like
verbs (en+ru) without dispatching them.
- Default MC_VERSION to "auto" in runtime/config.js; mineflayer
receives `false` to trigger version auto-detection.
- Update auto-escalation prompt's reflex chain summary.
Docs:
- AGENTS.md: product pivot notice up top; chat-driven scope-trust is
flagged as legacy/Pi-only.
- README.md / docs/runtime.md: replace operator-chat command list with
dialog-only description; update reflex chain summary.
- docs/roadmap.md: Phase 2/3 marked superseded by the PRD where they
assumed chat-driven control.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Operator PRDs and scratch implementation notes live under plans/ and
should not be committed to the repo.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the loop "стой и кидай proposals" → "копит ресурсы, строит,
работает к глобальной цели". Three pieces:
1. Crafting primitives (runtime/actions.js).
craftPlanks (4 per log, any wood type), craftSticks (4 per 2 planks),
placeCraftingTable (crafts a table from planks if needed + places at
reference block + reuses an existing table within 4 m), craftWoodenAxe,
craftWoodenPickaxe, craftWoodenSword. Each uses bot.recipesFor()
+ bot.craft() with a 15s timeout. Returns the same {ok, detail}
contract as the other actions.
inv.{getItemCount, getAnyPlanksCount, getAnyLogCount} helpers
exported so the reflex layer can read inventory cheaply without
pulling mineflayer state through every reducer.
2. Tech-tree reflex (runtime/reflex.js).
New techTreeReflex between sleep and autonomous. Inventory-driven
progression: log+0 planks → planks; planks+0 sticks → sticks;
planks+sticks+no axe → wooden_axe; +no pickaxe → wooden_pickaxe;
+no sword → wooden_sword. 5 s cooldown so we don't fire on every
tick.
Pure script, no LLM. The progression is exactly what a player
does in the first 10 min on a new world; making it scripted means
the bot never burns tokens on it.
3. LLM planner (runtime/planner.js).
Background timer (every 15 min, with a 30 s warm-up after start).
Reads goal.md + plan.md + a slim snapshot, prompts Pi to output a
fresh plan.md to stdout. Stripped of code fences and written
verbatim to state/<host>/plan.md. Capped at 16 KB.
The plan is markdown the operator can read or edit by hand. Numbered
milestones, ✓ prefix for completed ones, kept short. The reflex
layer doesn't auto-execute LLM text — but the planner sets the
long-horizon shape that future reflexes (build house, plant farm)
can read.
5 min timeout on the pi subprocess. If it crashes or times out, the
next 15-min tick just retries — no propagation to the reflex loop.
The progression now looks like, roughly:
chop log (autonomous) →
craft planks → craft sticks → wooden_axe (tech-tree) →
chop faster (autonomous, has axe now) →
wooden_pickaxe + wooden_sword (tech-tree) →
mine stone … (next PR: stone tools, farm site selection,
house frame)
Smoke-tested: all three modules import cleanly, exports check out.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Operator feedback: "бот должен быть полностью автономным — сам себя
улучшать и чинить, в этом и есть смысл; все что я вижу пока что он
стоит на месте и кидает proposals на каждый чих — это кардинально не
то что я хочу". Acted on:
1. Trigger filter — proposals only on real bugs.
runtime/bot.js classifies failure detail into bug / timeout /
feature-gap / other. The 5-in-a-row trigger fires only when the run
contains a bug (TypeError / Cannot read / is not defined …) OR is
entirely timeouts on the same operation. Feature gaps like "no
reachable log within 32 blocks", "no food in inventory", "no bed in
range", "no target in reach" are SKIPPED — the reflex layer routes
around them (noTreesUntil → wander, etc). The LLM has no business
patching code for missing inventory.
Threshold raised 3 → 5 in a row. Cooldown unchanged (30 min).
2. Auto-apply, no operator-in-the-loop.
New runtime/auto-improve.js polls proposals/ every 2s. When it sees
a new .md and 10s have passed since first sighting (debounce),
spawns scripts/auto-patch.js detached.
New scripts/auto-patch.js: refuses on dirty tree, moves proposal
pending → approved/, branches `auto/<slug>` off main, runs `pi -p`
with 10-min timeout. If Pi committed AND every changed file is
under runtime/ → cherry-picks onto main. Otherwise discards the
branch. No push, no PR. Audit trail in state/<host>/proposals/approved/.
Rate limit: 15-min cooldown between finished runs + 4/hour hard cap.
3. Auto-rollback on bad patches.
runtime/supervisor.js: when MAX_RESTARTS_PER_MINUTE is exceeded
AND `git log -1 HEAD` is younger than 15 min AND HEAD touched
runtime/, runs `git reset --hard HEAD~1`. Up to MAX_ROLLBACKS=3
lifetime, then exits 1 for manual investigation. Restart counters
are reset after a successful rollback so the next attempt isn't
immediately killed.
4. current-task.json slim.
No longer stores the full perception snapshot (was ~3 KB per write
× every action). Position only — sufficient as a resume anchor.
Slim snapshot still goes into the proposal markdown for context.
docs/runtime.md — rewrote the self-improvement section: full flow
diagram, classification rules, all rate-limit knobs, manual escape
hatches kept but documented as rarely-needed.
Also cleared 5 stale proposals from previous smoke tests so the first
production run isn't burning Pi tokens on stale bugs that have since
been fixed.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
User report 2026-05-25: launched 'npm run bot' fresh, MC server kicked
every login with "Игрок с данным никнеймом уже играет на сервере" and
the bot fell into a perpetual reconnect-then-kicked loop. Root cause:
a smoke-test supervisor from an earlier shell was still running in the
background, holding the pepa_bot session open. Two supervisors racing
on the same nickname is undefined behaviour from the server's side and
results in this exact failure mode.
Changes:
runtime/supervisor.js — acquires state/<host>/supervisor.pid before
spawning the child. If another supervisor is alive (kill -0 check), the
new one exits with a clear message telling the operator how to recover.
On SIGINT/SIGTERM/exit the lock is released; stale pidfiles are detected
when the recorded PID is no longer alive.
scripts/stop.sh — emergency cleanup helper:
- kills any supervisor or bot.js processes matching this repo
- removes pidfile + bot.sock
- reminds the operator to wait ~30s for the MC server to drop the old
session before re-launching
package.json — new `npm run stop` script.
Smoke-tested:
- first 'node runtime/supervisor.js' acquires lock, writes pid
- second call refuses with diagnostic message
- first SIGTERM ⇒ pidfile removed automatically
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the "bot stands on a tree doing nothing" problem reported live
when the operator launched the TUI after PR #6 landed. The reactive
chain (operator > defend > eat > sleep > idle) was passive by design:
day-time, full HP and food, no hostile within 4 m ⇒ every reflex
returned noop. The bot perched in dark-oak canopy and never moved.
Changes
runtime/actions.js:
- chopNearestTree: find any *_log within 32 blocks, equip best axe
(falls back to fists), path to the block, dig. Per-bot 5-min
blacklist of unreachable log positions so we don't grind on the
same impossible target.
- wander: pick a random offset 6-16 blocks away and path there.
- setMovementsForGather / setMovementsForTravel: every action that
uses pathfinder now sets its own Movements profile (canDig=true)
instead of inheriting whatever the previous caller left. The old
behaviour caused chop to inherit flee's canDig=false and get stuck
in the canopy.
- fleeFrom now uses canDig=true too — the user observed the bot
permanently stuck on a leaf block because escape required digging.
runtime/reflex.js:
- new autonomousReflex between sleep and idle. Cooldown 10s. Picks
chop when log count < 16, else wander. When chop reports "no
reachable log within 32 blocks" we switch to wander for 60s so we
don't re-fire chop against the same impossible position.
- defendReflex tightened: only flee when closest is ≤8m (or ≤12m
on low HP). Avoids the "82 distant hostiles ⇒ constant flee
loop" pathology observed at this spawn.
- flee cooldown: same mob name within 60s ⇒ noop, so we yield to
other reflexes if flee keeps timing out.
- sleepReflex retry cooldown raised 30s → 5min. Sleeping fails
permanently if no bed is around; the short retry blocked
autonomous behaviour every tick.
- ctx.lastReflex now records {name, label, ts} after each
dispatched/completed reflex so the TUI can show what the bot
just decided.
runtime/bot.js:
- per-tick snapshot adds lastReflex and busy fields for the TUI.
- maybeReplyToPlayer: light canned greetings (yo/hey/hi/привет)
to non-operators when they address the bot. 30s cooldown so we
don't spam.
tui/tui.tsx:
- status bar shows either "▸ busy: <label>" while an action is
in flight, or "last reflex: <name> (<label>) Ns ago" when idle.
Gives an at-a-glance answer to "what is the bot doing right now?"
Smoke-tested live (play.xmatic.team, 2026-05-25T13:30-13:39):
- bot did dispatch chop tree (oak_log at 617,82,95)
- real bug surfaced and 3-fail rule filed a proposal automatically
- after the runtime fix, chop returns "no reachable log" gracefully
- bot switched to wander on the next tick
- position moved from (623.31, 85, 95.12) to (623.41, 86.02, 96.7)
— the first observable movement in this dark-oak spawn
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Updates docs/runtime.md and README.md to match what's actually shipped:
- reflex chain priorities and what each body now dispatches
- operator chat command list (status, come, pause, resume, stop)
- automatic + manual Pi escalation paths and the no-code-change rule
- the full self-improvement loop end-to-end (detector → TUI approval
→ propose:apply → supervisor restart) with the rationale for the
manual propose:apply step
- new state files layout (proposals/, proposals/approved/, etc.)
- supervisor.js + bot:bare script flags
No code changes.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the self-improvement loop end-to-end:
reflex fails 3× → proposal file → operator approves in TUI →
`npm run propose:apply <file>` spawns Pi on a feature branch →
Pi commits the patch → supervisor watches runtime/*.js and
restarts the child on change.
runtime/state-store.js — atomic current-task.json writes, daily diary
append, proposals/ + proposals/approved/ helpers.
runtime/bot.js:
- ctx.dispatch writes current-task.json on start and updates it on
completion / failure / throw.
- failure tracker: 3 consecutive same-label failures → writeProposal()
with the snapshot, labels, and a suggested-next-step section.
30-min cooldown prevents proposal spam.
- on startup, surfaces resume info (previous task + pending proposal
count); on death, clears current-task.json + writes diary line.
- new IPC commands: PROPOSAL_LATEST returns the newest pending
proposal body; PROPOSAL_APPROVE moves it to proposals/approved/.
tui/tui.tsx — status bar shows `[proposals N, press y]` badge when
bot.pendingProposals > 0. Hotkey 'y' opens the proposal panel; 'y'
approves, 'n'/Esc closes.
scripts/propose-apply.js — given an approved proposal filename, creates
a `feat/proposal-<slug>` branch and spawns `pi -p` with the proposal
+ repo-conventions prompt. Refuses on dirty tree. No auto-push, no
auto-merge — operator reviews the diff and decides.
runtime/supervisor.js — forks bot.js as a child, watches runtime/*.js,
restarts on file change or on child exit code 42. Rate-limited at 5
restarts/minute. SIGINT/SIGTERM forward cleanly. `npm run bot` now
goes through the supervisor; `npm run bot:bare` skips it.
Smoke-tested: supervisor spawned, bot connected to MC, spawned at
expected coords, diary line written, state cleanup on SIGTERM correct.
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
runtime/actions.js — Mineflayer wrappers with hard timeouts and structured
{ok, detail} returns:
- attackNearest: equip best melee, lookAt, single swing per call
- fleeFrom: lazy-load pathfinder, walk N blocks away (canDig=false to
avoid burrowing through walls under panic)
- eatBestFood: scan inventory by FOOD_PRIORITY, equip + consume
- sleepInBed: find nearest placed bed within 16 blocks, path to it, sleep
- goTo: pathfinder.goto for operator come/follow
runtime/reflex.js — bodies now dispatch real actions via ctx.dispatch:
- operator-goal (highest): satisfy come/follow command
- defend: ≤4m attack, ≤12m + low HP/many hostiles flee
- eat: food < 16 + 5s cooldown
- sleep: night + safe + 30s retry cooldown
- idle: heartbeat every 20th tick
Reflex returns "skipped" when ctx.busy so we don't count busy ticks as
either productive or noop in the escalation counter.
runtime/bot.js:
- ctx.dispatch fire-and-forget wrapper with busy gate, onComplete hook
- consecutiveNoops counter; after ESCALATE_AFTER_NOOPS (=20, ~1 min at
tick=3s), askPi with the current snapshot. 10-min cooldown.
- operator chat handler: parses `<botname> <verb>` messages from
OPERATOR_USERNAMES. Verbs: status, pause, resume, stop, come.
- Death drops any pending operator goal.
Smoke-tested live against play.xmatic.team:25565: bot connected, logged
in via AuthMe, reflex chain dispatched flee/sleep, hard timeout fired
when pathfinder couldn't reach the flee target (expected — no usable
ground path in dark_forest at this spawn).
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(mindcraft-skills): hard timeout on every skill call
mc_avoid_enemies (and 7 other tools) wrapped only in safeCall without a
withTimeout. When mindcraft's underlying pathfinder/pvp goal couldn't be
satisfied, the call never resolved — the Pi tick loop blocked forever.
Observed live: mc_avoid_enemies pending >10 minutes after one mc_observe.
safeCall now takes timeoutMs (default 30s) and wraps withTimeout itself,
so every tool gets a hard ceiling. Per-tool overrides:
- goToPosition / goToNearestBlock: 120s / 90s (unchanged from before)
- defendSelf / avoidEnemies: 45s
- stay: secs*1000 + 10s
- craft / consume / pickup / place: 30s
- equip: 15s
collectBlock still uses its bespoke per-iter 75s loop.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(runtime): script-driven reflex daemon + Ink TUI dashboard
Pure-Pi runtime had three failure modes in practice:
- slow: 20-60s per decision because LLM was in the hot path
- expensive: every tick (defend, eat, idle) paid for a reasoning pass
- invisible: required tmux capture-pane to know what the bot was doing
New runtime/ layer is a long-running Node daemon that owns the MC
connection, ticks a priority-ordered reflex chain (defend > eat > sleep
> idle) with NO LLM in the hot path, and exposes status + commands over
a Unix-socket IPC. tui/ is an Ink dashboard that attaches over IPC and
can detach freely — multiple TUI clients can connect at once.
Pi/Codex are still available, but as on-demand escalation: TUI hotkey
'a' spawns `pi -p "<prompt>"` as a subprocess and streams stdout into
the dashboard. The self-improvement loop (proposals → operator approval
→ Pi-driven patch → hot reload) is documented in docs/runtime.md but
not yet wired.
Reflex bodies are stubs today — they log decisions but don't drive
Mineflayer actions yet. The priority chain, IPC contract, and TUI are
fully working; subsequent commits will fill in defend/eat/sleep bodies
and wire automatic escalation.
Run with `npm run bot` + `npm run tui`. Pi-only fallback stays at
`npm run agent`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
OpenClaw is a different paradigm than our Pi-based bridge — messaging-first,
marketplace of pre-built skills (astraopenclaw/minecraft-agent already exists),
self-extension via skill authoring. The operator wants to run an OpenClaw
instance side-by-side with pepa-pi-bot to compare which approach gets to a
visible village faster.
prompts/openclaw-seed.md: single founding-message prompt. Identity, server
credentials inline (secrets via .env, not echoed), goal (small village,
long-horizon), bootstrap checklist (install skill, connect, AuthMe handle,
30-60s autonomous tick, on-death recovery), self-extension permission, hard
safety rules duplicated from AGENTS.md, definition of success ("a week from
now there's a cluster of buildings attributable to you").
Includes operational notes on nickname conflict (only one bot can be on the
server at a time under pepa_bot; suggest pepa_claw for the OpenClaw side),
.env coordination, LLM provider mixing, and a short post-mortem comparison
checklist for after a few hours of both running.
Not invoked by anything in this repo — it's an artefact for cross-runtime
experimentation. Lives in prompts/ alongside the Pi prompts and the
codex-seed-knowledge.md prompt because that folder is the right home for
reusable prompts regardless of which agent consumes them.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Four targeted fixes for issues observed during the live autonomous test:
1. **bot.modes shim** (extensions/lib/mcdata.js)
Mindcraft skills.* call bot.modes.pause('cowardice') in 7+ places.
`mineflayer-modes` does not exist on npm — it's internal to Mindcraft.
attachPluginsAndInit now installs a no-op shim so skills.stay() /
.consume() / .defendSelf() etc. stop crashing with "pause undefined".
2. **Death + respawn handlers** (extensions/mineflayer-bridge.ts)
Subscribe to bot 'death' event: append diary line with position,
clear current-task.json so Pi doesn't resume a stale task referencing
inventory that no longer exists.
On 'spawn' within 5s of death: log the new respawn position to diary.
Live test had bot killed twice by zombies at night; the next Pi
prompt was unable to recover. With this it's now an explicit diary
line + clean task slate.
3. **Auto-defend reflex tick** (extensions/mineflayer-bridge.ts)
New setInterval(2s) that, when bot.health < 18 AND a hostile mob
(zombie/skeleton/creeper/spider/etc.) is within 6 blocks AND no
active world task, fires `bot.pvp.attack(nearest)` in the background.
No LLM call needed for instant self-defense — saves tokens and reacts
on mineflayer timescale (sub-second) rather than Pi loop timescale
(~10s+ to reason and dispatch). Throttled to once per 4s.
4. **mc_collect_block bulk rework** (extensions/mindcraft-skills.ts)
Live test showed count=1 succeeds in ~25s but count=8 hangs past
270s with identical blocks in range. Upstream collectBlock plugin
appears to drift on its block cache after the first dig in dense
terrain. Loop single-block collects in-tool instead (75s per iter,
re-pathfind on each iteration). Track per-iter success/failure,
abort after 3 consecutive failures, return aggregate count to the
agent so it can adapt instead of seeing a single failure.
Smoke test: both extensions load, bot connects, perception confirms
hostiles nearby and daylight safety check. No syntax/runtime errors.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Live test showed count=1 succeeds in ~25s but count=8 timed out at 90s
because gathering 8 logs naturally takes ~200s of pathing + dig + pickup
across the area. Fixed 90s ceiling was too tight for legitimate bulk
collection.
Scale timeout linearly: 30s overhead + 30s per requested block, capped
at 600s. count=1 → 60s, count=4 → 150s, count=8 → 270s, count=20 → 630s
(capped at 600). Catches real hangs (wrong name, unreachable) without
killing legitimate long collections in dense forest.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mineflayer's collectBlock plugin and pathfinder can hang indefinitely when:
- the requested block name doesn't match anything in range (e.g. asking for
"oak_log" when the only nearby logs are "dark_oak_log"),
- pathfinder cannot reach the goal but doesn't return a clean noPath,
- a path computation enters an infinite-search state in dense terrain.
This blocks the entire Pi tool loop — observed in a live test where
mc_collect_block("oak_log", 4) ran for 8+ minutes without ever returning,
leaving Pi unable to respond to chat or do anything else.
Add a 90s timeout for collectBlock/goToNearestBlock and 120s for goToPosition.
On timeout the tool throws a descriptive error so the agent learns to retry
with a different block name or position rather than waiting forever.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The bot was acting blind: it knew its own coordinates but nothing about
what was around it. mc_goto's over-strict safety guards refused every
real path. mc_dig was too low-level to drive a coherent farming loop.
The result: 4+ hours of "trying" with zero physical achievement.
This commit reframes the bridge around a perceive→decide→act loop using
proven primitives from Mindcraft (github.com/kolbytn/mindcraft, MIT —
LICENSE-MINDCRAFT vendored beside the library files).
Changes:
- extensions/lib/ (new, vendored from Mindcraft with attribution)
- world.js (431 LoC) — 21 perception functions: getNearbyBlockTypes,
getNearbyEntities, getInventoryCounts, getNearestBlock, getPosition,
getBiomeName, etc.
- skills.js (2093 LoC) — 30+ action primitives: collectBlock, placeBlock,
goToPosition, goToNearestBlock, craftRecipe, equip, consume,
defendSelf, avoidEnemies, pickupNearbyItems, stay, etc.
- mcdata.js (~600 LoC) — Mindcraft's mc-data adapter. Imports patched
to local paths; mineflayer-auto-eat removed (our installed 5.x has
a divergent API; skills.consume() falls back to bot.consume()).
Added attachPluginsAndInit(bot) export so mineflayer-bridge.ts can
wire plugins onto its externally-created bot.
- settings.js — minimal stub with farmer-bot defaults.
- extensions/mindcraft-skills.ts (new, 406 LoC) — Pi extension registering
15 high-level tools on top of the vendored library:
- Perception: mc_observe, mc_inventory, mc_nearby_blocks, mc_nearby_entities
- Action: mc_collect_block, mc_place_block, mc_go_to, mc_go_to_block,
mc_craft, mc_equip, mc_consume, mc_defend_self, mc_avoid_enemies,
mc_stay, mc_pickup_nearby
ES modules from extensions/lib/ are loaded via dynamic import() at
extension init so the cross-extension require()-race resolves cleanly.
- extensions/mineflayer-bridge.ts
- Expose the live Mineflayer bot on globalThis.__pepaPiBot so the
mindcraft-skills extension can use it (set on connect, cleared on
error/end/manual disconnect).
- Call attachPluginsAndInit(nextBot) right after createBot to load
pathfinder, pvp, collectblock, armorManager and prime
minecraft-data once login completes.
- package.json — new runtime deps: minecraft-data, vec3,
mineflayer-pvp, prismarine-item.
- AGENTS.md — new "Perception → decision → action" section before
"Your tools right now" with full tool catalog and a deprecation
note for the broken mc_goto / mc_build_pyramid_5x5 / low-level
mc_dig from the old bridge.
Smoke test (medium thinking): bot called mc_observe and received a
real JSON snapshot — nearbyBlocks listed coal_ore, oak_log, water,
sand; nearbyEntityTypes listed creeper, zombie, pillager, skeleton.
The bot can finally see what it could not see this morning.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
A long-form prompt for a long-context autonomous coding agent (Codex Pro,
Claude Sonnet w/ repo access, etc.) — NOT a Pi prompt. Goal: produce a
comprehensive docs/ knowledge base so the bot doesn't have to rediscover
Mineflayer API surface and core Minecraft mechanics (mobs, biomes,
recipes, ore Y-levels, farming, breeding) every time it tries something
new.
Design constraints baked into the prompt:
- PR-only workflow. Worker agent operates on a feat/knowledge-base
branch; main stays untouched so the live bot is unaffected until the
operator reviews and merges.
- docs/ only. Never write skills/ — that's the bot's notebook; pre-
writing procedural skills kills emergence. Reference material is the
textbook; the bot stays the author of its own procedures.
- One-line AGENTS.md addition pointing to docs/, no broader policy
rewrite. Behaviour change is "consult docs/ before I'll try to learn".
- package.json gets four universal-useful plugins added (collectblock,
auto-eat, tool, armor-manager); statemachine/pvp/blockfinder/viewer
are left for the bot to opt into.
- Concrete definition of done, scope estimate (10-30h), review
checklist for spot-checking hallucinations before merge.
- Re-run triggers documented (MC version bump, Mineflayer major,
new must-have plugin).
Included so future operators / forks can repeat this kind of one-off
seeding without re-deriving the prompt. Lives alongside the Pi prompts
in prompts/, even though it targets a different runtime — the
prompts/ folder is the right home for "reusable prompts" regardless of
which agent consumes them.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three things that together turn the bot from a reactive chat agent into
a goal-driven autonomous one.
1. Memory model (docs/memory-model.md, new). Formal split:
- SHARED knowledge — skills/, extensions/, prompts/, docs/,
.pi/settings.json — committed, community-improvable, portable to any
server.
- PERSONAL memory — state/<MC_HOST>/ — gitignored, per-instance, per-
server. Survives restarts (local disk), doesn't survive a re-clone
(deliberately). Holds goal.md, plan.md, current-task.json,
locations.json, diary/, inventory-log.jsonl, escalations.
Covers resume-after-restart protocol, what "abstract a lesson into a
skill" means, and the two anti-patterns (committing state, gitignoring
shared knowledge).
2. AGENTS.md changes:
- New section "Long-term goal and personal memory" wiring AGENTS.md
directly into state/<MC_HOST>/goal.md + current-task.json with a
pointer to docs/memory-model.md.
- Operating principle #4 ("I'll try to learn") rewritten with
**bias to action**: a pending stub is now a last resort, not a
default. Operator-trusted requests are themselves approval — bot
does not write a stub and wait for a separate "go".
Rationale: today's pyramid task got stuck because the bot wrote
a careful "pending" stub and waited; the operator had to send
"ты ждешь одобрения? можешь стартовать!" before any action. That
extra round-trip is the reflex this rewrite removes.
- Operating principle #5 ("live your best life when idle") expanded
to "goal-driven autonomy" with an explicit 5-level priority order
(operator task > non-op reply > resume current-task.json > next
plan milestone > decompose goal). Memory protocol made concrete:
write current-task.json before every meaningful action, append to
diary, keep locations.json fresh, tick off plan.md.
3. prompts/live-your-life.md (new). Canonical kickoff to switch the
bot into autonomous mode. Numbered concrete asks (re-read three
docs, write plan.md, implement memory protocol, implement
resume-on-restart, start). Includes a "plan.md draft for review"
gate so the operator can shape direction without micromanaging
execution. Designed to be sent after Phase 0/1/operator-trust are
stable and a goal.md exists for the target server.
Companion seed (local-only, NOT in this commit because gitignored):
state/play.xmatic.team_25565/goal.md — "build a small village and
survive long-term, live like a farmer". Lives only on the operator's
machine; a fresh clone won't see it.
README and roadmap updated with the new Phase 3 status (🌱 → 🌿
kickoff) and pointers to the new memory-model doc.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-tier chat trust:
- Anyone in OPERATOR_USERNAMES (comma-separated, .env-only) is SCOPE-trusted.
The bot skips the "out of scope / not sure where" escalation reflex for
these users and instead applies "I'll try to learn" (Operating principle
#4): attempt, codify into a new skill, or reply with a concrete reason.
- Hard safety rules (no OP, no breaking other players' builds, no .env
leak, no chat spam, no destructive bash) remain ABSOLUTE. Operators get
the same refusal + escalation as anyone else for safety-borderline
requests — with slightly pointed wording, because they should know better.
- No transitive trust: chat-based "trust X for the next hour" / "make Y
an op" requests are themselves safety escalations. Op membership only
flows through .env on disk.
Security caveat documented in .env.example: nickname-based trust is only
safe on servers with identity protection (online-mode UUID or AuthMe).
On pure cracked servers OPERATOR_USERNAMES must stay empty.
- AGENTS.md: new Identity field for OPERATOR_USERNAMES; new Operating
principle #6 "Trusted operators" with the scope-vs-safety split; old
escalation principle renumbered to #7; Control channel section
rewritten with primary/secondary trust distinction.
- .env.example: OPERATOR_USERNAMES placeholder with multi-paragraph
security note covering when the model is and isn't safe.
- prompts/grant-op-trust.md: canonical implementation prompt for the
next Pi pass — re-read AGENTS.md, wire isOperator() into the bridge's
escalation flow, codify into skills/operator-trust.md, reload bridge,
verify with two concrete chat replays (scope vs safety).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 0 (body) is done. The next steps shouldn't be guessed prompt-by-prompt —
write down the order, the judgement principles, and the next concrete
session prompt, so the bot has a coherent direction and the human can
hand it off in one message.
- docs/roadmap.md (new): six phases, each with status, scope, and stretch.
Phase 0 = 🌳 done, Phase 1 = 🌿 in progress, the rest = 🌱.
Explicit non-goals (no PvP, no OP, no cross-server identity).
- AGENTS.md: First-objective section collapsed to a pointer at the
onboarding skill (it's been done). New "What to do, in priority order"
summary citing the roadmap. New top-level "Operating principles"
section: presence, bounded reconnect, hold focus, "I'll try to learn"
reflex, idle = best-life mode, escalate destructive doubt with a
JSONL log under state/<host>/escalations.jsonl.
- prompts/awake-and-live.md (new): canonical kickoff prompt for the
next session. Scopes itself explicitly to phases 1+5+6 and excludes
locomotion (phase 2 needs care, separate session).
- README Status: 🌳 Phase 0 done / 🌱 Phase 1 in progress, links to
roadmap and operating principles.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The whole point of this project is that the agent's growth is shareable.
If skills and extensions silently land in ~/.pi/ on the maintainer's
laptop, every clone starts from zero and the repo becomes a fancy
README. Fix that with an explicit hard rule and a contributor guide.
- AGENTS.md: new "Artifact location — hard rule" subsection spelling out
that skills/extensions/prompts/state/.pi-settings ALL live in this repo,
never in ~/.pi/. Pi's own built-in skills (skill-creator, etc.) stay
user-global; the agent may use them, but their *output* must land here.
- README: new "Everything in the repo" subsection covering the same rule
in user-facing language, plus a pointer to CONTRIBUTING.md.
- CONTRIBUTING.md (new): skill/extension formats, server-agnostic and
no-secrets requirements, smoke-test recipe, PR checklist.
- .gitignore: switch from blanket `.pi/` ignore to `.pi/*` + explicit
un-ignore of `.pi/settings.json`, so project Pi config is reproducible.
- "Don't push without operator confirmation" → "without human
confirmation via the repo" — consistent with the new control model.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Pi only acts when prompted. New repo users (and the maintainer's future
self) shouldn't have to invent the kickoff message — pin it.
- prompts/bootstrap.md: the canonical first-run message, with rationale
for each clause and guidance for shorter subsequent prompts
- README quickstart: new "Send the first message" subsection that quotes
the bootstrap prompt verbatim and explains what the agent does next
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>