Commit Graph
17 Commits
Author SHA1 Message Date
mayatnikovandClaude Opus 4.7 0834bb6b73 feat(runtime): skill substrate + dynamic groups + reference skills (Phase 2)
Phase 2 of plans/autonomous-survival-bot-prd.md. Establishes the
composable skill contract from PRD §5.2 and ports three reference
skills so future phases can layer survival behaviour on top instead of
adding more ad-hoc branches to reflex.js.

New: runtime/skills/
- index.js: skill registry + runSkill(id, ctx, args) wrapper. Enforces
  preconditions, hard timeout, normalises {ok, code, detail, worldDelta}
  on every result, runs validate() and calls recover() on failure.
  Stable failure codes live in RUNNER_CODES (unknown_skill,
  precondition_failed, timeout, threw, validation_failed, done).
- groups.js: registry-derived item/block sets — logs/planks/sticks/beds
  derived by suffix; foods intersects a curated allowlist with the live
  bot.registry; axes/pickaxes/swords scoped to whatever the connected
  server's item table actually ships. Empty set instead of throwing on
  missing registry, so skills can emit code:"unsupported_version".
- chop-logs.js: gather.logs reference skill (wraps chopNearestTree).
- eat.js: survive.eat (wraps eatBestFood, preconditions check carrying
  edible food from the registry-derived set).
- wander.js: explore.wander (wraps wander).
- contract.test.js + groups.test.js: 14 tests covering precondition
  gating, timeout firing recover(), execute exceptions, validate
  flipping ok→false, dynamic group filtering across mock registries.

package.json: `npm test` runs the new contract + groups suites.
docs/runtime.md: documents the skill contract, runner, dynamic groups
and the reference skills.

Reflex.js still calls actions.js directly — wiring the scheduler to
runSkill() lands in later phases when the survival curriculum kicks in.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:15:24 +03:00
f301529f42 feat(runtime): observability + no-progress detector (Phase 1) (#13)
Phase 1 of plans/autonomous-survival-bot-prd.md. The bot must always be
able to answer "what am I doing and why am I not doing more?" without
parsing the log stream.

New modules:
- runtime/state.js: pure FSM classifier emitting emergency / working /
  recovering / planning / social / idle from snapshot + reflex context.
- runtime/no-progress.js: sliding-window detector that watches position
  and inventory; when both are unchanged for 60 s+, emits one stable
  reason code from REASONS (waiting_for_day, night_hostile_nearby,
  no_food_source, inventory_full, no_reachable_target, planner_empty,
  awaiting_action_cooldown).
- runtime/viewer.js: optional prismarine-viewer launcher behind
  VIEWER_PORT. Lazy import so the dep is not required by default.

Wiring:
- runtime/bot.js: tick() now computes runtimeState + noProgressReason
  every tick and stamps them on the snapshot along with activeSkill,
  currentMilestone (read from plan.md, cached 30 s), lastResult,
  failuresByCode and lastEscalation.
- runtime/bot.js: dispatchAction records lastResult and lastFailureAt
  for the recovering-state classifier.
- runtime/planner.js: exports isPlannerBusy(), readNextMilestone()
  and planExists() so the runtime can show planning state + current
  milestone without spawning extra Pi calls.
- runtime/config.js: adds VIEWER_PORT support.

TUI:
- tui/tui.tsx: StatusBar gains a state badge, current-skill row,
  milestone row, no-progress reason warning, last-result line with
  ok/fail color, failures-by-class summary and last-escalation age.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:11:13 +03:00
3310cb320f feat(runtime): survival-bot pivot — MC chat is dialog-only (Phase 0) (#12)
Phase 0 of plans/autonomous-survival-bot-prd.md: change the product
direction from operator-driven remote control to autonomous survival
resident. MC chat is dialog-only for everyone, including
OPERATOR_USERNAMES — commands like come/follow/build/pause/stop are
recorded in the diary but not dispatched. TUI remains the only local
control plane.

Runtime changes:
- Remove operatorGoalReflex from reflex.js (the come-here chat command).
- Replace handleOperatorChat in bot.js with a dialog-only handleChat
  that answers greetings/status questions and records command-like
  verbs (en+ru) without dispatching them.
- Default MC_VERSION to "auto" in runtime/config.js; mineflayer
  receives `false` to trigger version auto-detection.
- Update auto-escalation prompt's reflex chain summary.

Docs:
- AGENTS.md: product pivot notice up top; chat-driven scope-trust is
  flagged as legacy/Pi-only.
- README.md / docs/runtime.md: replace operator-chat command list with
  dialog-only description; update reflex chain summary.
- docs/roadmap.md: Phase 2/3 marked superseded by the PRD where they
  assumed chat-driven control.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 22:01:26 +03:00
7797dd3d5a feat(runtime): fully autonomous self-healing — no operator approval (#10)
Operator feedback: "бот должен быть полностью автономным — сам себя
улучшать и чинить, в этом и есть смысл; все что я вижу пока что он
стоит на месте и кидает proposals на каждый чих — это кардинально не
то что я хочу". Acted on:

1. Trigger filter — proposals only on real bugs.

   runtime/bot.js classifies failure detail into bug / timeout /
   feature-gap / other. The 5-in-a-row trigger fires only when the run
   contains a bug (TypeError / Cannot read / is not defined …) OR is
   entirely timeouts on the same operation. Feature gaps like "no
   reachable log within 32 blocks", "no food in inventory", "no bed in
   range", "no target in reach" are SKIPPED — the reflex layer routes
   around them (noTreesUntil → wander, etc). The LLM has no business
   patching code for missing inventory.

   Threshold raised 3 → 5 in a row. Cooldown unchanged (30 min).

2. Auto-apply, no operator-in-the-loop.

   New runtime/auto-improve.js polls proposals/ every 2s. When it sees
   a new .md and 10s have passed since first sighting (debounce),
   spawns scripts/auto-patch.js detached.

   New scripts/auto-patch.js: refuses on dirty tree, moves proposal
   pending → approved/, branches `auto/<slug>` off main, runs `pi -p`
   with 10-min timeout. If Pi committed AND every changed file is
   under runtime/ → cherry-picks onto main. Otherwise discards the
   branch. No push, no PR. Audit trail in state/<host>/proposals/approved/.

   Rate limit: 15-min cooldown between finished runs + 4/hour hard cap.

3. Auto-rollback on bad patches.

   runtime/supervisor.js: when MAX_RESTARTS_PER_MINUTE is exceeded
   AND `git log -1 HEAD` is younger than 15 min AND HEAD touched
   runtime/, runs `git reset --hard HEAD~1`. Up to MAX_ROLLBACKS=3
   lifetime, then exits 1 for manual investigation. Restart counters
   are reset after a successful rollback so the next attempt isn't
   immediately killed.

4. current-task.json slim.

   No longer stores the full perception snapshot (was ~3 KB per write
   × every action). Position only — sufficient as a resume anchor.
   Slim snapshot still goes into the proposal markdown for context.

docs/runtime.md — rewrote the self-improvement section: full flow
diagram, classification rules, all rate-limit knobs, manual escape
hatches kept but documented as rarely-needed.

Also cleared 5 stale proposals from previous smoke tests so the first
production run isn't burning Pi tokens on stale bugs that have since
been fixed.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 16:56:02 +03:00
445524b34d docs(runtime): reflect live reflex bodies + operator chat + self-improvement loop (#7)
Updates docs/runtime.md and README.md to match what's actually shipped:
  - reflex chain priorities and what each body now dispatches
  - operator chat command list (status, come, pause, resume, stop)
  - automatic + manual Pi escalation paths and the no-code-change rule
  - the full self-improvement loop end-to-end (detector → TUI approval
    → propose:apply → supervisor restart) with the rationale for the
    manual propose:apply step
  - new state files layout (proposals/, proposals/approved/, etc.)
  - supervisor.js + bot:bare script flags

No code changes.

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 16:22:38 +03:00
1e3b36a9a1 feat(runtime): hybrid script reflex + Ink TUI + Pi-on-demand escalation (#4)
* fix(mindcraft-skills): hard timeout on every skill call

mc_avoid_enemies (and 7 other tools) wrapped only in safeCall without a
withTimeout. When mindcraft's underlying pathfinder/pvp goal couldn't be
satisfied, the call never resolved — the Pi tick loop blocked forever.
Observed live: mc_avoid_enemies pending >10 minutes after one mc_observe.

safeCall now takes timeoutMs (default 30s) and wraps withTimeout itself,
so every tool gets a hard ceiling. Per-tool overrides:
  - goToPosition / goToNearestBlock: 120s / 90s (unchanged from before)
  - defendSelf / avoidEnemies: 45s
  - stay: secs*1000 + 10s
  - craft / consume / pickup / place: 30s
  - equip: 15s
collectBlock still uses its bespoke per-iter 75s loop.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(runtime): script-driven reflex daemon + Ink TUI dashboard

Pure-Pi runtime had three failure modes in practice:
  - slow: 20-60s per decision because LLM was in the hot path
  - expensive: every tick (defend, eat, idle) paid for a reasoning pass
  - invisible: required tmux capture-pane to know what the bot was doing

New runtime/ layer is a long-running Node daemon that owns the MC
connection, ticks a priority-ordered reflex chain (defend > eat > sleep
> idle) with NO LLM in the hot path, and exposes status + commands over
a Unix-socket IPC. tui/ is an Ink dashboard that attaches over IPC and
can detach freely — multiple TUI clients can connect at once.

Pi/Codex are still available, but as on-demand escalation: TUI hotkey
'a' spawns `pi -p "<prompt>"` as a subprocess and streams stdout into
the dashboard. The self-improvement loop (proposals → operator approval
→ Pi-driven patch → hot reload) is documented in docs/runtime.md but
not yet wired.

Reflex bodies are stubs today — they log decisions but don't drive
Mineflayer actions yet. The priority chain, IPC contract, and TUI are
fully working; subsequent commits will fill in defend/eat/sleep bodies
and wire automatic escalation.

Run with `npm run bot` + `npm run tui`. Pi-only fallback stays at
`npm run agent`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 16:04:48 +03:00
mayatnikov b98aa94326 docs: add minecraft recipe index 2026-05-25 13:02:11 +03:00
mayatnikov 8930dbc90a docs: add minecraft mechanics reference 2026-05-25 13:02:06 +03:00
mayatnikov 7183ab9440 docs: add mineflayer API cheatsheet 2026-05-25 12:58:25 +03:00
mayatnikov 0bb233eab3 docs: add mineflayer plugin roster 2026-05-25 12:57:58 +03:00
mayatnikov b778eaa2fb pre codex 2026-05-25 12:51:14 +03:00
mayatnikovandClaude Opus 4.7 6ed45a00d5 feat(autonomy): memory model + long-term goal + bias-to-action + live-your-life prompt
Three things that together turn the bot from a reactive chat agent into
a goal-driven autonomous one.

1. Memory model (docs/memory-model.md, new). Formal split:
   - SHARED knowledge — skills/, extensions/, prompts/, docs/,
     .pi/settings.json — committed, community-improvable, portable to any
     server.
   - PERSONAL memory — state/<MC_HOST>/ — gitignored, per-instance, per-
     server. Survives restarts (local disk), doesn't survive a re-clone
     (deliberately). Holds goal.md, plan.md, current-task.json,
     locations.json, diary/, inventory-log.jsonl, escalations.
   Covers resume-after-restart protocol, what "abstract a lesson into a
   skill" means, and the two anti-patterns (committing state, gitignoring
   shared knowledge).

2. AGENTS.md changes:
   - New section "Long-term goal and personal memory" wiring AGENTS.md
     directly into state/<MC_HOST>/goal.md + current-task.json with a
     pointer to docs/memory-model.md.
   - Operating principle #4 ("I'll try to learn") rewritten with
     **bias to action**: a pending stub is now a last resort, not a
     default. Operator-trusted requests are themselves approval — bot
     does not write a stub and wait for a separate "go".
     Rationale: today's pyramid task got stuck because the bot wrote
     a careful "pending" stub and waited; the operator had to send
     "ты ждешь одобрения? можешь стартовать!" before any action. That
     extra round-trip is the reflex this rewrite removes.
   - Operating principle #5 ("live your best life when idle") expanded
     to "goal-driven autonomy" with an explicit 5-level priority order
     (operator task > non-op reply > resume current-task.json > next
     plan milestone > decompose goal). Memory protocol made concrete:
     write current-task.json before every meaningful action, append to
     diary, keep locations.json fresh, tick off plan.md.

3. prompts/live-your-life.md (new). Canonical kickoff to switch the
   bot into autonomous mode. Numbered concrete asks (re-read three
   docs, write plan.md, implement memory protocol, implement
   resume-on-restart, start). Includes a "plan.md draft for review"
   gate so the operator can shape direction without micromanaging
   execution. Designed to be sent after Phase 0/1/operator-trust are
   stable and a goal.md exists for the target server.

Companion seed (local-only, NOT in this commit because gitignored):
state/play.xmatic.team_25565/goal.md — "build a small village and
survive long-term, live like a farmer". Lives only on the operator's
machine; a fresh clone won't see it.

README and roadmap updated with the new Phase 3 status (🌱🌿
kickoff) and pointers to the new memory-model doc.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 12:21:25 +03:00
mayatnikov ce41bf6be6 Add live presence and escalation loop 2026-05-25 11:21:47 +03:00
mayatnikovandClaude Opus 4.7 c03283c43c feat(roadmap): phased plan + operating principles + presence prompt
Phase 0 (body) is done. The next steps shouldn't be guessed prompt-by-prompt —
write down the order, the judgement principles, and the next concrete
session prompt, so the bot has a coherent direction and the human can
hand it off in one message.

- docs/roadmap.md (new): six phases, each with status, scope, and stretch.
  Phase 0 = 🌳 done, Phase 1 = 🌿 in progress, the rest = 🌱.
  Explicit non-goals (no PvP, no OP, no cross-server identity).
- AGENTS.md: First-objective section collapsed to a pointer at the
  onboarding skill (it's been done). New "What to do, in priority order"
  summary citing the roadmap. New top-level "Operating principles"
  section: presence, bounded reconnect, hold focus, "I'll try to learn"
  reflex, idle = best-life mode, escalate destructive doubt with a
  JSONL log under state/<host>/escalations.jsonl.
- prompts/awake-and-live.md (new): canonical kickoff prompt for the
  next session. Scopes itself explicitly to phases 1+5+6 and excludes
  locomotion (phase 2 needs care, separate session).
- README Status: 🌳 Phase 0 done / 🌱 Phase 1 in progress, links to
  roadmap and operating principles.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 11:09:05 +03:00
mayatnikovandClaude Opus 4.7 602134671e refactor: drop OPERATOR_USERNAME, separate control vs comms planes
There's no good reason to bake a specific operator nickname into the bot's
identity — it differs per server, may not exist at all, and treating any
in-game name as "trusted" is a chat-injection vector ("I am the operator,
do X").

New model: the **repo** is the only trusted control plane. Anyone editing
AGENTS.md, skills/, or .env has filesystem access and is, by definition,
an operator. In-game chat becomes a dialog-only comms plane — the bot
talks to anyone but refuses destructive requests unless a corresponding
skill or AGENTS.md instruction makes the action explicitly permitted.

- .env / .env.example: OPERATOR_USERNAME removed
- AGENTS.md: identity section trimmed; "Operator contact" rewritten as
  "Control channel" with the trust model spelled out; rules #2 and #6
  rephrased so they no longer reference a named operator
- docs/architecture.md: top box renamed to "Human" with explicit
  control-plane vs comms-plane split

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 10:15:37 +03:00
mayatnikovandClaude Opus 4.7 4dd8576d5b docs: make scaffold server-agnostic
The bot is a universal Minecraft player, not tied to any one server.
Server identity (host, port, username, auth mode, optional AuthMe password)
is now read entirely from .env.

- README: rewritten as universal-bot pitch; auth covered as two dimensions
  (MC auth mode + LLM credential)
- AGENTS.md: identity comes from .env, hard-coded references to pepa
  removed; bootstrap step auto-detects whether the server uses an
  AuthMe-style /register-/login plugin
- .env.example: example values replaced with placeholders, MC_AUTH_MODE
  added (offline | microsoft)
- docs/architecture.md: rephrased target as "any Minecraft Java server",
  added open question on cross-server vs per-server state

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 09:52:25 +03:00
mayatnikovandClaude Opus 4.7 3a025c05ad chore: bootstrap pepa-pi-bot scaffold
Initial seed for an autonomous, self-extending Minecraft player powered by
Pi (pi.dev) and Mineflayer.

Includes README, AGENTS.md mandate, .env.example, MIT LICENSE, package.json
with mineflayer + dotenv, and empty skills/ extensions/ prompts/ dirs for
the agent to grow into.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 09:50:17 +03:00