* v0.3.0-rc.1: live skill registry + fast advisor scaffold
Roots out the v0.2.x failure mode: Pi-extracted lessons routinely named
hallucinated skill ids (relocate.surface, choose.safe.surface,
survive.shelter, gather.visible_log, …). All 47 Pi-lessons in the live DB
had applied_count=0 because normalisePreferSkill couldn't find them.
Fix:
1. runtime/skill-registry.js — single source of truth derived from
skills/index.js. Exports listSkillIds, isRegistered, and a
prompt-ready block (skillRegistryPrompt) grouped by namespace.
2. Pi prompts (coach/postmortem, coach/reflect) embed the live registry
with a "USE ONLY THESE, never invent" instruction. Lessons are
filtered at write-time too — anything not in the registry and not a
known mode name gets dropped.
3. coach/advice.js — normalisePreferSkill now returns null for unknown
ids, hardening consult() against any hallucinations that slip
through. Warn-logged for visibility.
Also lays the LLM substrate for the rest of v0.3.0:
- runtime/llm/provider.js — OpenAI-compatible chat client. Configured
via PEPA_FAST_LLM_{BASE_URL,API_KEY,MODEL,TIMEOUT_MS}. Safe no-op
unless API_KEY is set. Supports JSON-mode.
- runtime/coach/fast-advisor.js — tactical advisor tier (scaffold).
Exposes advise() that asks the fast LLM what to do RIGHT NOW when
the reflex is wedged/stuck. Rejects hallucinated skill ids using the
registry. Rate-limited 6/h, 30s cooldown. Not auto-triggered yet —
wired into reflex in rc.3 (awareness layer).
Tests: 279 green (+24 vs rc.3): 5 registry, 9 provider, 10 advisor.
See dev/v0.3.0/PLAN.md for the full iteration design (manifesto needs
ladder, event-driven awareness, skill pre-emption) and STATUS.md for
shipped/pending tracking.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* v0.3.0-rc.2: manifesto / needs ladder L0-L10
Adds an explicit hierarchical needs catalogue that the reflex consults
on every tick. The bot now pursues tangible intermediate goals (food,
wood tools, shelter, stone tools, ...) instead of inheriting whatever
the curriculum thought was "next".
Ladder:
L0 alive HP>5, food>0, not in lava, not panic-near hostile
L1 food ≥6 food items in inventory (or sated + any food)
L2 tools_wood wooden_pickaxe + wooden_axe + wooden_sword
L3 shelter_basic bed placed nearby or in inventory
L4 tools_stone stone-tier triplet
L5 armor_basic any chestplate (pursue=null until craft.leather-*
lands; ladder gracefully skips)
L6 food_security ≥16 food items
L7 tools_iron iron-tier triplet (pursue=gather.stone for now)
L8 armor_iron iron chestplate (pursue=null for now)
L9 village_seed bed + chest in nearby blocks
L10 village_full never detected, falls through to curriculum
Each need has detect(snapshot) → bool and pursue(snapshot) →
{skillId, args} | null. The ladder picks the LOWEST unsatisfied
pursuable need. Needs whose pursue is null get recorded as
blockedNeeds and the walk continues — no stalling on missing skills.
Wired into curriculumReflex: manifesto takes precedence over
curriculum.plan when it has a concrete suggestion. Tests can pass
ctx.disableManifesto=true to exercise the curriculum branch
in isolation (existing reflex tests keep passing this way).
Pi self-reflection prompt now includes
"activeNeed (Maslow ladder L0-L10): L2 tools_wood → gather.logs"
so Pi advises at the right level instead of giving generic guidance.
skillId returned by pursue() is validated against the live registry
(rc.1 plumbing) — manifesto cannot accidentally dispatch a
hallucinated skill name.
Tests: 315 green (was 279 on rc.1, +36 new):
- runtime/manifesto/needs.test.js — 24 tests (per-need detect/pursue,
helper sums)
- runtime/manifesto/state.test.js — 10 tests (ladder walk, hostile
takeover at L0, armor skipping, caching)
- runtime/reflex.test.js — 2 integration tests (manifesto overrides
curriculum plan; well-fed bot pursues tools_stone)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* v0.3.0-rc.3: event-driven awareness + skill pre-emption
Adds a reactive layer on top of the polling reflex. The bot now
notices environmental shocks (forced moves, HP plunges, hostile
spawns) within ~100ms instead of waiting for the next DISPATCH tick,
and the in-flight skill is preempted so the next reflex cycle can
re-plan against the current world state.
This is the rc that wires the "rc.1 plumbing + rc.2 manifesto" into
a feedback loop:
- awareness fires preempt → dispatch aborts
- reflex tick re-evaluates → manifesto walks the ladder
- new dispatch picks the right skill for the new world state
Pieces:
- runtime/awareness/events.js (new) — bot.on listeners:
- move: single-tick Δposition ≥ 5 blocks → forced_move flag + preempt
- health: HP drop ≥ 2 → health_plunge flag + preempt
- entitySpawn: hostile mob within 12 blocks → hostile_added + preempt
- blockUpdate: nearby block change → env_changed flag (no preempt,
throttled 800ms; otherwise gather skills would self-preempt
every dig)
- runtime/skills/index.js — RUNNER_CODES.PREEMPTED + raceWithAbort()
wraps every execute() against ctx.abortSignal. Existing skills get
preemption for free; they don't have to check the signal manually.
- runtime/bot.js:
- dispatchAction creates a fresh AbortController per dispatch and
stores it on reflexCtx.currentAbort
- attachAwareness fires controller.abort() when something disrupts
the active skill; runSkill returns code: "preempted" and the
reflex moves on
- reflexCtx.lastPreempt records the most recent shock
Tests: 332 green (was 315 on rc.2, +17 new):
- runtime/awareness/events.test.js — 12 tests (each event type +
thresholds + throttling + passive-mob filter)
- runtime/skills/contract.test.js — 3 abortSignal tests
(mid-flight, pre-armed, clean signal)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(.env): add PEPA_FAST_LLM_* placeholders for v0.3.0 fast advisor
Empty values keep the fast-advisor tier disabled (safe no-op). Fill
in BASE_URL + API_KEY + MODEL to enable. TimeWeb-style endpoint
example included.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(v0.3.0): rename fast-LLM env vars to TIMEWEB_* (match other projects)
Aligns with the user's other repos (proso) which use TIMEWEB_API_GROK /
TIMEWEB_URL_GROK. Single naming convention across projects avoids the
'which env var was it for this repo' mental tax.
PEPA_FAST_LLM_BASE_URL → TIMEWEB_BASE_URL
PEPA_FAST_LLM_API_KEY → TIMEWEB_API_KEY
PEPA_FAST_LLM_MODEL → TIMEWEB_MODEL
PEPA_FAST_LLM_TIMEOUT_MS → TIMEWEB_TIMEOUT_MS
Provider still works with any OpenAI-compatible endpoint — TimeWeb is
the default but the variable name doesn't lock us in. Tests + docs +
.env / .env.example updated.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(scripts): TimeWeb smoke test + bump default LLM timeout to 20s
scripts/check-timeweb.js — three probes: plain text, JSON mode, full
fast-advisor stack (registry injection + skill validation). Loads .env,
prints {ok, latency, reply preview} for each. Doesn't touch bot state.
Bumped DEFAULT_TIMEOUT_MS 8s → 20s in runtime/llm/provider.js. TimeWeb's
hosted agent endpoint takes 5-15s for the fast-advisor prompt
(registry block + snapshot context), so 8s was producing spurious
timeouts. OpenAI direct returns much faster; env var TIMEWEB_TIMEOUT_MS
overrides if needed.
Smoke verified live (PR #27 branch):
probe 1: 6.3s, plain prompt → "pepa hears you"
probe 2: 5.4s, JSON mode → {"alive":true,"name":"pepa"}
probe 3: 14.9s, advise() → action=switch_skill, skill=recovery.tunnel-out
(correct registered skill, sensible rationale — registry
injection successfully prevents hallucination)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.0): auto-trigger fast-advisor + token usage tracking
Closes the awareness → LLM → action loop that the rc.1/2/3 sequence
left as a followup. When the bot is wedged, looping, or just suffered
a preempt-then-retry, the reflex fires advise() in the background;
when the recommendation lands it overrides the next dispatch.
Async by design: advise() takes 5-15s on TimeWeb's hosted endpoint —
too slow for a synchronous reflex tick. tickAdvisor() is fire-and-
forget, the result lands on ctx.advisorRecommendation, and the *next*
tick reads and consumes it. Recommendations age out after 60s.
Components:
- runtime/coach/advisor-trigger.js — policy + async fire path
- tickAdvisor(ctx, {plannedSkillId}) checks three triggers:
1. wedged > 60s (no significant move)
2. last 4+ dispatches are the same skill AND it's planned again
3. preempt within last 30s + same skill being retried
- 90s trigger cooldown, single-in-flight guard
- consumeFreshRecommendation(ctx) reads/clears the cache
- runtime/reflex.js — curriculumReflex calls tickAdvisor() every tick
and consumes a fresh recommendation BEFORE dispatching. ctx flag
disableAdvisor=true for tests.
- runtime/bot.js — dispatchAction maintains a rolling 8-slot
reflexCtx.recentSkillIds for the loop-detection trigger.
Token usage:
- runtime/llm/provider.js — normaliseUsage() reads OpenAI/TimeWeb-
style {prompt_tokens, completion_tokens, total_tokens} from the
response. Returned on every complete() result and logged at info
level as "in=Nt/out=Mt".
- runtime/coach/fast-advisor.js — getUsageSnapshot() aggregates
total tokens across all calls in the session.
Measured on live TimeWeb endpoint (gpt-5.4-mini agent):
per call: ~705 input + 45 output = ~750 tokens
rate limit: 6 calls/hour
worst case at full budget: ~108K tokens/day
estimated cost (OpenAI gpt-5-mini reference price): ~$0.60/month
Well within any reasonable budget — model can run hot 24/7.
Smoke verified: scripts/check-timeweb.js probe 4 produces
trigger fired: true (wedged_90s)
recommendation: recovery.tunnel-out
rationale: "Stuck wedged for 90s; exploration is failing."
latency: 5302ms
Tests: 345 green (was 332, +13 advisor-trigger).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.0): paradigm shift — TimeWeb-only LLM + persistent advisor trail + improvement queue
This is the rc.4 batch the user requested:
1. Emergency triggers (low HP + close hostile, lava-under-foot)
bypass the long cooldown so the LLM is consulted BEFORE the bot
dies, not after.
2. Active manifesto need is now included in the advisor user prompt
— the LLM picks suggestions that satisfy the bot's current
concrete need (L2 tools_wood → "gather logs nearby" not
"explore further").
3. Every advisor recommendation is persisted to SQLite
(advisor_recommendations table) with full token usage. The
reflex marks 'applied=1' when it dispatches and updates
outcome_ok/code when the dispatch completes. Ground truth for
"is the LLM actually helping" lives in the DB, not in logs.
4. Pi CLI is OUT of every background loop. coach/postmortem and
coach/reflect now go through the same TimeWeb endpoint
fast-advisor uses, via the shared coach/llm-call.js helper.
Pi is reserved for manual operator commands.
5. The LLM (postmortem, reflect, advisor) can flag "structural
gaps" — missing skills/features the operator should implement.
These land in the new improvement_requests table. Dedup by
title bumps `votes` instead of inserting duplicates so the
queue doesn't bloat. Operator views via
`node scripts/list-improvements.js`.
6. A deterministic trigger-tuner runs hourly: reads 24h of
recommendation stats, flags triggers whose success rate is
below 25% (sample ≥ 5) or whose prompts are expensive (>1000
input tokens) with mediocre payoff. Improvements get
source="tuner", category="tuning". No LLM call.
New files:
runtime/coach/llm-call.js — askAnalytical() helper
runtime/coach/trigger-tuner.js — stats → improvements
runtime/coach/trigger-tuner.test.js
scripts/list-improvements.js — operator CLI
Schema additions:
advisor_recommendations: id, ts, trigger_reason, planned_skill,
recommended_skill, action, rationale, active_need, tokens_in,
tokens_out, latency_ms, applied, outcome_ok, outcome_code, outcome_at
improvement_requests: id, ts, source, category, title, description,
context, priority, status, duplicate_of, votes, implemented_at, notes
Renamed env-var consumers:
Pi-coach drainOnce({ askPi }) → drainOnce({ askAnalyticalFn? })
Pi-reflect runOnce({ askPi }) → runOnce({ askAnalyticalFn? })
bot.js attachCoach/attachReflect no longer pass askPi
attachTuner() added to bot.js spawn handler
lessons.source 'pi-coach' → 'timeweb-coach'
lessons.source 'pi-reflect' → 'timeweb-reflect'
Token cost measured live:
~705 input + 45 output = ~750 total per advisor call
worst case @ 6 calls/hour rate cap = ~108K tokens/day
OpenAI gpt-5-mini reference price: ~$0.60/month
Operator usage:
node scripts/list-improvements.js # open queue
node scripts/list-improvements.js --stats # advisor performance
node scripts/list-improvements.js --done 17 "shipped in 0.3.1"
node scripts/list-improvements.js --reject 18 "duplicate"
Tests: 360 green (was 332, +28 new).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
pepa-pi-bot
A universal, autonomous, self-learning Minecraft player. Built on Mineflayer with a hybrid runtime: a fast script-driven reflex loop for the everyday, headless Pi escalation for the hard bits, a SQLite-backed knowledge base the bot reads and writes as it plays, and a git-as-evolution-substrate self-improvement loop where Pi writes new skills, runs
npm test, and opens a PR for operator review. Works against any Minecraft Java server — vanilla, Paper, Spigot, Fabric, Forge, online-mode or cracked, modded or vanilla.
The bot is not a finished application. It is a seed. v0.2.0 added the learning substrate: every death is recorded with full context and asynchronously analysed by Pi into generalised lessons; the dispatcher consults those lessons before each action and routes around mistakes; the bot narrates its life in Russian MC chat to feel like a player, not a script. Earlier substrate (v0.1.x): Mineflayer body, reflex chain, persistent world-journal / scenario-memory, Voyager-style critic + Mindcraft-style modes/skill-library, scoped auto-patch loop.
Related work: conceptually close to Voyager (NVIDIA, GPT-4) and Mindcraft (multi-agent LLM framework). The differentiators:
- The growing skill library lives as versioned source code on
main(not as JSON in RAM). Every Pi-written skill goes throughgit checkout -b → npm test → PR for operator review, making the loop auditable and rollback-safe. - Learned lessons live in a per-server SQLite knowledge base (
state/<host>/knowledge.db), with recipes, mob intel, block intel, lessons, deaths, post-mortems, points-of-interest, cached wiki pages, and a chat log. The dispatcher reads from it before acting.
The name pepa-pi-bot is just the project's name (pepa from the original test server, pi from the original runtime). The bot itself is server-agnostic.
Runtime modes
Two ways to run the bot. The hybrid runtime is the default — Pi-only is a fallback for experiments.
| Mode | Entry | When to use |
|---|---|---|
| Hybrid runtime (recommended) | npm run bot + npm run tui |
Day-to-day. Script-driven reflex tick + Ink TUI dashboard + Pi/Codex called only on demand. Fast, cheap, observable. |
| Pi-only (fallback) | npm run agent |
When you want every decision to go through an LLM (rare, but useful for experiments and code-writing sessions). |
See docs/runtime.md for the full hybrid runtime guide, IPC protocol, TUI hotkeys, and the self-improvement loop.
Concept
Most Minecraft AI bots ship as monolithic projects: hard-coded actions, fixed prompts, a single LLM provider, sometimes a single target server. This repo flips that around with a layered runtime.
┌──────────────────────────────────────────────────────────┐
│ TUI (Ink) — operator dashboard, attaches via Unix sock │
│ status / live log / MC chat / hotkeys / ask-Pi │
└──────────────────────┬───────────────────────────────────┘
│ newline-JSON
▼
┌──────────────────────────────────────────────────────────┐
│ runtime/bot.js — long-running Node daemon │
│ ├── Mineflayer client (MC TCP, AuthMe, chat, events) │
│ ├── Modes chain (self_preservation > hunger > shelter) │
│ │ priority interrupts before curriculum dispatch │
│ ├── Reflex loop (defend > eat > sleep > curriculum) │
│ │ consults coach/advice before every dispatch │
│ ├── perception.js — numeric-id findBlocks (VB-safe) │
│ ├── knowledge/ — SQLite store (recipes, mob/block intel,│
│ │ lessons, deaths, post-mortems, POI, chat log) │
│ ├── coach/postmortem — death → DB → Pi → lessons │
│ ├── coach/advice — recall lessons → override / avoid │
│ ├── persona/chatter — Russian narration in MC chat │
│ ├── world-journal + scenario-memory (persistent JSONL) │
│ ├── stuck-incident → critic (Pi) → proposal │
│ ├── auto-improve → auto-patch → npm test → PR (review) │
│ └── pi-bridge — spawn `pi -p` only on demand │
└──────────────────────┬───────────────────────────────────┘
│ TCP 25565 (any host/port)
▼
ANY Minecraft Java server (configured in .env)
The reflex loop is the brain stem. Pi is the cortex — called only when the reflex loop is genuinely stuck, or when the operator asks for help via the TUI. Mineflayer is the body. The skills, reflexes, and supervision loop are meant to grow over time — both by hand and by the bot itself proposing patches.
Prerequisites
| Tool | Why | How to get it |
|---|---|---|
Pi ≥ 0.75 |
The agent runtime. Reads AGENTS.md, loads skills, calls the LLM. |
curl -fsSL https://pi.dev/install.sh | sh |
Node.js ≥ 20 |
Required by Pi and by Mineflayer. | brew install node / nvm install 20 |
| An LLM credential | One of: OpenAI / Anthropic / Google API key, or an OAuth-authenticated subscription (/login inside Pi). ChatGPT Pro and Claude Max work via OAuth on supported providers. |
See Authentication |
| Access to some Minecraft server | The bot joins as a real player. Cracked or premium, online-mode or offline, doesn't matter — configure it in .env. |
— |
| Network access to that server | Direct TCP to host:port. |
— |
The bot does not need its own Minecraft client install, server admin access, RCON, or any server-side plugin. It joins as a vanilla player over the standard protocol.
Quickstart
# 1. Clone + configure
git clone git@github.com:xmatic-squad/pepa-pi-bot.git
cd pepa-pi-bot
cp .env.example .env
$EDITOR .env # set MC_HOST, MC_USERNAME, auth mode, AuthMe password, etc.
# 2. Install Node deps
npm install
# 3. (Optional) Authenticate Pi for the escalation hotkey
pi /login # OAuth flow — ChatGPT Pro / Claude Max
# or export OPENAI_API_KEY / ANTHROPIC_API_KEY
# 4. Run the bot — two terminals
# Terminal 1: the daemon (logs in stdout, persists state under state/<host>/)
npm run bot
# Terminal 2: the dashboard (Ink TUI). Hotkeys: p/s/r/c/a/k/v/!/y/q.
npm run tui
The TUI auto-reconnects to the bot if you restart it. Press q to leave the TUI; the bot keeps running.
Want the LLM-driven, single-process flavour?
npm run agentlaunches the original Pi runtime instead. Seedocs/runtime.mdfor the trade-offs.
Sending chat or asking Pi from the TUI
- Press
cin the TUI to enter chat mode — type, Enter sends into MC chat (rate-limited per.env). - Press
ato enter ask-Pi mode — type a prompt, Enter spawnspi -p "<prompt>". Output streams into the Pi panel without leaving the TUI.
TUI hotkeys cheatsheet
| Key | Effect |
|---|---|
p |
Pause / resume the reflex loop (MC connection stays). |
s |
Stop the bot (graceful disconnect + cleanup). |
r |
Force a fresh status snapshot. |
c |
Send a chat message into MC. |
a |
Ask Pi (one-shot subprocess). |
k |
Run one registered skill with optional JSON args. |
v |
Capture a viewer screenshot for debugging. |
! |
Force a critic-backed incident/proposal. |
y |
Open the latest pending proposal (badge appears in status bar). In the panel: y approve, n/Esc close. |
q |
Quit the TUI — bot keeps running. |
Self-improvement loop (short version)
The bot proposes its own patches when something repeatedly fails. As of v0.2.0 the loop ends in a PR for operator review, not a direct merge to main:
- Reflex action fails 5× with the same label, or
noProgressReasonstays stuck → markdown proposal lands understate/<host>/proposals/. runtime/auto-improve.jswatcher (2 s poll, 10 s debounce) picks it up and spawnsscripts/auto-patch.jsdetached.auto-patch.js: refuses on dirty tree, moves proposalpending → approved/, creates branchauto/<slug>off main, runspi -p "<patch prompt>"with 10-min timeout.- If Pi committed and only touched the proposal's
editScope, ANDnpm run lint-patch+npm testboth pass →git push origin auto/<slug>+gh pr createagainstmain. - Operator reviews the PR on GitHub and merges. Branch protection on
mainrequires approval — auto-patch cannot merge itself. - Supervisor watches
runtime/**/*.jsand hot-restarts the child on file change once the merged commit is pulled.
Manual paths still work:
- TUI hotkey
yopens the latest pending proposal for inspection. npm run propose:apply <filename>— attended version: same flow but leaves the branch local without opening a PR.PEPA_AUTO_PATCH_MERGE=cherry-pickenv reverts to the legacy direct-to-main behaviour (not recommended; bypasses review).
See docs/runtime.md for the full lifecycle. Note (2026-05-25): MC chat is now dialog-only — operator commands (come, pause, stop, …) are no longer dispatched from chat; use the TUI for local control. See plans/autonomous-survival-bot-prd.md for the survival-bot pivot.
Authentication
Two dimensions:
1. Minecraft auth. Configured in .env via MC_AUTH_MODE:
offline— cracked servers. Any nickname works. No external auth call.microsoft— premium / online-mode servers. Mineflayer handles the device-code flow on first connect and caches the token in~/.minecraft-auth/.
2. LLM auth. Pi supports 15+ providers and two credential modes:
- OAuth subscription login —
pithen/logininside the TUI. Suitable for ChatGPT Plus/Pro, Claude Max, and other subscriptions that ship an OAuth flow. No metered API billing. - API key environment variables —
OPENAI_API_KEY,ANTHROPIC_API_KEY,GOOGLE_API_KEY, etc. Metered, but no UI prompt.
You can mix providers via --provider openai --model gpt-5 at launch — cheaper models for idle ticks, smarter ones for hard decisions.
How the agent extends itself
Pi has first-class support for three growth surfaces:
skills/— Markdown-defined capabilities Pi can invoke. The agent canWritenew ones at runtime when it discovers a missing capability.extensions/— TypeScript modules registering new tools, commands, or UI tweaks. Installed project-locally viapi install -l npm:<pkg>/pi install -l git:<url>, or written in-tree.prompts/— Reusable prompt templates. Useful for cron-driven tick prompts ("what should I do next minute?").
The opening AGENTS.md instructs the agent to start by writing a mineflayer-bridge extension that can:
- connect to the configured MC server (any host/port/version)
- handle the configured auth mode (offline or microsoft)
- if a login plugin like AuthMe is present, perform
/registerand/loginfrom a password supplied in.env - emit world events back into the agent loop
- expose
chat / move / dig / place / equip / attackas Pi tools
Everything beyond that — farming, exploration, base-building, player interaction, server-specific quirks — should emerge from the agent itself.
Everything in the repo
A hard rule, mirrored in AGENTS.md: every artefact the agent produces lives in this repo, never in the user's ~/.pi/ directory. That includes skills, extensions, prompt templates, project Pi settings (.pi/settings.json), and per-server state (state/<MC_HOST>/).
The point is reproducibility and community growth: a fresh git clone should bring along every skill any contributor has written. Pi's own built-in skills (skill-creator, agent-browser, etc.) stay user-global — the agent is allowed to use them, but anything it authors lands under ./skills/ or ./extensions/ here.
See CONTRIBUTING.md for the skill format and how to propose changes.
Project layout
pepa-pi-bot/
├── README.md ← you are here
├── AGENTS.md ← seed prompt, loaded by the Pi-only runtime
├── .env.example ← all required env vars, no secrets
├── package.json ← node deps + scripts (`bot`, `tui`, `agent`)
├── runtime/ ← hybrid runtime (script reflex + IPC server)
│ ├── bot.js long-running Mineflayer daemon
│ ├── reflex.js priority-ordered behaviours, no LLM
│ ├── perceive.js snapshot builder
│ ├── ipc-server.js Unix-socket server
│ ├── ipc-protocol.js shared IPC contract
│ └── pi-bridge.js spawn `pi -p` on demand
├── tui/ ← Ink TUI dashboard
│ ├── tui.tsx
│ └── ipc-client.js
├── skills/ ← markdown skills (grown by bot or operator)
├── extensions/ ← Pi extensions (mindcraft-skills, mineflayer-bridge)
├── prompts/ ← reusable prompt templates
└── docs/
├── runtime.md hybrid runtime guide (start here)
├── architecture.md longer-form design notes
├── memory-model.md per-server state layout
├── roadmap.md phased plan
└── …
Safety boundaries
Server-agnostic but with hard defaults the agent must respect on any server it joins:
- Never request OP / admin rights in chat.
- Never break or modify other players' builds without explicit human request.
- Never spam chat — built-in rate limit (
CHAT_RATE_LIMIT_PER_MINin.env). - Never leak secrets from
.env(auth passwords, API keys) into chat, world signs, books, commits, or web fetches. - No destructive bash in the repo (
rm -rf, force pushes) without operator confirmation. - Stop and wait if kicked or banned — do not auto-reconnect indefinitely.
These are mirrored in AGENTS.md and re-stated at the top of any system prompt that overrides it.
Status
🌳 Phase 0 — Body done. Bridge online, AuthMe handled, hello sent. See skills/server-onboarding.md.
🌳 Phase 1 — Presence implemented: bridge stays online with bounded reconnects, rolling chat buffer, status/recent-chat tools, sparing replies.
🌿 Phase 2 — Locomotion/build rails in progress: mineflayer-pathfinder is wired with guarded mc_goto, plus mc_build_pyramid_5x5. Phase 5 self-extension and Phase 6 escalation logging are implemented.
🌱 Phase 3 — Goal-driven autonomy seeded: docs/memory-model.md defines shared-knowledge vs personal-memory; per-server goal.md / plan.md / current-task.json / diary/ shape autonomous behaviour.
🌿 Survival-bot pivot (2026-05-25) — the bot is becoming a self-sufficient survival resident of the configured server. MC chat is dialog-only; operator/player chat commands are recorded but not dispatched (TUI is the only local control plane). The hybrid runtime now has enriched perception, priority modes, a skill-driven curriculum, food acquisition, bed/sleep, base/chest/shelter/farm skills, persistent skill metrics, scenario memory, and a scoped auto-patch loop with npm test smoke gating. Full plan: plans/autonomous-survival-bot-prd.md (local-only, gitignored).
🌿 v0.2.0 — self-learning agent (2026-05-27). SQLite-backed knowledge base at state/<host>/knowledge.db with seeded recipes / mob intel / block intel / 12 starter lessons. Every death is captured with full context into a deaths table; a periodic Pi-coach loop (≤3 calls/hour, 12-min cooldown) batches unanalysed deaths and extracts generalised lessons. A 30-min self-reflection loop asks Pi "are you in a loop or making progress?" and records the verdict + lessons to state/<host>/reflections/. The reflex chain consults coach/advice before every dispatch (including fallback paths) — high-confidence lessons can override (attack creeper → survive.flee) or back-off. Bot narrates major events in Russian MC chat (≤8 lines/hour, ≥75 s gap). New escape skill survive.pillar-up for pit terrain; wedged-emergency reflex fires it after 60s of no horizontal progress. Auto-patch now opens a PR for operator review instead of cherry-picking onto main — main is protected, the operator is the only approver. Currently shipped: rc.1 (substrate) + rc.2 (P0 hardening) + rc.3 (escape + advice-in-fallback). Current state, next steps, and known issues: dev/v0.2.0/STATUS.md. Original design: docs/v0.2.0-self-learning.md.
Full plan: docs/roadmap.md. Memory layout: docs/memory-model.md. Day-to-day judgement: "Operating principles" in AGENTS.md.