* docs(v0.3.1): PRD — LLM prompt cost optimization
Design-only commit; no runtime changes. Spec for the next patch iteration.
Goal: cut per-advise() input tokens from ~800 to ≤300, preserving the
LLM's ability to produce valid registered skill ids and useful rationale.
Five proposed changes ranked by impact:
P1 Compact registry format (saves ~350t/call) — group by namespace,
comma-list ids, drop human titles. Default mode for advisor;
verbose mode kept for postmortem/reflect.
P2 Need-scoped registry (~50t additional) — show LLM only skills
relevant to the active Maslow need + always-available safety
skills (survive.flee, pillar-up, recovery.tunnel-out, explore.*).
P3 Snapshot pruning (~50t) — drop weather/experience/dimension/biome/
players from the user prompt; the LLM doesn't consult them.
P4 Prompt caching probe — check if TimeWeb passes through
prompt_tokens_details.cached_tokens. If yes, restructure prefix
to maximize cache hits (cached input is ~10x cheaper at OpenAI).
P5 Per-trigger cost telemetry in scripts/list-improvements.js --stats:
avg_in / avg_out / cost_₽ / share% per trigger_reason, using
TIMEWEB_PRICE_IN_RUB_PER_M and TIMEWEB_PRICE_OUT_RUB_PER_M env.
Trigger: TimeWeb admin panel after first day of v0.3.0 live showed
34K tokens / day at low activity. At cap budget that projects to
~480₽/month (101₽/M in, 608₽/M out for gpt-5.4-mini). Manageable
but the savings are mostly free — repeated infra tokens, not signal.
All changes are additive; runtime behaviour stays the same. If the
LLM produces worse advice with the compact registry, flip back via a
single constant in fast-advisor.js.
Acceptance: re-run scripts/check-timeweb.js probe 3 — expect
tokens_in ≤ 300 (was ~800). Live for 1h, check --stats: avg_in ≤ 300
per trigger group. Existing 360 tests still green.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.1): storyline — canonical Minecraft survival quest
The bot has been stuck in a loop for two days:
acquire-food (fail: no nearby food) → explore.far → pillar-up (fail) → repeat
Diagnosis: manifesto + LLM advisor both correctly identify "you need
food" but neither expresses *what concretely to do next*. Manifesto is
a priority ladder (need-detection), not a narrative arc.
This commit adds the missing narrative layer — an ordered list of
operational steps that mirror the vanilla Minecraft survival path:
1. orient_self — Понять где я
2. first_wood — Собрать 8 поленьев
3. crafting_basics — Сделать верстак и палки
4. first_tools — Деревянные орудия
5. first_food — Найти первую еду
6. shelter_minimal — Простой шелтер с кроватью
7. stone_tier — Каменные орудия
8. food_security — Запас еды на 16+
9. iron_age — Железо и печь
10. settle_base — Постоянная база
11. village_grow — Развивать деревню (ongoing)
Each step has:
- completed(snapshot) → bool — detects achievement from snapshot
- suggestSkill(snapshot) → { skillId, args? } — concrete next dispatch
- emergencyPause(snapshot) → bool — defers to manifesto L0 alive
emergencies (low HP near hostile, lava under foot, food = 0)
- narration_ru — chat-friendly Russian one-liner spoken on entry
Components:
- runtime/goal/storyline.js — 11-step canonical quest catalogue
- runtime/goal/state.js — pickCurrentStep(snapshot) walks the list,
returns first non-completed step + its suggestion. 3s cache.
Validates suggestSkill's skillId against the live registry.
- runtime/reflex.js — curriculumReflex dispatch priority is now:
1. manifesto (L0 alive emergencies always win)
2. storyline (concrete operational subgoal)
3. curriculum plan (legacy fallback)
Tests pass ctx.disableStoryline=true for isolation.
- runtime/bot.js — snapshot.storyStep populated each tick so
chatter/advisor/reflect observers see the same view.
- runtime/coach/fast-advisor.js — buildUserPrompt now embeds the
current step + its suggested skill, so LLM advice is anchored
("step 5 first_food, storyline wants survive.acquire-food, but
recent dispatches show it's failing — try explore.far + scout").
- runtime/coach/advisor-trigger.js — forwards ctx.storyStep into
advise() and logs step id at trigger time.
- runtime/coach/reflect.js — reflection prompt includes storyline
progress so 30-min self-assessment is anchored.
- runtime/persona/chatter.js — narrates step.narration_ru on
transition. Rate-limited via existing maybeNarrateRaw().
New operator CLI:
- scripts/show-story.js — fetches the live snapshot via IPC sock and
prints step progress with ✓/→/ markers, current skill, inventory.
Falls back to --plain catalogue view when bot offline.
Token cost impact: ~+30 input tokens per advise() call (one extra
line in user prompt). Trivial vs the value of grounding LLM advice
in a concrete narrative.
Operator usage:
node scripts/show-story.js # live progress + which step + why
node scripts/show-story.js --plain # static catalogue of all 11 steps
Tests: 376 green (was 360, +16 storyline tests).
Also in this branch (already committed): dev/v0.3.1/PRD.md —
LLM prompt cost optimization design doc.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(v0.3.1): storyline beats manifesto L1+ (only L0 alive emergencies override)
Found in live logs after the previous commit deployed:
storyline: step 1/11: orient_self → explore.wander
advisor-trigger: firing because wedged (planned=survive.acquire-food, ...)
Manifesto was still picking survive.acquire-food (L1 food) over the
storyline's orient_self → explore.wander. That's the wrong precedence —
storyline expresses a *concrete operational subgoal* and L1+ manifesto
needs are just "you'd benefit from food" priorities, not emergencies.
New dispatch precedence in curriculumReflex:
1. manifesto L0 (alive emergencies: lava, low-HP+hostile, food=0)
2. storyline (concrete narrative subgoal — beats L1+ manifesto)
3. manifesto L1+ (fallback when storyline has no concrete suggestion)
4. curriculum plan (legacy fallback)
This way the bot starts following the narrative arc even while
manifesto's L1 food is technically unsatisfied — orient_self runs to
completion before pursuing food explicitly. Storyline already handles
food as step 5 (first_food), so we're not skipping it.
Tests: 378 green (+2 priority-ordering tests):
- L0 manifesto emergency: upstream reflex (defend/modes) catches before
curriculum dispatch
- storyline beats manifesto when both have suggestions: well-fed bot
with logs → craft.planks (storyline crafting_basics), not gather.logs
(manifesto L2)
- updated "manifesto fallback" test to require disableStoryline=true
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(tui): fullscreen monitor-only TUI (opencode-style)
Replaces the old tui/tui.tsx hotkey-heavy dashboard with a read-only
observability screen. Operator actions live in scripts/* now —
TUI is for watching, not driving.
Layout (top to bottom, all auto-resizing to terminal):
1. Header — MC/IPC status, pos, HP, food, day/night, hostiles
2. Storyline — current step + 11-step quest map (✓/→/○)
3. Activity — last N skill dispatches (colour by outcome)
4. MC Chat — last N chat lines (cyan for bot, yellow for players)
5. Advisor — last N LLM recommendations (trigger + outcome + tokens)
6. Improvements — open requests from knowledge.improvement_requests
7. Footer — 24h token usage + cost in ₽ + q-to-quit
Data sources:
- IPC sock: snapshot frames, log frames, chat frames (push)
- SQLite knowledge.db: advisor_recommendations + improvement_requests
polled every 5s (pull)
Token cost displayed live using TIMEWEB_PRICE_IN_RUB_PER_M /
TIMEWEB_PRICE_OUT_RUB_PER_M env vars (defaults: 101 / 608 for
gpt-5.4-mini).
Switches:
- npm run tui → new monitor (this file)
- npm run tui:legacy → old action-driven tui/tui.tsx (kept for now)
Implementation notes:
- Uses ink + alternate-screen-buffer ANSI for proper "opencode-feel"
fullscreen behaviour; restores prior terminal contents on quit.
- Skips alt-screen and useInput when stdin/stdout isn't a TTY
(smoke tests, piped output) — both gracefully degrade.
- Stable React keys via per-event uid counter, avoids reconciler
duplicate-key warnings as logs/chat/dispatches stream in.
- Resize handled via 1s stdout-dimension poll, NOT direct
'resize' listener (which conflicts with ink's own listener and
triggers MaxListenersExceededWarning).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* ui(tui): compact 4-section monitor (was 6) — fits 1080p without zoom
Operator reported the TUI overflowed the screen unless terminal was
zoomed way out. The 11-step storyline list alone was eating ~13
rows, and each advisor/improvement entry took 2-3 rows. Now:
- Header + storyline collapsed into one panel (2 lines):
line 1: pepa · ●MC ●IPC · 1m50s · pepa_bot · (697,61,702) · HP 20 · food 5 · ☀ · ⚔60(creeper@58b)
line 2: story ▓▒░░░░░░░░░ 1/11 orient_self · Понять где я → explore.wander
The 11-step ladder is now a unicode progress bar (▓ done, ▒ current,
░ pending) — same info, fits in one row.
- Advisor entries: one line each instead of two.
✓ wedged_60s → survive.flee 802t 1900ms
(outcome mark / trigger / target skill / tokens / latency)
- Improvements entries: one line each instead of two.
#1 P2 ×3 Add craft.iron-pickaxe skill
Description dropped from the row — use `node scripts/list-improvements.js`
for full text.
- Sections: 4 (was 6).
[header+story] · [activity | chat] · [advisor | improvements] · [footer]
Tested on a typical 1080p terminal — fits comfortably without zoom.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(v0.3.1): real survival patterns — biome-aware scout, wedge-relocate, escape-pit-safe
Operator reported the bot wandered the same 50×50 patch for 2 hours
without making any progress toward food. Diagnosis showed three root
causes; this commit addresses all five open improvement_requests
the LLM (postmortem + tuner) flagged automatically.
Research basis (`Voyager`, `Plan4MC`, `GITM`, `Mindcraft`):
- Coverage / commit-to-cardinal exploration when local scan fails
- Biome-aware strategy switching using a static affordance table
- Wedge detector above the skill layer that triggers RELOCATE not
RETRY (per-skill stuck checks reset on re-entry — useless)
- Time-in-region bbox heuristic + need-duration AND skill-cycle gate
Concrete changes:
1. `runtime/goal/storyline.js`
- orient_self.completed: added timeout fallback (HP=full + session
>120s → done) so barren biomes don't block the bot on step 1.
Closes improvement #2 'Нет навыка оценки когда сменить район'.
- first_food.suggestSkill: now picks survive.scout-food (new) when
no passive mob is nearby; falls back to survive.acquire-food only
when something is in immediate range.
2. `runtime/biome-affordances.js` (new)
- Static table: 40+ biomes → {has_passive_mobs, has_trees,
has_water, has_crops, livable}.
- Unknown biomes return optimistic defaults to avoid regressions.
- Closes improvement #1 'Нет навыка целевого поиска еды по биому'.
3. `runtime/skills/scout-food.js` (new — survive.scout-food)
- Tiered strategy: biome check → scan 32 → scan 64 → commit a
cardinal for 200 blocks rescanning every 16. On cardinal
exhaustion, returns code:"exhausted" so the curriculum can
escalate to village.relocate.
- In barren biomes (desert/ocean/snowy_plains) the scan is
SKIPPED — bot walks straight toward the nearest neighbour
biome that affords passive mobs (8-direction biome probe at
radius 64).
4. `runtime/awareness/wedge-detector.js` (new)
- Rolling 10-min position bbox tracker. observe() called every
tick; isWedged() returns true when bbox<50 AND active need
unmet >5min AND skill cycles ≥3.
- markRelocationStarted() suppresses further wedge firings
until the bot has displaced ≥200b — prevents stack overflow
of relocate calls.
- Lives ABOVE the skill layer (in runtime/reflex.js), because
any per-skill stuck check resets on re-entry.
5. `runtime/skills/relocate.js` (new — village.relocate)
- 300-block walk in least-recently-used cardinal (per-incident
memory in ctx.recentRelocations).
- Re-paths every 32 blocks, soft-tolerates pathfinder failures
(3 consecutive throws → exit with code:"stuck_in_place").
- Closes improvement #2 + #4 ('low success rate trigger').
6. `runtime/skills/escape-pit-safe.js` (new — recovery.escape-pit-safe)
- Surveys 4 cardinals AND ceiling height before committing.
Picks the direction with most open blocks (≥3, no lava).
Falls through to pillar-up only if ceiling clear ≥4b. Returns
code:"no_strategy" if both blocked so curriculum can escalate
to relocate.
- Closes improvement #3 'Нет навыка для безопасного выхода'.
7. `runtime/reflex.js`
- Wedge detector wired before manifesto/storyline. If wedge.wedged
is true, dispatches village.relocate directly and returns —
bypasses every other branch.
- ctx.disableWedge flag for tests.
8. `runtime/coach/advisor-trigger.js`
- LLM provider outage backoff: 3 consecutive http_400 / timeout /
network_error → suppress advisor for 10 min. Today's TimeWeb
gpt-5.4-mini was 400'ing for an hour straight; we were spending
trigger budget on dead calls. Closes improvement implicit gap
in #5.
9. `runtime/bot.js`
- Tracks botSpawnedAt; snapshot._sessionMs exposed for storyline
orient_self timeout fallback.
Tests: 396 green (was 378, +18):
- runtime/biome-affordances.test.js — 8 tests
- runtime/awareness/wedge-detector.test.js — 9 tests
- runtime/goal/storyline.test.js — 1 new test (orient_self timeout)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(coach/trigger-tuner): crash after ~1h — runOnce is sync, not a Promise
The live bot died overnight with:
TypeError: runOnce(...).catch is not a function
at trigger-tuner.js:42 → [supervisor] child exited code=1
attach() wrapped the timer body as `runOnce().catch(...)` but
runOnce() returns a plain {ok, flagged, ...} object (pure SQL, no
await). The first tuner tick (60min after spawn) threw → killed the
whole bot process. Never surfaced before because the bot rarely ran
uninterrupted for a full hour during development.
Fix: guard the synchronous call with try/catch, matching how
persona/chatter.js already does its sync tick. (postmortem.drainOnce
and reflect.runOnce ARE async, so their .catch is correct — audited.)
Regression test added: captures the setInterval callback and invokes
it synchronously, asserting it does not throw.
Tests: 397 green.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(v0.3.1): mechanical food/stuck fixes — bot reaches the chicken now
The wedge wasn't only in the manifesto layer; several mechanical bugs
kept the bot in a dead random-walk:
- storyline / manifesto / curriculum: "local food" now means an edible
passive mob within <=32 blocks. A distant chicken or a cod no longer
fools the bot into dispatching acquire-food (which then fails on
no_path). Long-range food goes through scout-food instead.
- scout-food: partial approach to a target now counts as progress
(approached_target, e.g. moved:14); a blocked heading is NOT counted
as movement; added blind/tunnel fallback so it doesn't die when the
pathfinder can't route cleanly.
- acquire-food: on no_path it now also tries a blind/tunnel approach to
the animal; no_drop routes back into food scouting instead of giving
up.
- explore.far / relocate / flee: fewer false "done" results (micro-steps
no longer counted as success), more genuine escapes from stuck.
- scripts/show-story.js: live IPC now actually renders the current
storyline step.
Verification: scripts/lint-patch.js clean; npm test 404/404 green; bot
relaunched in tmux `pepa`. Live logs show real progress — bot switched
to survive.scout-food, approached the chicken (approached_target
moved:14), then reached survive.acquire-food: hunting chicken. Food
isn't fully closed yet but the remaining issue is concrete pickup/drop,
not dead random-walk.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(v0.3.1): sated bot stops chasing food + perf-leak + fuzzy improvement dedup
Third day of "bot just walks back and forth burning tokens". Root
causes were mechanical, not the manifesto:
1. SATED BOT CHASING FOOD (the big one)
Bot had food=17 (nearly full) but storyline first_food + manifesto
L1 required 2+ food ITEMS in inventory, so it looped scout-food /
acquire-food for hours instead of working. Now both treat a hunger
bar >= 14 (SATED_FOOD) as satisfied even with empty food inventory —
a full bot chops wood / makes tools and grabs food opportunistically,
only hard-pursuing food when actually hungry (< 14).
manifesto/needs.js foodDetect + goal/storyline.js first_food.completed.
2. perf_hooks MEMORY LEAK (overnight OOM suspect)
"MaxPerformanceEntryBufferExceededWarning: 1,000,001 measure entries".
mineflayer/pathfinder emit perf marks we never consume. Added a
60s reaper in bot.js (performance.clearMeasures/clearMarks). unref'd.
3. IMPROVEMENT QUEUE SELF-DUPLICATING
The LLM re-filed closed gaps with reworded titles (#5/#8/#9 were
dupes of implemented #1/#2/#3). Exact-title dedup missed them.
Replaced with token-set fuzzy match (isDuplicateTitle): jaccard>=0.75
OR >=3 shared meaningful tokens with jaccard>=0.5. Also: a re-filed
gap that's already implemented/rejected is NOT resurrected as a new
open row. Cleared all 5 open requests (now genuinely implemented).
Also confirmed (no change needed):
- canDig=true is a DELIBERATE codebase-wide choice ("without it the bot
gets permanently stuck", actions.js). The stale memory recommending
canDig=false is updated. ViaBackwards dig works partially (dug:1
moved:1.8 observed); false would trap the bot in every pit.
- scout-food already has blind/tunnel fallback + 12s step timeout
(operator's earlier edits) so trapped-pathfinder degrades instead of
hanging 30s.
Tests: 407 green (was 404). Updated needs/state/storyline tests for the
SATED_FOOD threshold; added fuzzy-dedup + tokenize/jaccard tests.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
pepa-pi-bot
A universal, autonomous, self-learning Minecraft player. Built on Mineflayer with a hybrid runtime: a fast script-driven reflex loop for the everyday, headless Pi escalation for the hard bits, a SQLite-backed knowledge base the bot reads and writes as it plays, and a git-as-evolution-substrate self-improvement loop where Pi writes new skills, runs
npm test, and opens a PR for operator review. Works against any Minecraft Java server — vanilla, Paper, Spigot, Fabric, Forge, online-mode or cracked, modded or vanilla.
The bot is not a finished application. It is a seed. v0.2.0 added the learning substrate: every death is recorded with full context and asynchronously analysed by Pi into generalised lessons; the dispatcher consults those lessons before each action and routes around mistakes; the bot narrates its life in Russian MC chat to feel like a player, not a script. Earlier substrate (v0.1.x): Mineflayer body, reflex chain, persistent world-journal / scenario-memory, Voyager-style critic + Mindcraft-style modes/skill-library, scoped auto-patch loop.
Related work: conceptually close to Voyager (NVIDIA, GPT-4) and Mindcraft (multi-agent LLM framework). The differentiators:
- The growing skill library lives as versioned source code on
main(not as JSON in RAM). Every Pi-written skill goes throughgit checkout -b → npm test → PR for operator review, making the loop auditable and rollback-safe. - Learned lessons live in a per-server SQLite knowledge base (
state/<host>/knowledge.db), with recipes, mob intel, block intel, lessons, deaths, post-mortems, points-of-interest, cached wiki pages, and a chat log. The dispatcher reads from it before acting.
The name pepa-pi-bot is just the project's name (pepa from the original test server, pi from the original runtime). The bot itself is server-agnostic.
Runtime modes
Two ways to run the bot. The hybrid runtime is the default — Pi-only is a fallback for experiments.
| Mode | Entry | When to use |
|---|---|---|
| Hybrid runtime (recommended) | npm run bot + npm run tui |
Day-to-day. Script-driven reflex tick + Ink TUI dashboard + Pi/Codex called only on demand. Fast, cheap, observable. |
| Pi-only (fallback) | npm run agent |
When you want every decision to go through an LLM (rare, but useful for experiments and code-writing sessions). |
See docs/runtime.md for the full hybrid runtime guide, IPC protocol, TUI hotkeys, and the self-improvement loop.
Concept
Most Minecraft AI bots ship as monolithic projects: hard-coded actions, fixed prompts, a single LLM provider, sometimes a single target server. This repo flips that around with a layered runtime.
┌──────────────────────────────────────────────────────────┐
│ TUI (Ink) — operator dashboard, attaches via Unix sock │
│ status / live log / MC chat / hotkeys / ask-Pi │
└──────────────────────┬───────────────────────────────────┘
│ newline-JSON
▼
┌──────────────────────────────────────────────────────────┐
│ runtime/bot.js — long-running Node daemon │
│ ├── Mineflayer client (MC TCP, AuthMe, chat, events) │
│ ├── Modes chain (self_preservation > hunger > shelter) │
│ │ priority interrupts before curriculum dispatch │
│ ├── Reflex loop (defend > eat > sleep > curriculum) │
│ │ consults coach/advice before every dispatch │
│ ├── perception.js — numeric-id findBlocks (VB-safe) │
│ ├── knowledge/ — SQLite store (recipes, mob/block intel,│
│ │ lessons, deaths, post-mortems, POI, chat log) │
│ ├── coach/postmortem — death → DB → Pi → lessons │
│ ├── coach/advice — recall lessons → override / avoid │
│ ├── persona/chatter — Russian narration in MC chat │
│ ├── world-journal + scenario-memory (persistent JSONL) │
│ ├── stuck-incident → critic (Pi) → proposal │
│ ├── auto-improve → auto-patch → npm test → PR (review) │
│ └── pi-bridge — spawn `pi -p` only on demand │
└──────────────────────┬───────────────────────────────────┘
│ TCP 25565 (any host/port)
▼
ANY Minecraft Java server (configured in .env)
The reflex loop is the brain stem. Pi is the cortex — called only when the reflex loop is genuinely stuck, or when the operator asks for help via the TUI. Mineflayer is the body. The skills, reflexes, and supervision loop are meant to grow over time — both by hand and by the bot itself proposing patches.
Prerequisites
| Tool | Why | How to get it |
|---|---|---|
Pi ≥ 0.75 |
The agent runtime. Reads AGENTS.md, loads skills, calls the LLM. |
curl -fsSL https://pi.dev/install.sh | sh |
Node.js ≥ 20 |
Required by Pi and by Mineflayer. | brew install node / nvm install 20 |
| An LLM credential | One of: OpenAI / Anthropic / Google API key, or an OAuth-authenticated subscription (/login inside Pi). ChatGPT Pro and Claude Max work via OAuth on supported providers. |
See Authentication |
| Access to some Minecraft server | The bot joins as a real player. Cracked or premium, online-mode or offline, doesn't matter — configure it in .env. |
— |
| Network access to that server | Direct TCP to host:port. |
— |
The bot does not need its own Minecraft client install, server admin access, RCON, or any server-side plugin. It joins as a vanilla player over the standard protocol.
Quickstart
# 1. Clone + configure
git clone git@github.com:xmatic-squad/pepa-pi-bot.git
cd pepa-pi-bot
cp .env.example .env
$EDITOR .env # set MC_HOST, MC_USERNAME, auth mode, AuthMe password, etc.
# 2. Install Node deps
npm install
# 3. (Optional) Authenticate Pi for the escalation hotkey
pi /login # OAuth flow — ChatGPT Pro / Claude Max
# or export OPENAI_API_KEY / ANTHROPIC_API_KEY
# 4. Run the bot — two terminals
# Terminal 1: the daemon (logs in stdout, persists state under state/<host>/)
npm run bot
# Terminal 2: the dashboard (Ink TUI). Hotkeys: p/s/r/c/a/k/v/!/y/q.
npm run tui
The TUI auto-reconnects to the bot if you restart it. Press q to leave the TUI; the bot keeps running.
Want the LLM-driven, single-process flavour?
npm run agentlaunches the original Pi runtime instead. Seedocs/runtime.mdfor the trade-offs.
Sending chat or asking Pi from the TUI
- Press
cin the TUI to enter chat mode — type, Enter sends into MC chat (rate-limited per.env). - Press
ato enter ask-Pi mode — type a prompt, Enter spawnspi -p "<prompt>". Output streams into the Pi panel without leaving the TUI.
TUI hotkeys cheatsheet
| Key | Effect |
|---|---|
p |
Pause / resume the reflex loop (MC connection stays). |
s |
Stop the bot (graceful disconnect + cleanup). |
r |
Force a fresh status snapshot. |
c |
Send a chat message into MC. |
a |
Ask Pi (one-shot subprocess). |
k |
Run one registered skill with optional JSON args. |
v |
Capture a viewer screenshot for debugging. |
! |
Force a critic-backed incident/proposal. |
y |
Open the latest pending proposal (badge appears in status bar). In the panel: y approve, n/Esc close. |
q |
Quit the TUI — bot keeps running. |
Self-improvement loop (short version)
The bot proposes its own patches when something repeatedly fails. As of v0.2.0 the loop ends in a PR for operator review, not a direct merge to main:
- Reflex action fails 5× with the same label, or
noProgressReasonstays stuck → markdown proposal lands understate/<host>/proposals/. runtime/auto-improve.jswatcher (2 s poll, 10 s debounce) picks it up and spawnsscripts/auto-patch.jsdetached.auto-patch.js: refuses on dirty tree, moves proposalpending → approved/, creates branchauto/<slug>off main, runspi -p "<patch prompt>"with 10-min timeout.- If Pi committed and only touched the proposal's
editScope, ANDnpm run lint-patch+npm testboth pass →git push origin auto/<slug>+gh pr createagainstmain. - Operator reviews the PR on GitHub and merges. Branch protection on
mainrequires approval — auto-patch cannot merge itself. - Supervisor watches
runtime/**/*.jsand hot-restarts the child on file change once the merged commit is pulled.
Manual paths still work:
- TUI hotkey
yopens the latest pending proposal for inspection. npm run propose:apply <filename>— attended version: same flow but leaves the branch local without opening a PR.PEPA_AUTO_PATCH_MERGE=cherry-pickenv reverts to the legacy direct-to-main behaviour (not recommended; bypasses review).
See docs/runtime.md for the full lifecycle. Note (2026-05-25): MC chat is now dialog-only — operator commands (come, pause, stop, …) are no longer dispatched from chat; use the TUI for local control. See plans/autonomous-survival-bot-prd.md for the survival-bot pivot.
Authentication
Two dimensions:
1. Minecraft auth. Configured in .env via MC_AUTH_MODE:
offline— cracked servers. Any nickname works. No external auth call.microsoft— premium / online-mode servers. Mineflayer handles the device-code flow on first connect and caches the token in~/.minecraft-auth/.
2. LLM auth. Pi supports 15+ providers and two credential modes:
- OAuth subscription login —
pithen/logininside the TUI. Suitable for ChatGPT Plus/Pro, Claude Max, and other subscriptions that ship an OAuth flow. No metered API billing. - API key environment variables —
OPENAI_API_KEY,ANTHROPIC_API_KEY,GOOGLE_API_KEY, etc. Metered, but no UI prompt.
You can mix providers via --provider openai --model gpt-5 at launch — cheaper models for idle ticks, smarter ones for hard decisions.
How the agent extends itself
Pi has first-class support for three growth surfaces:
skills/— Markdown-defined capabilities Pi can invoke. The agent canWritenew ones at runtime when it discovers a missing capability.extensions/— TypeScript modules registering new tools, commands, or UI tweaks. Installed project-locally viapi install -l npm:<pkg>/pi install -l git:<url>, or written in-tree.prompts/— Reusable prompt templates. Useful for cron-driven tick prompts ("what should I do next minute?").
The opening AGENTS.md instructs the agent to start by writing a mineflayer-bridge extension that can:
- connect to the configured MC server (any host/port/version)
- handle the configured auth mode (offline or microsoft)
- if a login plugin like AuthMe is present, perform
/registerand/loginfrom a password supplied in.env - emit world events back into the agent loop
- expose
chat / move / dig / place / equip / attackas Pi tools
Everything beyond that — farming, exploration, base-building, player interaction, server-specific quirks — should emerge from the agent itself.
Everything in the repo
A hard rule, mirrored in AGENTS.md: every artefact the agent produces lives in this repo, never in the user's ~/.pi/ directory. That includes skills, extensions, prompt templates, project Pi settings (.pi/settings.json), and per-server state (state/<MC_HOST>/).
The point is reproducibility and community growth: a fresh git clone should bring along every skill any contributor has written. Pi's own built-in skills (skill-creator, agent-browser, etc.) stay user-global — the agent is allowed to use them, but anything it authors lands under ./skills/ or ./extensions/ here.
See CONTRIBUTING.md for the skill format and how to propose changes.
Project layout
pepa-pi-bot/
├── README.md ← you are here
├── AGENTS.md ← seed prompt, loaded by the Pi-only runtime
├── .env.example ← all required env vars, no secrets
├── package.json ← node deps + scripts (`bot`, `tui`, `agent`)
├── runtime/ ← hybrid runtime (script reflex + IPC server)
│ ├── bot.js long-running Mineflayer daemon
│ ├── reflex.js priority-ordered behaviours, no LLM
│ ├── perceive.js snapshot builder
│ ├── ipc-server.js Unix-socket server
│ ├── ipc-protocol.js shared IPC contract
│ └── pi-bridge.js spawn `pi -p` on demand
├── tui/ ← Ink TUI dashboard
│ ├── tui.tsx
│ └── ipc-client.js
├── skills/ ← markdown skills (grown by bot or operator)
├── extensions/ ← Pi extensions (mindcraft-skills, mineflayer-bridge)
├── prompts/ ← reusable prompt templates
└── docs/
├── runtime.md hybrid runtime guide (start here)
├── architecture.md longer-form design notes
├── memory-model.md per-server state layout
├── roadmap.md phased plan
└── …
Safety boundaries
Server-agnostic but with hard defaults the agent must respect on any server it joins:
- Never request OP / admin rights in chat.
- Never break or modify other players' builds without explicit human request.
- Never spam chat — built-in rate limit (
CHAT_RATE_LIMIT_PER_MINin.env). - Never leak secrets from
.env(auth passwords, API keys) into chat, world signs, books, commits, or web fetches. - No destructive bash in the repo (
rm -rf, force pushes) without operator confirmation. - Stop and wait if kicked or banned — do not auto-reconnect indefinitely.
These are mirrored in AGENTS.md and re-stated at the top of any system prompt that overrides it.
Status
🌳 Phase 0 — Body done. Bridge online, AuthMe handled, hello sent. See skills/server-onboarding.md.
🌳 Phase 1 — Presence implemented: bridge stays online with bounded reconnects, rolling chat buffer, status/recent-chat tools, sparing replies.
🌿 Phase 2 — Locomotion/build rails in progress: mineflayer-pathfinder is wired with guarded mc_goto, plus mc_build_pyramid_5x5. Phase 5 self-extension and Phase 6 escalation logging are implemented.
🌱 Phase 3 — Goal-driven autonomy seeded: docs/memory-model.md defines shared-knowledge vs personal-memory; per-server goal.md / plan.md / current-task.json / diary/ shape autonomous behaviour.
🌿 Survival-bot pivot (2026-05-25) — the bot is becoming a self-sufficient survival resident of the configured server. MC chat is dialog-only; operator/player chat commands are recorded but not dispatched (TUI is the only local control plane). The hybrid runtime now has enriched perception, priority modes, a skill-driven curriculum, food acquisition, bed/sleep, base/chest/shelter/farm skills, persistent skill metrics, scenario memory, and a scoped auto-patch loop with npm test smoke gating. Full plan: plans/autonomous-survival-bot-prd.md (local-only, gitignored).
🌿 v0.2.0 — self-learning agent (2026-05-27). SQLite-backed knowledge base at state/<host>/knowledge.db with seeded recipes / mob intel / block intel / 12 starter lessons. Every death is captured with full context into a deaths table; a periodic Pi-coach loop (≤3 calls/hour, 12-min cooldown) batches unanalysed deaths and extracts generalised lessons. A 30-min self-reflection loop asks Pi "are you in a loop or making progress?" and records the verdict + lessons to state/<host>/reflections/. The reflex chain consults coach/advice before every dispatch (including fallback paths) — high-confidence lessons can override (attack creeper → survive.flee) or back-off. Bot narrates major events in Russian MC chat (≤8 lines/hour, ≥75 s gap). New escape skill survive.pillar-up for pit terrain; wedged-emergency reflex fires it after 60s of no horizontal progress. Auto-patch now opens a PR for operator review instead of cherry-picking onto main — main is protected, the operator is the only approver. Currently shipped: rc.1 (substrate) + rc.2 (P0 hardening) + rc.3 (escape + advice-in-fallback). Current state, next steps, and known issues: dev/v0.2.0/STATUS.md. Original design: docs/v0.2.0-self-learning.md.
Full plan: docs/roadmap.md. Memory layout: docs/memory-model.md. Day-to-day judgement: "Operating principles" in AGENTS.md.