Adds an explicit hierarchical needs catalogue that the reflex consults
on every tick. The bot now pursues tangible intermediate goals (food,
wood tools, shelter, stone tools, ...) instead of inheriting whatever
the curriculum thought was "next".
Ladder:
L0 alive HP>5, food>0, not in lava, not panic-near hostile
L1 food ≥6 food items in inventory (or sated + any food)
L2 tools_wood wooden_pickaxe + wooden_axe + wooden_sword
L3 shelter_basic bed placed nearby or in inventory
L4 tools_stone stone-tier triplet
L5 armor_basic any chestplate (pursue=null until craft.leather-*
lands; ladder gracefully skips)
L6 food_security ≥16 food items
L7 tools_iron iron-tier triplet (pursue=gather.stone for now)
L8 armor_iron iron chestplate (pursue=null for now)
L9 village_seed bed + chest in nearby blocks
L10 village_full never detected, falls through to curriculum
Each need has detect(snapshot) → bool and pursue(snapshot) →
{skillId, args} | null. The ladder picks the LOWEST unsatisfied
pursuable need. Needs whose pursue is null get recorded as
blockedNeeds and the walk continues — no stalling on missing skills.
Wired into curriculumReflex: manifesto takes precedence over
curriculum.plan when it has a concrete suggestion. Tests can pass
ctx.disableManifesto=true to exercise the curriculum branch
in isolation (existing reflex tests keep passing this way).
Pi self-reflection prompt now includes
"activeNeed (Maslow ladder L0-L10): L2 tools_wood → gather.logs"
so Pi advises at the right level instead of giving generic guidance.
skillId returned by pursue() is validated against the live registry
(rc.1 plumbing) — manifesto cannot accidentally dispatch a
hallucinated skill name.
Tests: 315 green (was 279 on rc.1, +36 new):
- runtime/manifesto/needs.test.js — 24 tests (per-need detect/pursue,
helper sums)
- runtime/manifesto/state.test.js — 10 tests (ladder walk, hostile
takeover at L0, armor skipping, caching)
- runtime/reflex.test.js — 2 integration tests (manifesto overrides
curriculum plan; well-fed bot pursues tools_stone)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
6.8 KiB
pepa v0.3.0 — status
Live tracking document for the v0.3.0 iteration ("Maslow + Awareness").
See PLAN.md for the full design.
Shipped
rc.1 — Live skill registry + Fast advisor scaffold
Root problem solved: 47/47 Pi-extracted lessons in v0.2.x had
applied_count = 0 because Pi was hallucinating skill ids
(relocate.surface, choose.safe.surface, survive.shelter,
gather.visible_log, …) that don't exist in the registry. Both halves
fixed: (a) Pi now sees the real registry in its system prompt,
(b) anything that still slips through gets rejected at consult time.
runtime/skill-registry.js— single source of truth wrappingskills/index.js. Exports:listSkillIds()— live id listisRegistered(id)— bool checkdescribeSkill(id)— id/title/timeoutMsskillRegistryPrompt({ limit })— prompt-ready block grouped by namespace, with "USE ONLY THESE, never invent" instruction
runtime/llm/provider.js— OpenAI-compatible chat client, env-driven:PEPA_FAST_LLM_BASE_URL(defaulthttps://api.openai.com/v1)PEPA_FAST_LLM_API_KEY(required to enable; safe no-op otherwise)PEPA_FAST_LLM_MODEL(required)PEPA_FAST_LLM_TIMEOUT_MS(default 8000)- Supports JSON-mode via
response_format: { type: "json_object" } - Surfaces
not_configured,no_model,http_<status>,network_error,timeout,bad_jsoncodes
runtime/coach/fast-advisor.js— tactical "what now?" tier. Scaffold only in rc.1; auto-trigger comes in rc.3.advise({snapshot, reason, recentSkillIds, lessonsTail})→{action: 'switch_skill'|'continue'|'wait', skillId?, rationale}- Rejects any returned
skill_idnot in the live registry - Rate-limit: 6 calls/hour, 30s cooldown between calls
- System prompt embeds registry; user prompt carries snapshot + trigger
runtime/coach/advice.js:normalisePreferSkill()now returnsnullfor anything not in registry/mode-map (was: passed through unchanged → dispatcher crashed atrunSkill())- Logs
warnline when a hallucinated prefer_skill is dropped
runtime/coach/postmortem.js:- Pi prompt includes the live registry block (
skillRegistryPrompt) with a "CRITICAL: USE ONLY THESE" instruction - On insert, drops
prefer_skill/avoid_skillthat's neither a registered id nor a known mode name; warn-logs the count
- Pi prompt includes the live registry block (
runtime/coach/reflect.js— same treatment as postmortem (registry in prompt + write-time filter)
Tests: 279 green (was 257 on rc.3). Added:
runtime/skill-registry.test.js— 5 testsruntime/llm/provider.test.js— 9 testsruntime/coach/fast-advisor.test.js— 10 tests
rc.2 — Manifesto / Needs ladder L0-L10
Root problem solved: pre-v0.3.0 the bot had no notion of intermediate
goals. The curriculum produced a single "next milestone" but no
hierarchy. So when the bot was wedged with no pickaxe, it kept trying
explore.far instead of recognising "I need wood → planks → pickaxe
first". Lessons from Pi couldn't help because there was no
internal-state language to express "L2 not satisfied".
The needs ladder gives the bot an explicit, ordered list of survival concerns. Each reflex tick picks the LOWEST unsatisfied need and dispatches a concrete skill toward it.
L0 alive HP>5, food>0, no lava, no creeper@close
L1 food ≥6 food items in inventory (or hungry+have any)
L2 tools_wood wooden_pickaxe + wooden_axe + wooden_sword
L3 shelter_basic bed placed nearby or in inventory
L4 tools_stone stone tier (pickaxe + axe + sword)
L5 armor_basic any chestplate equipped (pursue=null for now)
L6 food_security ≥16 food items
L7 tools_iron iron tier (pursue=gather.stone until craft.iron-* lands)
L8 armor_iron iron chestplate (pursue=null for now)
L9 village_seed bed + chest nearby
L10 village_full global goal (never detected, falls through to curriculum)
runtime/manifesto/needs.js— catalogue of 11 needs. Each hasdetect(snapshot)andpursue(snapshot). Pursue can returnnull(e.g. armor levels) and the ladder gracefully skips, recording the level as "blocked".runtime/manifesto/state.js—pickActiveNeed(snapshot)walks the ladder, picks the first unsatisfied + pursuable need. Returns{need, skillId, args, blockedNeeds}. 3-second cache to avoid re-walking the ladder on every micro-tick. ValidatesskillIdagainst the live registry (rc.1 piece) before returning — manifesto can't ship a hallucinated id.runtime/reflex.js:curriculumReflexnow consults manifesto FIRST. If a need dictates a skill, that's what gets dispatched. The curriculum plan is the fallback when manifesto has no concrete pursue.- Tests can pass
ctx.disableManifesto = trueto exercise the curriculum branch in isolation.
runtime/coach/reflect.js— Pi self-reflection prompt now includes the active need (L2 tools_wood → gather.logs (Деревянные орудия)) so Pi can give level-appropriate advice instead of generic suggestions.
Tests: 315 green (was 279 on rc.1, +36 new):
runtime/manifesto/needs.test.js— 24 tests (one per need detect/pursue)runtime/manifesto/state.test.js— 10 tests (ladder walk, caching, skipping)runtime/reflex.test.js— 2 new integration tests (manifesto-on overrides curriculum; well-fed bot pursues tools_stone)
rc.3 — (pending) Event-driven awareness + skill pre-emption
Next session quick start
- Read PLAN.md for the full design and per-rc breakdown.
- Check live DB to see if Pi-lesson application is improving:
After rc.1 deploys, expect Pi-coach/Pi-reflect
sqlite3 state/play.xmatic.team_25565/knowledge.db \ "SELECT source, COUNT(*) AS n, SUM(applied_count > 0) AS applied FROM lessons GROUP BY source ORDER BY n DESC;"appliedcount to start growing as the registry feedback closes the loop. - Set fast-advisor env when ready to test:
The advisor still isn't auto-triggered in rc.1 — it's wired in rc.3.
export PEPA_FAST_LLM_BASE_URL="https://<timeweb-endpoint>/v1" export PEPA_FAST_LLM_API_KEY="<key>" export PEPA_FAST_LLM_MODEL="gpt-5-mini" - Pick the next rc from PLAN.md.
Workflow notes
- main is protected — only operator merges PRs
- Tests:
npm test(279 green at last check), isolated under/tmp/ - The bot supervisor hot-restarts on file changes in
runtime/**/*.js - If something regresses badly, revert to v0.2.0-rc.3 commit
865aae1