v0.2.0-rc.2: P0 hardening — Pi headless, test state isolation, advice fixes

P0 (correctness):

1. PEPA_HEADLESS=1 guard in extensions/mineflayer-bridge.ts. When `pi -p`
   spawns a subprocess (banter, coach, planner, reflect, auto-patch), the
   bridge no longer attempts a second MC connect — the hybrid runtime
   already owns the nickname. runtime/pi-bridge.js sets the env var on
   every spawn. Root cause of the "two pepa_bot's racing for the slot"
   bug seen in reply-pi stderr.

2. Test state isolation in runtime/config.js. When running under the node
   test runner (detected via execArgv/argv) — or when PEPA_STATE_DIR is
   set — stateDir redirects to /tmp/pepa-test-state-<pid>/. log.js,
   scenario-memory, world-journal, and knowledge.db all follow.
   `npm test` no longer pollutes live scenarios.jsonl, world-journal.jsonl,
   or daily log files. Verified empirically: post-fix run added 0 test
   rows to the live scenarios file. Cleaned ~550 historical test rows
   from live state in the same change.

3. defendReflex outcome reporting (runtime/reflex.js). Previously a
   creeper-rule override marked the lesson succeeded=false BEFORE the
   flee skill returned. Now dispatchDefendFlee accepts {lessonId} and
   the onComplete fires reportAdviceOutcome with the actual flee result.

4. Mode-name → skill-id translation in runtime/coach/advice.js. Pi-coach
   occasionally returns prefer_skill values that are mode names
   ("night_shelter", "self_preservation", "hunger"). normalisePreferSkill
   maps these to SAFE_OVERRIDES entries before dispatch. Also handles
   "tunnel-out", "survive_flee", "survive flee" shapes.

New behavior:

5. Self-reflection loop (runtime/coach/reflect.js). Every 30 min, the
   bot asks Pi: "Are you making progress, or stuck in a loop? What
   should you do differently?" Pi answers with a verdict
   (progress/loop/recovering/idle/emergency), summary, next-action, and
   0-N new lessons. The reflection is written to
   state/<host>/reflections/<ts>.md and lessons land in the DB with
   source="pi-reflect". Rate-limited to 2 calls/hour. Wired through
   bot.js with the existing askPi + lastSnapshot accessor.

Tests: 246/246 green (+9 new: 6 advice mode-name + 4 reflect).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
2026-05-27 13:56:37 +03:00
co-authored by Claude Opus 4.7
parent 042793a53f
commit 552c67f61e
10 changed files with 517 additions and 13 deletions
+13 -4
View File
@@ -168,9 +168,13 @@ function dispatchDefendFlee(ctx, hostile, dist, opts = {}) {
ctx.lastFleeAttempt = { name: hostile.name, ts: Date.now() };
const fromEntity = matchingHostileEntity(ctx, hostile.name, dist);
const onComplete = opts.lessonId
? (res) => reportAdviceOutcome({ lessonId: opts.lessonId, succeeded: !!res?.ok })
: undefined;
ctx.dispatch(
() => fleeFrom(ctx.bot, fromEntity, 16),
`flee from ${hostile.name}`,
onComplete ? { onComplete } : {},
);
return { action: "dispatched", kind: "defend-flee", label: hostile.name };
}
@@ -203,12 +207,17 @@ function defendReflex(ctx) {
}
// v0.2.0 — consult learned lessons. If knowledge says "do not
// attack <hostile> in this state" (e.g. creeper rule, or no-weapon
// rule learned from post-mortems), flee instead. This is the
// closing of the learning loop for emergency combat.
// rule learned from post-mortems), flee instead. The lesson outcome
// is reported AFTER the flee skill finishes (via dispatchDefendFlee
// onComplete), not before — flee's success/failure is what proves
// or disproves the lesson, not the act of consulting it. This is
// the closing of the learning loop for emergency combat.
const advice = consultAdvice({ plannedSkillId: `attack ${hostile.name}`, snapshot: s });
if (advice.action === "avoid" || advice.action === "override") {
if (advice.lessonId) reportAdviceOutcome({ lessonId: advice.lessonId, succeeded: false });
return dispatchDefendFlee(ctx, hostile, dist, { ignoreCooldown: true });
return dispatchDefendFlee(ctx, hostile, dist, {
ignoreCooldown: true,
lessonId: advice.lessonId,
});
}
ctx.dispatch(
() => attackNearestUntilClear(ctx.bot, hostile.name, {