v0.4.0 vNext — closed-loop world model + settlement contract (#29)

* feat(v0.4.0): vNext — closed-loop world model + settlement contract

Implements the vNext architecture from the research doc: demote the noisy
multi-rail planner in favour of a closed loop (world truth → invariant check)
plus a single utility-driven goal authority.

L1 services (fix no_drop / silent pathfinder hang first):
- InventoryLedger: diff-based "did I actually get it" verifier; acquire-food
  now confirms via ledger.gainedSince instead of the unreliable count/event.
- MotionService.gotoSafe: wall-clock timeout + progress watchdog +
  path_update(noPath/timeout) → structured {reached|stuck|timeout|nopath}.

L3 plan — unify the three competing rails (curriculum/manifesto/storyline):
- Settlement Contract: ordered M0–M9 milestones, each invariant-checked
  against an authoritative world view (early steps delegate to the proven
  curriculum; late game adds farming).
- InvariantChecker + predicate library; GoalManager selects the lowest unmet
  milestone via utility argmax (food-urgency preempts, DEPS-style).
- Wired into the scheduler: bot.js precomputes snapshot.contract; reflex.js
  consumes it in place of the storyline rail. Manifesto L0 still preempts.

Eval + robustness:
- Village Score (single 0..1 metric) on the snapshot + TUI "build" line.
- survive.dig-in skill + dusk_dig_in mode (exposed at night, no bed → cover).
- approach_block helper (GoalNear + lookAt, avoids GoalLookAtBlock #341).

+28 new tests (450 total green). LLM remains entirely off the tick path.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.4.0): finish vNext plan — anti-loop, skill-graph, worldDelta diff, flee→motion

Completes the remaining v0.4.0 plan items and one fix motivated by a live
in-game observation (flee hanging 30s against a persistent zombie).

- flee → MotionService.gotoSafe: structured {stuck|timeout|nopath} in ~4s with
  a blind-retreat fallback, instead of the observed 30s pathfinder hang + 3
  watchdog replans. Movements setup guarded so it is unit-testable.
- QW5 anti-loop (runtime/anti-loop.js): same skill failing >=3x in 5min →
  30min blacklist (reflex shouldSkip) + one-shot improvement_request
  (bot.js drainFired -> writeProposal).
- 4.1 closed-loop worldDelta: runSkill snapshots inventory before execute and
  attaches the real delta (_invObserved) to every successful result; opt-in
  skill.expectGain asserts the claimed gain or returns world_unchanged.
- 3.6 skill-graph (Plan4MC): declarative requires/produces for ~20 skills;
  prerequisitesMet/canRun/runnableFrontier; GoalManager annotates suggestions
  with blockedBy when prereqs are unmet.

+22 tests (472 total green). Live smoke confirmed dig-in works and no new
errors; flee loop is what this commit's flee migration addresses.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
This commit was merged in pull request #29.
This commit is contained in:
Yuriy Mayatnikov
2026-05-28 11:29:39 +03:00
committed by GitHub
co-authored by Claude Opus 4.7 mayatnikov
parent 15b6c11002
commit c910457817
31 changed files with 2236 additions and 61 deletions
+45
View File
@@ -30,6 +30,7 @@ import { skill as flee } from "./flee.js";
import { skill as sleep } from "./sleep.js";
import { skill as tunnelOut } from "./recovery-tunnel-out.js";
import { skill as pillarUp } from "./pillar-up.js";
import { skill as digIn } from "./dig-in.js";
import { skill as escapePitSafe } from "./escape-pit-safe.js";
import { skill as diagPhysics } from "./diagnose-physics.js";
import { skill as diagScan, matchSkill as diagMatch } from "./diagnose-scan.js";
@@ -77,6 +78,7 @@ register(flee);
register(sleep);
register(tunnelOut);
register(pillarUp);
register(digIn);
register(escapePitSafe);
register(diagPhysics);
register(diagScan);
@@ -128,6 +130,18 @@ export const RUNNER_CODES = Object.freeze({
DONE: "done",
});
// Signed inventory diff between two count Maps (from InventoryLedger.mark/
// snapshot). Used to attach the real world change to a skill result.
function invDiff(before, after) {
const out = {};
const names = new Set([...(before?.keys?.() ?? []), ...(after?.keys?.() ?? [])]);
for (const n of names) {
const d = (after?.get?.(n) ?? 0) - (before?.get?.(n) ?? 0);
if (d !== 0) out[n] = d;
}
return out;
}
function normaliseResult(res, fallbackCode) {
const ok = !!res?.ok;
return {
@@ -220,6 +234,12 @@ export async function runSkill(id, ctx, args = {}) {
return result;
}
// WorldDelta diff layer (research §TL;DR): snapshot the inventory before
// execute so we can attach the REAL inventory change to the result and,
// for skills that opt in via `expectGain`, assert the claimed gain actually
// happened instead of trusting the skill's own bookkeeping.
const ledgerBefore = ctx?.ledger?.mark?.() ?? null;
const timeoutMs = skill.timeoutMs ?? 30_000;
let raw;
try {
@@ -275,6 +295,31 @@ export async function runSkill(id, ctx, args = {}) {
return failed;
}
}
// Closed loop: compare the inventory now vs the pre-execute baseline.
if (result.ok && ledgerBefore && ctx?.ledger) {
try { if (ctx.bot) ctx.ledger.update(ctx.bot); } catch {}
const observed = invDiff(ledgerBefore, ctx.ledger.snapshot());
if (Object.keys(observed).length > 0) {
result.worldDelta = { ...(result.worldDelta ?? {}), _invObserved: observed };
}
// Opt-in strict check: the world must show the claimed gain.
if (skill.expectGain) {
const gain = ctx.ledger.gainedSince(ledgerBefore, skill.expectGain.matcher);
if (gain < (skill.expectGain.min ?? 1)) {
const failed = {
ok: false,
code: "world_unchanged",
detail: `${id} reported ok but ${skill.expectGain.label ?? "expected items"} did not increase (gain ${gain})`,
worldDelta: result.worldDelta,
};
if (typeof skill.recover === "function") {
try { failed.recovery = skill.recover(ctx, failed) ?? null; } catch {}
}
return failed;
}
}
}
if (!result.ok && typeof skill.recover === "function") {
try {
result.recovery = skill.recover(ctx, result) ?? null;