v0.4.0 vNext — closed-loop world model + settlement contract (#29)

* feat(v0.4.0): vNext — closed-loop world model + settlement contract

Implements the vNext architecture from the research doc: demote the noisy
multi-rail planner in favour of a closed loop (world truth → invariant check)
plus a single utility-driven goal authority.

L1 services (fix no_drop / silent pathfinder hang first):
- InventoryLedger: diff-based "did I actually get it" verifier; acquire-food
  now confirms via ledger.gainedSince instead of the unreliable count/event.
- MotionService.gotoSafe: wall-clock timeout + progress watchdog +
  path_update(noPath/timeout) → structured {reached|stuck|timeout|nopath}.

L3 plan — unify the three competing rails (curriculum/manifesto/storyline):
- Settlement Contract: ordered M0–M9 milestones, each invariant-checked
  against an authoritative world view (early steps delegate to the proven
  curriculum; late game adds farming).
- InvariantChecker + predicate library; GoalManager selects the lowest unmet
  milestone via utility argmax (food-urgency preempts, DEPS-style).
- Wired into the scheduler: bot.js precomputes snapshot.contract; reflex.js
  consumes it in place of the storyline rail. Manifesto L0 still preempts.

Eval + robustness:
- Village Score (single 0..1 metric) on the snapshot + TUI "build" line.
- survive.dig-in skill + dusk_dig_in mode (exposed at night, no bed → cover).
- approach_block helper (GoalNear + lookAt, avoids GoalLookAtBlock #341).

+28 new tests (450 total green). LLM remains entirely off the tick path.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(v0.4.0): finish vNext plan — anti-loop, skill-graph, worldDelta diff, flee→motion

Completes the remaining v0.4.0 plan items and one fix motivated by a live
in-game observation (flee hanging 30s against a persistent zombie).

- flee → MotionService.gotoSafe: structured {stuck|timeout|nopath} in ~4s with
  a blind-retreat fallback, instead of the observed 30s pathfinder hang + 3
  watchdog replans. Movements setup guarded so it is unit-testable.
- QW5 anti-loop (runtime/anti-loop.js): same skill failing >=3x in 5min →
  30min blacklist (reflex shouldSkip) + one-shot improvement_request
  (bot.js drainFired -> writeProposal).
- 4.1 closed-loop worldDelta: runSkill snapshots inventory before execute and
  attaches the real delta (_invObserved) to every successful result; opt-in
  skill.expectGain asserts the claimed gain or returns world_unchanged.
- 3.6 skill-graph (Plan4MC): declarative requires/produces for ~20 skills;
  prerequisitesMet/canRun/runnableFrontier; GoalManager annotates suggestions
  with blockedBy when prereqs are unmet.

+22 tests (472 total green). Live smoke confirmed dig-in works and no new
errors; flee loop is what this commit's flee migration addresses.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Yuriy Mayatnikov <mayatnikov@me.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
This commit was merged in pull request #29.
This commit is contained in:
Yuriy Mayatnikov
2026-05-28 11:29:39 +03:00
committed by GitHub
co-authored by Claude Opus 4.7 mayatnikov
parent 15b6c11002
commit c910457817
31 changed files with 2236 additions and 61 deletions
+78
View File
@@ -0,0 +1,78 @@
// Village Score (L5 eval) — one number for "is the bot actually building a
// settlement, or just walking?" (research §2). The research's formula mixes
// milestones, food stock, fence closure, lit tiles, uptime, distinct skills
// and dialog quality. We compute the subset that is *observable* today; fence
// polygon / lit-tile fraction stay at 0 until those skills exist (the score is
// honest about what it can measure rather than faking precision).
//
// Pure: snapshot + a few derived inputs in, { score, components } out. Score is
// normalised to 0..1 so a dashboard / TUI can show a single percentage.
const TWO_HOURS_MS = 2 * 60 * 60 * 1000;
const DISTINCT_SKILL_TARGET = 12;
function clamp01(n) {
if (!Number.isFinite(n)) return 0;
return n < 0 ? 0 : n > 1 ? 1 : n;
}
function milestoneFraction(contract) {
if (!contract || !contract.total) return 0;
return clamp01(contract.completed / contract.total);
}
function foodSecurity(snapshot) {
if (snapshot?.hasFood) return 1;
return clamp01((snapshot?.food ?? 0) / 18);
}
function baseEstablished(snapshot) {
const loc = snapshot?.locations ?? {};
const want = ["base", "shelter", "chest"];
const have = want.filter((k) => loc[k]).length;
return clamp01(have / want.length);
}
function distinctSkillsSucceeded(metrics) {
if (!metrics) return 0;
const n = Object.values(metrics).filter((m) => (m?.ok ?? 0) > 0).length;
return clamp01(n / DISTINCT_SKILL_TARGET);
}
function uptimeFraction(uptimeMs) {
return clamp01((uptimeMs ?? 0) / TWO_HOURS_MS);
}
function survival(snapshot) {
return clamp01((snapshot?.health ?? 0) / 20);
}
const WEIGHTS = Object.freeze({
milestones: 0.35,
food: 0.15,
base: 0.15,
distinctSkills: 0.15,
uptime: 0.10,
survival: 0.10,
});
export function computeVillageScore(snapshot, { contract, uptimeMs = 0, metrics = null } = {}) {
const components = {
milestones: milestoneFraction(contract),
food: foodSecurity(snapshot),
base: baseEstablished(snapshot),
distinctSkills: distinctSkillsSucceeded(metrics),
uptime: uptimeFraction(uptimeMs),
survival: survival(snapshot),
};
let score = 0;
for (const [k, w] of Object.entries(WEIGHTS)) score += w * components[k];
return {
score: Math.round(clamp01(score) * 1000) / 1000,
components,
milestonesCompleted: contract?.completed ?? 0,
milestonesTotal: contract?.total ?? 0,
};
}
export const _internal = { clamp01, WEIGHTS, milestoneFraction, foodSecurity, baseEstablished };