Operator feedback: "бот должен быть полностью автономным — сам себя
улучшать и чинить, в этом и есть смысл; все что я вижу пока что он
стоит на месте и кидает proposals на каждый чих — это кардинально не
то что я хочу". Acted on:
1. Trigger filter — proposals only on real bugs.
runtime/bot.js classifies failure detail into bug / timeout /
feature-gap / other. The 5-in-a-row trigger fires only when the run
contains a bug (TypeError / Cannot read / is not defined …) OR is
entirely timeouts on the same operation. Feature gaps like "no
reachable log within 32 blocks", "no food in inventory", "no bed in
range", "no target in reach" are SKIPPED — the reflex layer routes
around them (noTreesUntil → wander, etc). The LLM has no business
patching code for missing inventory.
Threshold raised 3 → 5 in a row. Cooldown unchanged (30 min).
2. Auto-apply, no operator-in-the-loop.
New runtime/auto-improve.js polls proposals/ every 2s. When it sees
a new .md and 10s have passed since first sighting (debounce),
spawns scripts/auto-patch.js detached.
New scripts/auto-patch.js: refuses on dirty tree, moves proposal
pending → approved/, branches `auto/<slug>` off main, runs `pi -p`
with 10-min timeout. If Pi committed AND every changed file is
under runtime/ → cherry-picks onto main. Otherwise discards the
branch. No push, no PR. Audit trail in state/<host>/proposals/approved/.
Rate limit: 15-min cooldown between finished runs + 4/hour hard cap.
3. Auto-rollback on bad patches.
runtime/supervisor.js: when MAX_RESTARTS_PER_MINUTE is exceeded
AND `git log -1 HEAD` is younger than 15 min AND HEAD touched
runtime/, runs `git reset --hard HEAD~1`. Up to MAX_ROLLBACKS=3
lifetime, then exits 1 for manual investigation. Restart counters
are reset after a successful rollback so the next attempt isn't
immediately killed.
4. current-task.json slim.
No longer stores the full perception snapshot (was ~3 KB per write
× every action). Position only — sufficient as a resume anchor.
Slim snapshot still goes into the proposal markdown for context.
docs/runtime.md — rewrote the self-improvement section: full flow
diagram, classification rules, all rate-limit knobs, manual escape
hatches kept but documented as rarely-needed.
Also cleared 5 stale proposals from previous smoke tests so the first
production run isn't burning Pi tokens on stale bugs that have since
been fixed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>