Thesis tested

This is the first phase that supports the complete core claim: a hidden local change makes a deployed learned model wrong, and a small verified edit restores multi-step competence earlier and more reliably than incremental fine-tuning.

One common interaction budget

Sixty fresh worlds covered two independently gated change classes: action rate and repair delay. Frozen, incremental fine-tuning, symbolic repair and hybrid composition consumed the same visible transition prefixes at budgets 0, 8, 16, 32 and 64.

The hybrid could additionally spend counted expert answers. Evaluation measured one-step accuracy, H5 and H8 recursive rollouts, bounded planning, false plans, update time, invariant violations, version retention and leak boundaries.

Recovery across the budget

The result was not a last-checkpoint win. Hybrid-minus-fine-tuning confidence intervals stayed strictly positive for H5 recovery AUC in both shift families, while the longer H8 rollout improved and planning remained non-inferior.

Normalized area under competence versus visible-transitions curve
ShiftMetricFrozenFine-tunedHybridHybrid − FT 95% CI
RateH5 rollout0.0790.2040.581[0.265, 0.491]
RatePlanning0.5600.5950.809[0.132, 0.298]
DelayH5 rollout0.0850.1220.762[0.594, 0.687]
DelayPlanning0.4360.5030.599[0.044, 0.148]

Final multi-step competence

At budget 64, the hybrid preserved unaffected one-step competence while materially improving recursive rollout and planning. No strategy violated a planning invariant, but candidate plans were still checked against the reference environment.

Final competence at budget 64
ShiftStrategyOne-stepH5H8PlanningFalse plans
RateFine-tuned0.4580.3710.2670.64050
RateHybrid0.8960.7090.5710.88717
DelayFine-tuned0.2730.1220.0820.50959
DelayHybrid0.9800.9180.8860.62329

What the resource comparison means

Local verified repair used 69–75% less measured update CPU than fine-tuning on the tiny 1,468-parameter model. That ratio was implementation evidence, not a production claim. The more important comparison was the resource vector: equal trajectories, about one answer, far more rollout recoveries.

The scale question—whether the cost advantage grows with learned-model capacity—was explicitly left to later phases.

The milestone

Phase 6A turned the project from a rule-graph prototype into an evidenced systems thesis. The asset is the loop: gap detection, targeted teaching, typed compilation, verification, hybrid execution, provenance and versioning.

It remained a controlled synthetic result. Industrial generalization, million-parameter competence and uncertainty-calibrated planning were still open.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.