This is the first phase that supports the complete core claim: a hidden local change makes a deployed learned model wrong, and a small verified edit restores multi-step competence earlier and more reliably than incremental fine-tuning.
One common interaction budget
Sixty fresh worlds covered two independently gated change classes: action rate and repair delay. Frozen, incremental fine-tuning, symbolic repair and hybrid composition consumed the same visible transition prefixes at budgets 0, 8, 16, 32 and 64.
The hybrid could additionally spend counted expert answers. Evaluation measured one-step accuracy, H5 and H8 recursive rollouts, bounded planning, false plans, update time, invariant violations, version retention and leak boundaries.
Recovery across the budget
The result was not a last-checkpoint win. Hybrid-minus-fine-tuning confidence intervals stayed strictly positive for H5 recovery AUC in both shift families, while the longer H8 rollout improved and planning remained non-inferior.
| Shift | Metric | Frozen | Fine-tuned | Hybrid | Hybrid − FT 95% CI |
|---|---|---|---|---|---|
| Rate | H5 rollout | 0.079 | 0.204 | 0.581 | [0.265, 0.491] |
| Rate | Planning | 0.560 | 0.595 | 0.809 | [0.132, 0.298] |
| Delay | H5 rollout | 0.085 | 0.122 | 0.762 | [0.594, 0.687] |
| Delay | Planning | 0.436 | 0.503 | 0.599 | [0.044, 0.148] |
Final multi-step competence
At budget 64, the hybrid preserved unaffected one-step competence while materially improving recursive rollout and planning. No strategy violated a planning invariant, but candidate plans were still checked against the reference environment.
| Shift | Strategy | One-step | H5 | H8 | Planning | False plans |
|---|---|---|---|---|---|---|
| Rate | Fine-tuned | 0.458 | 0.371 | 0.267 | 0.640 | 50 |
| Rate | Hybrid | 0.896 | 0.709 | 0.571 | 0.887 | 17 |
| Delay | Fine-tuned | 0.273 | 0.122 | 0.082 | 0.509 | 59 |
| Delay | Hybrid | 0.980 | 0.918 | 0.886 | 0.623 | 29 |
What the resource comparison means
Local verified repair used 69–75% less measured update CPU than fine-tuning on the tiny 1,468-parameter model. That ratio was implementation evidence, not a production claim. The more important comparison was the resource vector: equal trajectories, about one answer, far more rollout recoveries.
The scale question—whether the cost advantage grows with learned-model capacity—was explicitly left to later phases.
The milestone
Phase 6A turned the project from a rule-graph prototype into an evidenced systems thesis. The asset is the loop: gap detection, targeted teaching, typed compilation, verification, hybrid execution, provenance and versioning.
It remained a controlled synthetic result. Industrial generalization, million-parameter competence and uncertainty-calibrated planning were still open.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.