A bounded correction can preserve strong predictive recovery, but selective competence is not demonstrated unless approximate knowledge creates enough consequential mistakes for asking and abstention to be falsifiable.
Replacing an exact answer with a bounded one
Earlier CSTR evidence gave the repair system the exact hidden actuator scale. Phase 6E prospectively replaced that easy answer with a width-0.20 interval and at most one conditional width-0.10 refinement.
Nine scale hypotheses were propagated through H25 prediction and H60 planning. A plan could execute only if every member satisfied the existing target and temperature certificate. Otherwise the system asked once more or selected a neutral fallback.
The approximate patch retained useful prediction
All eight changed streams triggered at observation 8. Every coarse and refined interval contained the hidden scale, passed the visible-window verifier and produced an evidence-linked immutable version.
The final midpoint recovered 95.49% of lost H25 competence. Fine-interval coverage reached 99.78% over 64 fresh stress episodes while mean normalized uncertainty width fell 45.29%.
| Strategy | Mean H25 error | Evidence status |
|---|---|---|
| Frozen | 0.052701 | Development comparator |
| Full fine-tuning | 0.018805 | Development comparator |
| Rank-2 adapter | 0.027443 | Development comparator |
| Coarse midpoint | 0.005899 | Approximate answer |
| Final midpoint | 0.003443 | After conditional refinement |
| Exact reference | 0.000587 | Evaluation only |
Why 21 successful executions still did not pass
The exact evaluation reference could reach 23 of 32 tasks. The selective policy executed 21 and succeeded on all 21 with no safety violation.
But the naïve coarse midpoint already succeeded on 22 of those 23 reachable tasks. There was only one naïve failure to avoid, below the preregistered minimum of three. One avoided mistake is positive descriptive evidence, not task-level dominance or a validated abstention policy.
A perfect execution rate among selected tasks cannot validate selection when the underlying workload contains almost no failures.
Four failed gates remain part of the result
Two of four unchanged controls crossed the frozen detector threshold. The harness did not issue answers on those known controls, but a deployed system would not know the regime label, so they remain false positives.
Reference-reachable abstention precision was 50%, below the registered 60%. Every world requested the second answer, and two interval compile-and-verify operations took a median 0.13664 seconds versus 0.10901 seconds for the small full-FT baseline.
| Gate group | Observed | Decision |
|---|---|---|
| Nominal controls | 2/4 clean | Fail |
| Prediction and interval calibration | 95.49% recovery · 99.78% coverage | Pass |
| Naïve failure count | 1/23 reachable | Fail |
| Selective executions | 21/21 success · 0 violations | Pass |
| Abstention precision | 50% | Fail |
| Update latency | 0.13664 s vs 0.10901 s FT | Fail |
The next move is a consequential world—not a wider interval
Confirmation seeds remain unopened. We will not widen uncertainty, move success tolerances or retune CSTR until it produces the desired failures.
The next prospective milestone is an embodied multi-step benchmark audit. Before repair is evaluated, it must independently show a competent state-based world model, clean nominal trigger calibration and missions whose success naturally changes with payload, friction or actuator authority.
That audit moves the project toward a visual flagship demonstration while preserving the same scientific rule: prove that the world is consequential before claiming that repair helped.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.