Thesis tested

A bounded correction can preserve strong predictive recovery, but selective competence is not demonstrated unless approximate knowledge creates enough consequential mistakes for asking and abstention to be falsifiable.

Replacing an exact answer with a bounded one

Earlier CSTR evidence gave the repair system the exact hidden actuator scale. Phase 6E prospectively replaced that easy answer with a width-0.20 interval and at most one conditional width-0.10 refinement.

Nine scale hypotheses were propagated through H25 prediction and H60 planning. A plan could execute only if every member satisfied the existing target and temperature certificate. Otherwise the system asked once more or selected a neutral fallback.

The approximate patch retained useful prediction

All eight changed streams triggered at observation 8. Every coarse and refined interval contained the hidden scale, passed the visible-window verifier and produced an evidence-linked immutable version.

The final midpoint recovered 95.49% of lost H25 competence. Fine-interval coverage reached 99.78% over 64 fresh stress episodes while mean normalized uncertainty width fell 45.29%.

Development H25 prediction
StrategyMean H25 errorEvidence status
Frozen0.052701Development comparator
Full fine-tuning0.018805Development comparator
Rank-2 adapter0.027443Development comparator
Coarse midpoint0.005899Approximate answer
Final midpoint0.003443After conditional refinement
Exact reference0.000587Evaluation only

Why 21 successful executions still did not pass

The exact evaluation reference could reach 23 of 32 tasks. The selective policy executed 21 and succeeded on all 21 with no safety violation.

But the naïve coarse midpoint already succeeded on 22 of those 23 reachable tasks. There was only one naïve failure to avoid, below the preregistered minimum of three. One avoided mistake is positive descriptive evidence, not task-level dominance or a validated abstention policy.

The hard boundary

A perfect execution rate among selected tasks cannot validate selection when the underlying workload contains almost no failures.

Four failed gates remain part of the result

Two of four unchanged controls crossed the frozen detector threshold. The harness did not issue answers on those known controls, but a deployed system would not know the regime label, so they remain false positives.

Reference-reachable abstention precision was 50%, below the registered 60%. Every world requested the second answer, and two interval compile-and-verify operations took a median 0.13664 seconds versus 0.10901 seconds for the small full-FT baseline.

Phase 6E development conjunction
Gate groupObservedDecision
Nominal controls2/4 cleanFail
Prediction and interval calibration95.49% recovery · 99.78% coveragePass
Naïve failure count1/23 reachableFail
Selective executions21/21 success · 0 violationsPass
Abstention precision50%Fail
Update latency0.13664 s vs 0.10901 s FTFail

The next move is a consequential world—not a wider interval

Confirmation seeds remain unopened. We will not widen uncertainty, move success tolerances or retune CSTR until it produces the desired failures.

The next prospective milestone is an embodied multi-step benchmark audit. Before repair is evaluated, it must independently show a competent state-based world model, clean nominal trigger calibration and missions whose success naturally changes with payload, friction or actuator authority.

That audit moves the project toward a visual flagship demonstration while preserving the same scientific rule: prove that the world is consequential before claiming that repair helped.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.