Thesis tested

The same local action-effect repair works in a recognized nonlinear continuous-control environment without reading simulator coefficients. It restores continuous prediction and plan quality, but not yet task-level reliability dominance.

Why four-tank changes the evidence

The coupled four-tank benchmark introduces nonlinear continuous dynamics, two causal action channels and longer plan-once horizons. The learned backbone has 17,924 parameters and is trained for recursive competence rather than exact symbolic execution.

A hidden actuator degradation changes one action channel. Acquisition may read only visible trajectories and public affected-state footprints—not simulator equations, hidden coefficients or evaluation tasks.

Failures that shaped the protocol

The first nonlinear workload failed because the chosen shift barely changed consequential behavior. A causal-sensitivity protocol then failed because the shift signal was smaller than nominal model error. A stronger rollout-trained backbone fixed competence, but reactive control masked model differences.

Plan-once evaluation exposed the causal gap. A prospective 1% non-inferiority margin replaced an over-precise cost ordering that had failed by 0.0036%. R4 development then passed on both channels: 8/8 exact repairs, 31/32 hybrid successes versus 0/32 frozen, and 32/32 paired non-inferiority.

Two independently frozen evaluation partitions

Across 24 hidden changes, 96 held-out prediction episodes and 72 plan-once tasks, every change triggered exactly one correction and exact local commit. Every case recovered the registered H25 prediction gap and mean planning gap.

Descriptive aggregate of B2-R4 and B2-R5; not a pooled preregistered gate
EvidenceResult
Actionable trigger + exact commit24/24
H25 prediction-gap recovery24/24
Mean planning-gap recovery24/24
Within 1% of better learned update72/72 tasks
Incorrect / cross-channel commits0
Invariant violations / leaks0

Prediction, planning and update cost

On the final replication, hybrid final H25 error was 0.000180 versus 0.001046 for full fine-tuning. Mean planning cost fell from 0.002989 frozen to 0.000465 hybrid. The 178–179 byte patch updated in 2.51 ms median versus 100.27 ms for full fine-tuning.

Final replication, selected continuous metrics
MetricFrozenFull FTHybrid
Final H25 error0.0010460.000180
H25 AUC0.0027340.001264
Mean planning cost0.0029890.000465
Median update latency0100.27 ms2.51 ms

The one gate that failed twice

B2-R4 and the feasibility-aware B2-R5 replication each passed nine of ten gates. The final replication covered 17 of 20 tasks that the shifted reference could reach—85% versus the frozen 90% requirement—and achieved 17 total successes versus 18 for the adapter.

This blocks any claim of task-level dominance or industrial reliability. The defensible claim is restoration of causal predictions and mean plan quality, materially faster than parameter updating, with locality and integrity preserved.

No more four-tank tuning

Further threshold, target, tolerance, seed, planner or backbone tuning was explicitly closed as a rabbit hole.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.