Thesis tested

Interactive repair is only economical if the system can identify the smallest correction with the largest downstream effect. These phases establish that acquisition policy before any learned world model is introduced.

The scientific question

A model can be incomplete in thousands of places, but not every unknown matters to the task at hand. The first hypothesis was that question selection should follow the behavioral frontier: reveal the law that blocks the most currently reachable scenarios.

The protocol strictly separated acquisition trajectories from held-out and stress trajectories. Question selection, verification and stopping could not inspect evaluation outcomes.

Early discipline

The first run missed its answer-reduction gate by 0.16 percentage point. The threshold was not moved. A pre-registered replication passed, and the pooled rule opened the next phase.

From exact contexts to structural questions

Exact state/action identity fragmented one reusable law into many sparse questions. A domain-neutral structural identity compressed those duplicates using public action schemas, known dependencies and qualitative relations to thresholds.

The gain grew with topology size: larger worlds created more places where choosing the right question mattered.

Phase 3A frozen confirmation
RegimeFrontier compressionAUC vs randomAnswer reductionGate
Small52.4%+0.040313.75%Pass
Medium45.1%+0.053321.80%Pass
Large39.1%+0.088426.13%Pass

One correction, multiple grounded instances

A confirmed concrete correction was then compiled into a typed transition motif and instantiated across compatible grounded components. Every instance still passed the existing verifier; ambiguous bindings abstained.

This is not open-ended schema induction. Public action schemas and typed role bindings were supplied. The result is narrower and more useful: validated knowledge can be reused without adding domain branches to the runtime.

Phase 3B motif-transfer confirmation
RegimeConcrete answersMotif answersReductionExact final worlds
Small6.005.872.2%30/30
Medium10.008.2317.7%30/30
Large16.009.8338.5%30/30

What this established—and what it did not

The evidence supports a bounded acquisition claim: structural question identity and verified motif transfer reduce teaching cost on fresh controlled worlds, with the benefit increasing across the tested topology range.

It does not yet say that a deployed neural dynamics model can be repaired. It establishes the acquisition and compilation machinery that later phases must integrate with learned prediction and planning.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.