A repair system must know when not to repair. Phase 6B-T closes that trigger contract; Phase 6C-A then shows the mechanism can execute at million-parameter scale, while honestly failing an absolute-competence gate at the lower tier.
The exact trigger contract
An exposed action becomes actionable only when visible evidence contradicts both the public symbolic version and the deployed learned predictor, and typed edit provenance is available. Exposure alone is insufficient.
Invalid provenance abstains before asking. Unchanged controls remain silent. Every accepted correction creates a replayable question → answer → proposal → verification → commit chain.
| Cohort | Actionable | Public-only | Exact commits | Hybrid H5 targets | Full FT |
|---|---|---|---|---|---|
| 8 nodes / 12 relations | 29/30 | 1/30 | 29/29 | 30/30 | 14/30 |
| 16 nodes / 24 relations | 29/30 | 1/30 | 29/29 | 30/30 | 15/30 |
Minimal intervention is part of correctness
Two fresh worlds exposed a public-model contradiction while the deployed learned predictor already matched observations. Both were labeled public-only gaps, consumed zero answers and created zero commits.
That distinction prevents the symbolic layer from overwriting a competent learned representation merely because a public rule is stale.
The million-parameter bridge
The same repair function was executed with 1.009M and 10.069M active parameters. Repair update latency scaled with the patch, not the full backbone: 87× and 182× faster than full fine-tuning, with 1.51% and 0.23% inference overhead.
| Tier | Active parameters | Repair vs full FT | Inference overhead | Outcome |
|---|---|---|---|---|
| Primary | 1.009M | 87× faster | 1.51% | Failed absolute competence |
| Systems | 10.069M | 182× faster | 0.23% | All tier gates passed; n=4 |
Why 6C-A is still a formal failure
The primary ≥1M tier recovered every relative target but finished below frozen absolute bars: one-step 0.849 versus 0.900, H5 0.631 versus 0.700, and H8 0.559 versus 0.650. The overall conjunction therefore failed 18/19.
A capability-matched follow-up trained a native 1.013M backbone. One-step and finite-budget planning passed, but old-regime recursive competence did not. The stop rule prevented post-shift repair curves from being run, so that study is a backbone failure—not evidence against or for repair.
The cost curve is real and useful. It must always be paired with the failed ≥1M absolute-competence gate.
What the scale result establishes
The repair mechanism is not intrinsically tied to a 23k-parameter toy model. It has executed on backbones above one and ten million active parameters with shrinking relative inference overhead and a widening update-time advantage.
But scale of execution is not scale of intelligence. A capable large backbone remains a prerequisite, and the next evidence had to move to recognized external nonlinear environments.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.