Thesis tested

The LLM stops at semantic normalization. A deterministic compiler owns types, provenance and the approval boundary; a verifier owns behavioral admissibility; a human remains the authority for mutation.

The compiler boundary

A sentence never becomes a law directly. It becomes either a typed draft, a clarification request or a rejection. A draft retains source authority, character span and digest; any proposed change is marked approval-required.

Neither the language normalizer nor the compiler receives repository commit capability, hidden law identifiers or evaluation trajectories.

Execution path

Utterance → untrusted normalization → typed draft → local proposal → behavioral verification → human approval → versioned commit.

Controlled and synthetic language

Phase 5A validated the typed compiler on explicit action-scoped statements. Phase 5B then varied surface form through disjoint synthetic paraphrase families while preserving the same semantic and authority contracts.

Frozen compiler evaluations
EvaluationExact draftsApproval-ready proposalsIncorrect draftsNegative controls
Controlled language60/6057/600360/360 safe
Synthetic paraphrases240/240216/2400480/480 safe

Five independent operator proxies

Five English-speaking participants wrote corrections from grounded liquid-storage situations without seeing parameter names, intermediate representation or expected patches. Their pseudonymized sentences were later processed by a frozen OpenAI structured-output adapter under explicit consent.

The selected development snapshot produced 39 exact drafts from 40 correction dialogues. The remaining case asked for clarification; no incorrect draft or proposal was produced. All ten human challenge cases safely clarified or rejected.

Phase 5C frozen development snapshot
MetricResult
Exact correction drafts39/40 (97.5%)
Delay corrections20/20 exact
Rate corrections19/20 exact; one clarification
Human challenge safety10/10 no draft or proposal
Original-source provenance39/39 exact

Why this is not yet a language-generalization claim

The human cohort is small and development-only. Participants were operator proxies, not industrial experts. Proposal-review timing was not measured, and the planned ten-person confirmation was deliberately deferred when language ceased to be the core research bottleneck.

The adapter is therefore a credible qualitative interface and a safety-structured development result—not the primary quantitative evidence for world-model repair.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.