R12 S4 Pointer-Anchored Event Tape Repair
Status
FORMALLY REJECTED. This was frozen as a zero-fit public-development repair after S4 v1 treatment evaluation and before any repaired score. No model weight, corpus row, optimizer, update count, seed, threshold, or confirmation input changed.
Failure diagnosis
S4 v1 predicts exact event count on 2,048/2,048 held-out rows and is fully correct on every valid tape, but strict decoding invalidates 116 rows: 66 initial-roster cardinality errors, 47 event-role component-cardinality errors, and three entity-identity errors. Gold initial/query boundaries lift exact programs from 94.336% to 97.217%, every depth at least 96.471%. Shuffled supervision remains zero. The remaining miss is hard argmax span fragmentation, not count or event semantics.
Sole repair
Build a structural lexicon from the admitted training split only:
- exact known direction token patterns and class;
- exact amount and query-literal token patterns and value;
- the set of training entity-span token widths.
At inference:
- Each of the three schema-fixed initial-role global pointer anchors expands to the highest-scoring training-width window that contains it.
- An event exists only when an exact direction pattern contains a token whose model argmax role is
event.kind. These anchored patterns, ordered by source position, define event count. - Inside each adjacent anchored-event interval, exact occurrences of the three model-predicted
initial token sequences compete under
event.entityrole score; exact known literals compete underevent.literalscore. - The query-role global anchor expands only to an exact training query-literal pattern.
- Any missing, overlapping, ambiguous, or duplicate structural selection is invalid. No gold depth, count, span, entity, event, state, or answer enters inference.
This is a deterministic structured decoder over model logits, equivalent to lexicon-constrained semantic parsing. It is not a new reasoning primitive.
Frozen gates
The original S4 gates remain unchanged: at least 98% exact count overall and 95% every depth; at least 95% exact programs overall and 90% every held-out depth; at least 95% answers overall and 90% at depth eight; gold-count rescue below two points; shuffled exact programs at most 40%; locked S3 gold sanity; total parameters below 150M; zero confirmation access.
V1.1 may run once on the same public development rows after source, lexicon builder, evaluator, and this repair are committed. A pass authorizes only a separately frozen fresh confirmation protocol.
Pre-evaluation builder receipt
The first post-commit training-only lexicon build failed closed at SHA-256
f487d1cb98bebd84137c1b0b7839e2241603cc4f920f4f1a09205f502e9015e6. Its sole failed gate
incorrectly required one entity token width. The admitted training spans contain 3,061 width-four,
130,847 width-five, and 10,092 width-six occurrences because contextual BPE boundaries vary. The
frozen repair above already specified the set of training entity-span widths, and the decoder was
implemented to accept that set. Before any development score, the builder gate is therefore
repaired to require a nonempty bounded width set and exact accounting of all 144,000 training intro
spans. The failed receipt is retained as s4_structural_lexicon_v1.failed_one_width.json.
Result
Jobs 693160 and 693161 completed cleanly. Treatment retains 2048/2048 exact event counts but
falls to 25/2048 exact programs and 300/2048 answers; shuffled remains 0/2048 exact programs.
Training-width expansion selects the wrong 4/5/6-token roster boundaries and creates 1,176
event_entity failures. Frozen assessment SHA-256
fd0479b0737af49313b0cebf1863c4826c21de336f51e240ece3e4d60d11d587 records
reject_s4_v1_1_public_development. Do not repair or rescore v1.1 on this board. The lawful next
test is a newly preregistered event-relative start/end pointer architecture on fresh development
data.