← Complete research archive
Pretraining & dataResearch record97 lines

R12 Factorized-Language Complete-Compiler Corpus Admission

The generator and tests were committed as e6d957e before any production seed existed. The schema-aware compiler evaluator was committed as 6c416cb, and the matched ordinary-parser control plus common Slurm contract were committed as 9b85463. Production seeds were chosen only afte…

R12_REFERENTIAL_LITERAL_POINTER_COMPILER_FACTORIZED_CORPUS_ADMISSION.mdOpen original Markdown ↗

R12 Factorized-Language Complete-Compiler Corpus Admission

Status: development matrix complete; absolute gates pass; attribution fails; confirmation sealed

Frozen identity

The generator and tests were committed as e6d957e before any production seed existed. The schema-aware compiler evaluator was committed as 6c416cb, and the matched ordinary-parser control plus common Slurm contract were committed as 9b85463. Production seeds were chosen only after the generator commit.

SplitRowsQuartetsSource SHA-256
train96,00024,000e6feb311c37f34a88ce7bda59ebb4f968c9ce3b4052cb5c0f6c2ef2e3fca44a8
development compositional2,048512e69fb70bddfb827a428c297352a72e45612ff3528a9fa107dec38c04189e1922
development lexical OOD2,04851240a059024770d3785ac27f7b02365d7741f631f7413aa5e69631efbc3af73dc0
confirmation8,1922,048e2bc25d8d95bb48c8d2915e6b966f3b96a01d3726177e6469f7824d0cf4b1a0f

The full local report SHA-256 is fd2c26580a1b164ad1095e0ad7940ffc2420c16f4882cbb0372c923c32bdc8f7. The Newton/GitHub development report removes both the confirmation seed and the confirmation artifact entry; its SHA-256 is d481114232e438294bd1ea7f5b739f6068c2bf10fe02c1ee3c216c2e56aa3be3. Confirmation JSONL remains local only and is absent from Newton.

Accounting

QuantityValue
Total rows108,288
Source tokens10,858,878
Source UTF-8 bytes38,214,563
Model-owned pointer labels1,082,880
Teacher/model calls during generation0
Checkpoint reads during generation0
Production evaluation answer reads0

All rows contain ten nonempty source-owned spans: three initial entities, two operation kind/entity/literal triples, and one query position. Independent pop/insert and adjacent-swap CPU executors agree on every row.

Structural gates

Every frozen CPU gate passes:

  • all IDs are unique;
  • every semantic quartet preserves canonical/paraphrase behavior and separates both order and binding twins;
  • canonical/order/binding token bags are identical;
  • train covers every known atomic language factor;
  • compositional development uses only known atoms in unseen combinations;
  • lexical OOD direction words are absent from all known-lexicon strata;
  • exact prompt, word-13-gram, nonce-name, and full factor-combination overlap are zero across every split pair;
  • token-bag shortcut accuracy is exactly chance at 1/3; absolute-position and source-length shortcut ceilings remain below the admitted chance tolerance.

The factor cross-product independently varies intro frame, list style, operation ordinal vocabulary, operation frame, argument order, direction lexicon, distractor frame/location, query frame, and punctuation/case style. A split-disjoint neutral nonce anchor prevents shared generic prose from creating cross-split 13-grams; it is sampled independently of every semantic label and is included in the name-overlap audit.

Matched development arms

Both initial arms use immutable raw Shohin 300k, 96,000 examples, one epoch, 1,517 optimizer updates, seed 2026071810, source token IDs plus length mask as their only inference inputs, and the same evaluator.

ArmAdapter parametersTotal parametersNewton job
v1.3 parameter islands8,658,701133,740,365693048 on evc25
favorable ordinary bidirectional tagger8,607,886133,689,550693049 on evc28

The adapter-budget difference is 50,815 parameters, or 0.587% of the islands adapter. The ordinary control directly labels source tokens with five bidirectional Transformer encoder layers and pools direction classes at the predicted kind spans. It receives no learned program slots or parameter islands. Both jobs completed cleanly. Parameter islands and the ordinary tagger each reach 100% exact programs and full pointers on compositional development, so the absolute compiler gates pass but the required +5-point islands advantage is 0. The free-slot and structured controls reach 98.242% and 100% exact programs; shuffled-label islands reach 0.146%. The committed assessor returns retain_as_conventional_compiler_baseline_confirmation_sealed. Full results and oracle localization are in R12_REFERENTIAL_LITERAL_POINTER_COMPILER_FACTORIZED_DEVELOPMENT_RESULT.md.

Decision boundary

V1.3 passes every frozen primary development gate but ties the favorable ordinary parser at 100% exact semantic programs. The attribution gate therefore fails and confirmation must remain sealed. Lexical OOD is diagnostic and must never be pooled into the compositional score. No result on this board alone establishes state transition, halting, autonomous rollout, or native reasoning.