Completing a 300K-step model from scratch
A complete 125M-parameter language-model stack: tokenizer, data, model, long-run training, checkpoints, and evaluation.
Architecture, answers, failures
A chronology of the decisions that changed the project—from the original scratch model, through the first recurrent-workspace lift, through verified source-deleted mechanism tests, failed broad gates, and the first protected product pass from trained internal revision and coherent whole-trajectory commitment.
A complete 125M-parameter language-model stack: tokenizer, data, model, long-run training, checkpoints, and evaluation.
Recurring instability was isolated to long homogeneous data blocks and removed with deterministic domain-interleaved batches.
A 4,934-parameter module learned six primitives and composed all 36 withheld two-step transitions on a bounded protocol.
Sixteen learned slots, one tied cross-attentive cell, eight internal updates, and a gated residual back into the prompt are now public.
T2 initially improved small Qwen reasoning boards and produced a coherent AIME-only solution, establishing useful specialist computation.
Longer Qwen and SmolLM3 comparisons erased the broad lead. T2 retained hard-math evidence but lost enough general capability to stop scaling.
At 3,000 updates C2 became the strongest single arm, adding 58 GSM8K and 41 MATH answers over B1 while losing code and logic.
On 3,392 previously unopened problems, B1 and C2 solve 1,782 and 1,781. The workspace shifts capability toward math and science rather than raising the total.
A prompt-only domain router reached 1,790 unopened answers—just eight above B1. Subject identity is not a reliable proxy for per-problem expert advantage.
On 200 fresh verified prompts B1 solves 32, C2 solves 15, and a perfect selector solves only 38. The complementary-expert hypothesis is closed.
A verified 16M-target comparison selected the 1,000-update late-layer checkpoint; simply training longer degraded its broad capability.
Long math rollouts and exact answer verification produced 2,313 math and 1,800 science traces from fresh, evaluation-disjoint prompts.
A 400-update warm start with equal protective replay raised the matched five-domain macro from 45.4% to 48.9% while preserving code.
Expanded GPQA improves from 27/198 to 34/198 and HumanEval holds at 45/164; the remaining large boards are still running.
The active hard-negative wave targets 1,783 unsolved math and 2,296 unsolved science prompts, admitting only newly verified solutions.
A source-sealed factorized packet preserved 252 coherent worlds with zero support loss, zero false certificates, and a 16.90× aggregate storage advantage.
The tiny compiler won only two of five seeds; the frozen SmolLM2 residual successor missed exact-packet promotion by orders of magnitude. The next design must preserve token roles and source order.
A hierarchical structured compiler reached 96.09% exact train packets but fell to 8.59% lexical and zero renderer/composition exactness when forced to emit one irreversible parse.
A frozen 1,024-episode audit found that semantic templates and alignments survived shift; retaining only two coherent interpretations preserved the valid world in every episode.
ULC1 recovered 256/256 fresh episodes across train, lexical, renderer, and composition cohorts while matched top-1, particle, recurrence, and soft-mixture controls remained far behind.
EIC1 reaches 768/768 on normal and mapped-swap boards while an equal-FLOP duplicate control falls to 384/768, qualifying the typed identity owner.
NPL2 reaches 85.6104% late-query exactness across five source-disjoint seeds—exactly the typed oracle—while the strongest non-oracle arm reaches 3.9185%.
MZE1 learns all 75,272 finite transitions with 400 parameters, executes held programs through depth 32, and preserves the full NPL2 confirmation score.
EAL2 reads natural before/after evidence, identifies eight unseen laws per episode, deletes the demonstrations, and answers all 40,960 confirmation queries.
NCP1 maps raw variable-length commands into ordered episode-local operation pointers and stays exact under renaming and reversed command order across five boards.
OQB1 quotients repeated occurrences into an episode-local basis while a neural owner attaches values and queries; coherent reindexing remains exact and broken identity collapses.
SVE1 reads raw bytes into complete value events with no numeric-span scanner, confirming 30,720/30,720 event sequences and 40,960/40,960 answers.
SNL1 composes the frozen byte readers, identity bus, command compiler, and neural law synthesizer: 1,280/1,280 laws, 20,480/20,480 states, and 40,960/40,960 answers.
OPB1 later confirmed exact evidence-to-operation binding over five source-disjoint seeds, closing another controlled compiler interface.
OPB1 confirms 30,720 operation bindings, 20,480 terminal states, and 40,960 answers; source scrub and decoy controls fall to zero state and answer accuracy.
The Qwen3.5-4B QPT1 intervention raised macro accuracy from 55.630% to 62.588% and solved 69 more problems, but its code score fell from 30/40 to 26/40, closing the exact gate.
SAG1 retained 30/40 executable-code problems and improved broad performance over B1, yet a 12-answer MATH regression versus the continued control rejected promotion.
IDR1 uses one Qwen3.5-9B owner to draft and a later same-family owner to revise. On source-disjoint holdout, trained revision solves 625/1,279 versus 495 for the matched untrained second pass.
A model-owned commit policy raises source-disjoint holdout to 652/1,279 with exact order consistency. Its one-answer edge over an independent scorer is too small to attribute the gain specifically to antisymmetry.
Same-family draft, trained revision, and whole-trajectory commitment solve 383/538 at 75.815% macro—67 more problems and +8.552 points over the matched original second pass while retaining code.
Qwen3.6-35B-A3B reaches 143/256 versus 111 unchanged. Mixtral-8x22B reaches 448/1,023 versus 147 unchanged and 356 matched self-refinement; its release gate remains closed because code and baseline retention regress.
The Qwen3.5-9B draft/revision/commit system is confirmed at 383/538 and 75.815% macro on the protected product board. The distinct antisymmetric-scoring claim is closed because its matched control is only one answer behind. Sparse-MoE transfer now reaches 143/256 on a source-disjoint Qwen3.6-35B-A3B screen versus 111/256 unchanged. On Mixtral-8x22B it reaches 448/1,023 versus 147 unchanged and 356 matched self-refinement. That gain is real, but the release remains closed because executable code and baseline retention regress.