← Complete research archive
Theory & no-go resultsClosed / no-go544 lines

R12 Endogenous Typed Theory Reactor Preregistration

Frozen successor protocol. The trainable architecture, causal episode lifecycle, composite objective, optimizer/checkpoint contract, common transaction schema, three exact ontology boards, seven genuine structural/semantic variants, and the complete 2,688-execution primary matrix…

R12_ENDOGENOUS_TYPED_THEORY_REACTOR_PREREG.mdOpen original Markdown ↗

R12 Endogenous Typed Theory Reactor Preregistration

Status

Frozen successor protocol. The trainable architecture, causal episode lifecycle, composite objective, optimizer/checkpoint contract, common transaction schema, three exact ontology boards, seven genuine structural/semantic variants, and the complete 2,688-execution primary matrix are implemented. No continuation pretraining, post-training, or capability claim exists.

The user pretraining hold remains active. The present phase is architecture construction and qualification only. Synthetic resource profiling may execute forward, backward, and bounded optimizer mechanics against an immutable copy of the step-300k checkpoint, but it may not read pretraining shards, write model state, or constitute continuation pretraining.

Objective

Test whether one actual-Shohin, raw-token, source-deleted architecture can infer and execute a previously unseen typed theory rather than operate inside a supplied finite-machine ontology.

The candidate must induce an anonymous episode object

Theta = (Q, tau, R, F, Gamma, Omega)

where:

  • Q binds physical mentions to episode-local objects;
  • tau assigns latent object types;
  • R defines relation symbols, arities, argument roles, and initial facts;
  • F defines operator preconditions and graph effects;
  • Gamma defines sequential, synchronous, saturation, branching, and halt semantics; and
  • Omega defines observers available to a late query.

Behavioral Identifiability Gate

For evidence D, bounded theory class H, and admissible late challenges Q, define:

V(D) = {Theta in H : Theta satisfies D} / isomorphism

Two theories are behaviorally equivalent when they answer every admissible late challenge identically. Exact deterministic reasoning is identifiable if and only if the behavioral quotient of V(D) has one class.

Every scored episode must receive an independent exact version-space receipt:

  • singleton behavioral class: candidate must commit and answer;
  • multiple behavioral classes: candidate must abstain;
  • empty version space: candidate must reject; and
  • coherent alternate singleton: candidate must commit to the alternate world's behavior, not reject it.

Three Ontologies

Three leave-one-ontology-out folds are mandatory.

OntologyHidden structureRequired execution
Horn closureobjects, typed predicates, asymmetric rolesmonotone least fixed point
Typed term rewritingconstructors, variables, ordered child rolesdeletion, replacement, branching normal form
Guarded resource processplaces, resource types, multiplicitiesguarded consume/produce, sequence, deadlock/halt

These families may share only the generic typed-transaction substrate. No family identifier, family head, domain opcode, or host semantic callback may enter the candidate.

Architecture Implementation

train/endogenous_typed_theory_reactor.py implements the checkpoint-compatible architecture intended for later continued pretraining:

  1. an endogenous compiler cross-attends anonymous object slots to raw-token Shohin residuals;
  2. the compiler emits only bounded categorical value codes, type probabilities, a sparse relation ledger, activity, root, commit, and halt state; deployed packets reject continuous values and relation counts above the frozen cap. Production geometry uses 64 slots, 16 relation roles, 256 categorical symbols, and reified ordered hyperedge/value-byte nodes;
  3. a shared recurrent reactor cross-attends a separate post-seal command stream, reads exact directed endpoint identity through an edge-aware typed relation-message bus, emits eight structural/terminal choices plus a distinct REJECT, and applies differentiable graph updates;
  4. an exact-forward straight-through path supports discrete transactions while exposing pre-discretization probabilities for corrective training gradients; and
  5. a separate causally masked query reader consumes every declared typed-state field, including endpoint-aware incoming/outgoing neighbor messages, and query residuals without seeing future query tokens.

The corrected architecture adds 67,697,771 parameters: 21,466,377 in the compiler, 29,757,217 in the reactor, and 16,474,177 in the query reader. With the immutable 125,081,664-parameter Shohin base, the complete system contains 192,779,435 parameters and leaves 7,220,565 below the 200M ceiling.

The actual protected checkpoint hash matches, step 300,000 loads strictly with zero missing or unexpected tensors, and the wrapper parameter receipt passes. Focused tests require nonzero first-batch gradients, autoregressive prefix invariance, causal use of every state field, commit freezing, categorical deployed packets, exact sparse transactions, parameter accounting, and independent base freezing.

train/ettr_episode.py, train/ettr_objectives.py, train/ettr_data_contract.py, train/ettr_optimization.py, train/ettr_checkpoint.py, and train/ettr_train_step.py now close the previously missing continuation contract:

  • every row has independent WORLD -> COMMAND -> QUERY streams and explicit reset boundaries;
  • language-model targets cannot cross a segment boundary;
  • initial-packet, free-running terminal-packet, transaction, initial/terminal equivariance, commit/halt, sparsity, and anti-bypass supervision share one device-resident composite objective;
  • training snapshots and every batch are immutable/hash-bound, with opaque content-hash episode IDs and no live-writer or family-routing field;
  • protected-base and added-architecture Muon/AdamW groups are disjoint, with an embedded WSD update cursor; and
  • atomic checkpoints bind model, optimizer, schedule, RNG, data cursor, source manifest, and protected-base provenance, and are admitted only at an optimizer/between-episode boundary that can be resumed exactly.

The complete ETTR/cross-ontology architecture and custody inventory passes 165/165. Terminal-packet supervision is connected through the recurrent reactor to the compiler, while joint normalization preserves the original packet-family weight scale. A degree-preserving edge-swap falsifier proves that the reactor and query reader distinguish graphs with identical per-slot relation counts but different endpoints. This establishes continuation-contract mechanics, not reasoning capability. A healthy-node BF16 H100 memory/throughput profile remains mandatory. Continued pretraining remains under the explicit user hold.

G0 Horn Mechanics

The first offline board is implemented in pipeline/cross_ontology_horn_board.py, with a separately implemented exact oracle in pipeline/audit_cross_ontology_horn_board.py.

  • 20 three-rule theories occupy 20 behavioral equivalence classes.
  • The challenge space contains 27 typed atoms and 378 initial states.
  • Independent closure engines agree on all 7,560 theory/challenge pairs.
  • Exact evidence yields singleton, ambiguous, contradictory, and coherent alternate dispositions for every target theory.
  • Four opaque renderers change source bytes while preserving semantics.
  • Reference packets use only the generic transaction schema.

Full mechanics disposition: R12_ETTR_G0_HORN_BOARD_RESULT.md.

The typed-rewrite board adds 15 behaviorally distinct two-rule theories, with eight exact rule combinations held out while every primitive rule remains seen. Two independent normal-form engines agree on all 960 theory/term pairs, including repeated-variable, ordered-child, and nonconfluent cases.

The guarded-resource board adds 60 behaviorally distinct three-operator theories, 81 typed markings, and 36 unseen length-two/three programs. Two independent engines agree on all 174,960 held-out executions, including 13,362 normal halts and 161,598 deadlocks.

Across all three boards there are 183,480 exact independent-oracle comparisons and 352 exact identifiability episodes.

Frozen Primary Matrix

pipeline/cross_ontology_qualification_matrix.py materializes the exact preregistered geometry:

  • 3 leave-one-ontology-out folds;
  • 8 behaviorally distinct theories per fold;
  • 7 genuine variants per theory;
  • 16 aligned late challenges per variant;
  • 168 source worlds, 384 canonical late challenges, and 2,688 primary executions;
  • 1,472 declared exact-invariance executions;
  • 750 exact outcome/directive separations from semantic or identifiability twins;
  • 384 required abstentions;
  • zero candidate-visible family-label leaks; and
  • 2,688 unique row hashes and 24 disjoint theory hashes.

Matrix payload SHA-256: d1904b54a0fab8e59cfcb0b0dd464f5c8778e5b828907028ec8614aeae76d5d5. Rules, alignments, expected outputs, and exact oracles remain assessor-side.

Process Custody Implementation

The architecture now has a non-pickle, allowlisted safetensors state wire and four detached process surfaces:

  1. run_ettr_world_compiler.py receives world tokens, compiler weights, and the hash-bound Shohin checkpoint, then emits only immutable typed state.
  2. run_ettr_state_executor.py receives only typed state, immutable post-seal command tokens, reactor weights, the hash-bound Shohin checkpoint, geometry, and a step budget. It encodes the command through the protected base before invoking the generic reactor.
  3. run_ettr_late_query.py receives only terminal state, late-query tokens, query-reader weights, and the hash-bound Shohin checkpoint.
  4. run_cross_ontology_assessor.py is model-free and receives only immutable candidate and independently generated expected outputs.

The serial custody test destroys each previous stage directory before the next process starts and creates assessor expectations only after candidate exit. This establishes process mechanics, not hostile-kernel isolation or capability.

Architecture

Raw-token compiler

  • Load and execute the protected Shohin checkpoint.
  • Accept only tokenizer output and masks from raw source text.
  • Use one shared compiler/adaptor across all ontologies and renderers.
  • Emit an immutable anonymous typed-theory packet.
  • Do not receive exact spans, numbers, entity equality, relation roles, family labels, program graphs, schedules, answers, or assessor products.

Generic reactor

One recurrent controller emits only:

ALLOC WRITE CLEAR LINK UNLINK SET_ROOT COMMIT HALT REJECT

A rule-blind committer may enforce bounds, pointer validity, type shape, capacity, and transaction atomicity. It may not match a semantic rule, choose a redex, compute closure, perform arithmetic, repair a transaction, select an answer, or retry after assessor feedback.

Late-query reader

The reader receives only the committed terminal object and raw late-query tokens. It cannot access source tokens, compiler residuals, KV state, parser state, execution trajectory, or assessor data.

Four-Process Custody

  1. Compiler process: reads source and writes an immutable packet.
  2. Executor process: starts fresh, receives only packet and command stream, commits terminal state, then exits.
  3. Query process: starts fresh, receives terminal state and raw late query, writes an answer or abstention, then exits.
  4. Assessor process: starts only after all candidate processes exit and uses an independent implementation.

Packets may contain no source-derived digest, source offsets, raw names, residuals, hidden caches, answer labels, or executable host callbacks. Post-seal source poisoning must be bit-invariant.

Smallest Decisive Board

  • 3 leave-one-ontology-out folds;
  • 8 independently generated held-out theories per fold;
  • 7 versions per theory: base, alpha/reorder, alias split, relation reification, type twin, execution-semantics twin, and ambiguity-deleted twin;
  • 16 independently generated post-seal challenges per version;
  • 2,688 primary scored executions;
  • at most 6 objects, 3 inferred types, 3 relations of arity at most 3, 3 opaque operators, depth 8, and branch width 2;
  • at least 4 renderers, with one fully held out;
  • disjoint canonical and isomorphism hashes across splits; and
  • no abstract operator/effect/control program overlap between fitting and the held-out ontology.

Hybrid confirmation must include at least:

  • arithmetic index selecting a rewrite location;
  • relation result selecting a resource operator; and
  • resource state controlling a Horn query.

The frozen hybrid receipt implements exactly those three couplings with 16 cases each. Two independent executors agree on 96/96 factual and counterfactual executions; all 48 interventions change the causal signal and the final output; candidate-visible payloads contain no family labels. Payload SHA-256: d155f868494f9379b214028c8d7475cc2cde08192c9b3a5bbdea5a73b29f98e2. This is an offline mechanics receipt, not a learned score.

Matched Controls

  1. actual Shohin trunk;
  2. zeroed trunk residuals;
  3. parameter-permuted trunk;
  4. example-swapped frozen trunk;
  5. equal-parameter generic recurrent classifier;
  6. fixed-ontology typed reactor;
  7. family-routed executors with matched aggregate parameters;
  8. random-label control;
  9. ambiguous, contradictory, and coherent-alternate evidence; and
  10. independent type, role, effect, control-semantic, state, order, and query transplants.

Every learned arm must share update count, data access, initialization lineage, and parameter budget where structurally possible.

Promotion Gates

  • actual checkpoint loaded and hash-verified;
  • fewer than 200,000,000 unique participating parameters;
  • actual Shohin treatment beats every zeroed/randomized/swapped-trunk control;
  • raw-token end-to-end custody passes with no semantic host parser;
  • 100% independent oracle agreement and packet-schema validation;
  • at least 95% exactness on identifiable cases in every held-out ontology;
  • 100% abstention on behaviorally ambiguous cases;
  • 100% rejection on contradictory cases;
  • 100% coherent-alternate-world behavior;
  • 100% alpha/reorder/alias/reification invariance after canonical alignment;
  • at least 95% execution-twin and noncongruent-twin separation;
  • at least 20 points over every qualified matched learned control;
  • all three leave-one-ontology-out folds pass individually; and
  • hybrid confirmation reaches at least 85%.

A pass establishes bounded cross-ontology typed-theory induction. It does not establish unrestricted natural-language reasoning. Natural-language and public benchmark promotion remains governed by G4 in R12_GENERAL_REASONING_GATE.md.

Architecture-Phase Amendment: Factorial Interchange

Effective: 2026-07-26 EDT. Source: commit 5771c64.

This amendment qualifies training mechanics only. It does not authorize continuation pretraining, post-training, capability attribution, or a native reasoning claim.

Every causal training unit is an immutable 2x2 rectangle over two semantic WORLD factors and two semantic COMMAND factors. Equivalent factors must use different raw token renderings. WORLD-equivalent rows must share every initial packet target field and mask, while distinct WORLD factors must produce different initial packet targets. Terminal support geometry is identical within a rectangle, and changing either WORLD or COMMAND must change the terminal target at both settings of the orthogonal factor.

Intervention predictions may not replay their factual target row. The WORLD arm composes a packet compiled from a distinct rendering of the required WORLD factor with a distinct COMMAND row carrying the required COMMAND semantics. The COMMAND arm performs the orthogonal interchange. For both arms:

  • packet source row, command source row, and target row are explicit;
  • source rows differ in raw bytes from the target row;
  • targets are gathered only from immutable factual rectangle rows;
  • terminal packet and complete transaction trace are supervised;
  • WORLD and COMMAND losses and receipts remain separate; and
  • hard-forward gradients must reach the compiler and command path in isolated tests.

The continuation boundary must independently replay each labeled generic transaction from the initial packet and reproduce the complete terminal packet. It must reject contradictory values, types, relations, roots, activity, edge capacity, commit/halt state, or disposition. Initial status is the compiler's open reset. A right-padded row is valid only if its final supervised step has committed or halted, because deployed recurrence executes the fixed step width and relies on terminal state to freeze later mutation.

Decisive state checks run in evaluation mode with hard=True and must pass validate_deployed_state for initial, factual terminal, WORLD-intervention terminal, and COMMAND-intervention terminal packets.

The systems profile is schema v3. One exact-source H100 run must execute factual episodes, both intervention arms, the full composite objective, backward, and Muon/AdamW update in matched eager and compiled arms. It may read only the hash-bound protected checkpoint and synthetic rectangles, may write only one isolated JSON receipt, and may never read shards or write model state.

Architecture-Phase Amendment: Sealed Packet Sufficiency and Consumer Gate

Effective: 2026-07-26 EDT. Source: commit cf56818.

This amendment freezes the last pre-training architecture controls. It does not authorize continuation pretraining or claim learned capability.

The continuation manifest must bind the complete train and validation populations independently: canonical context identities, canonical full-batch payload digests, row and context cardinalities, split payload hashes, combined dataset hash, and packet-sufficiency receipt. Admissions must use sealed independent copies that cannot be changed by mutating visible manifest or index fields. Validation contexts and payloads are never train-admissible.

Every deployed terminal-packet field must receive nonzero factual or intervention supervision across the admitted training population. A support mask may describe objective support but may not waive deployed-state sufficiency. Optimizer ownership is checked against live parameter groups; any exception after the first optimizer mutation permanently poisons that optimizer, including after wrapper reconstruction or serialization.

The late-query consumer gate uses an identical causal query prefix for all four corners of a WORLD x COMMAND rectangle. Every WORLD and COMMAND edge must change the factual next-token target. Intervention execution receives target row indices, never answer labels. Correct and foil logits are read through the actual source-deleted query reader from distinct terminal states under the same query prefix.

Before any later reasoning promotion, the learned architecture must be evaluated against this frozen control matrix:

ControlRequired constructionFailure diagnosed
Query-onlyCanonical empty terminal packet with the original queryFrozen language path can answer without state
Zero-readerRemove the state-derived query residualReader contribution is unnecessary
Shuffled-statePermute terminal packets within matched query strata with no fixed pointsPacket identity is not causally consumed
Wrong-state factorial foilsSubstitute each orthogonal WORLD/COMMAND corner under the identical prefixFactorial interchange is not compositionally bound
Wrong-queryHold packet fixed and use a semantically different matched queryPacket stores only a single answer shortcut
Target derangementDerange immutable factual targets after all inputs are sealedObjective or assessor leaks labels
Query twinsAt least two semantic questions and two paraphrases per sealed stateReader cannot reuse one state for independent queries
Packet sufficiency ablationRemove one deployed state-field supervision family at a timeA declared packet field is decorative
Physical source deletionCompiler source artifacts are deleted before executor/query stagesHidden source or residual channel remains
Autonomous codebook readoutGenerate the categorical answer token without teacher-forced divergent prefixOne-token binding does not survive deployment

All controls must share examples, update count, parameter budget, and initialization lineage where structurally possible. The treatment must show a positive causal packet effect and must beat query-only, zero-reader, shuffled-state, wrong-state, and deranged-target controls on each held-out ontology, not merely in aggregate. Query twins must both answer correctly from one sealed state, and physical source deletion must be bit-invariant.

The exact hardware receipt is schema shohin-ettr-h100-profile-v5. Full-objective timing and memory are measured before separate eager-BF16 isolated gradient attribution. Compiler, reactor core, command projection, and query reader groups are disjoint. Encoded work is exactly WORLD + 2*COMMAND + 3*QUERY; no synthetic arm may be silently omitted from throughput accounting.

Executable Learned-Qualification Harness

The assessor-side inference controls are implemented in train/ettr_qualification.py under schema shohin-ettr-causal-qualification-v1. This implementation does not authorize training and does not claim a learned result.

An immutable qualification batch binds every deployed terminal-state tensor, query token and mask, autonomous read index, factual target, semantic-factor identity, paraphrase identity, and control permutation into one SHA-256. Packet identities separately hash every field of each deployed state row. Every state permutation is a no-fixed-point matched derangement. Wrong-WORLD and wrong-COMMAND controls change exactly one factorial identity; query twins hold state and paraphrase identity fixed while changing query semantics and factual answer.

The harness zeros and masks every token after the read position before any forward, passes targets=None, seals all read-position logits, and only then scores factual and deranged labels. It executes treatment, query-only, zero-reader, shuffled-state, wrong-WORLD, wrong-COMMAND, and matched wrong-query/query-twin arms. The scorer rejects a different batch receipt or any post-forward mutation of factual, query-twin, or deranged targets.

Packet-sufficiency ablation remains an equal-budget training comparison: each declared state-field supervision family must be removed in a separate arm, not simulated by zeroing a trained packet at evaluation. Physical source deletion remains process-enforced by the four-process custody runner. The complete integrated architecture/custody inventory, including 19 hostile harness tests and the direct staged qualification board, passes 240/240 in 156.18 seconds.

The sealed readout receipt binds every logit and assessor target to both the exact batch SHA-256 and an exact model-state SHA-256. Every candidate input is cloned and hash-checked around its forward; model and batch receipts must be identical before and after the complete arm sequence. State-control donors must change the factual target, and state-control readouts are scored against both the original negative-control label and the donor's correct counterfactual label. Arbitrary changed output is not a causal pass.

The supported evaluator is atomic and never returns logits. A separately frozen semantic-role manifest and exact model receipt must match preregistered SHA-256 values. The model receipt includes every named child module's class implementation; subclasses, instance method overrides, and hooks are inadmissible. Arm execution uses a secret-random permutation and returns its receipt. Independent final review found no supported-public-API P0/P1.

Frozen Three-Stage Qualification Board

pipeline/ettr_factorial_qualification_board.py is the direct learned- qualification source of truth. It does not relabel the older hybrid challenge rows. It constructs Horn closure, typed rewriting, and guarded-resource episodes with three genuinely distinct stages:

  1. WORLD establishes only the initial typed state;
  2. COMMAND arrives only after that state is sealed and transforms it; and
  3. QUERY arrives only after command execution and asks one of two independent semantic questions through one of two paraphrases.

Each ontology is an exact 2x2 WORLD x COMMAND rectangle. The frozen board has 12 terminal packets and 48 autonomous one-token query rows. Independent oracles agree on all 12 executions. Every semantic/paraphrase cell changes on both WORLD edges and both COMMAND edges, yielding 24 answer-changing WORLD edges and 24 answer-changing COMMAND edges. Every terminal packet supports two distinct factual query targets.

The payload SHA-256 is 18686ff7f0476b5a4432830f2a301f693833cf867656d3997a010cf17bb0149a. WORLD, COMMAND, QUERY, and assessor packages have separate immutable receipts. Candidate-visible packages omit answers, oracle outputs, ontology labels, and assessor targets. Tests physically delete earlier packages before later-stage materialization and reject cross-stage leakage.

train/ettr_factorial_qualification.py binds externally produced hard terminal states to the board, model, packet, world-factor, and command-factor receipts. It tokenizes only answer-free late-query prefixes and emits the production ETTRQualificationManifest and ETTRQualificationBatch. The shuffled control is a no-fixed-point four-cycle around each factorial rectangle, so every donor changes exactly one factor at each edge rather than using a diagonal that could preserve an XOR-like answer. Wrong-WORLD, wrong-COMMAND, query-twin, and target-derangement controls are exact.

The supported admission path requires four independently preregistered identities: complete model, execution manifest, compiler receipt, and executor receipt. The execution manifest binds the board and stage-package receipts, pretokenized WORLD/COMMAND files, configuration, protected checkpoint and step, and compiler/reactor weights. The compiler receipt binds WORLD input to its immutable initial-state file and canonical tensor receipt. The executor receipt must name that exact parent, bind the post-seal COMMAND, and bind both the immutable terminal-state file and canonical terminal tensor receipt. Directly supplying a valid tensor and an asserted model hash is unsupported and rejected. Checkpoint deserialization is weights-only.

The final trust root is implemented. Canonical tokenization receipts recompute all ordered WORLD, COMMAND, selected process-level QUERY, and all 48 qualification-query token rows and masks from the raw packages and exact immutable tokenizer JSON. Complete-model assembly strictly reconstructs the protected checkpoint plus compiler, reactor, and query-reader safetensors and recomputes a model identity that binds weights, behavioral configurations, module sources, runtime, all named parameters, and all named buffers including non-persistent RoPE buffers. The execution manifest binds runner sources, hard mode, and executor steps. The detached late-query process validates the executor terminal receipt and emits its own reader/query/answer receipt. An assessor-held Ed25519 key signs the entire chain plus the exact qualification batch and token codebook, and claim-bearing materialization requires an externally pinned authority preregistration, public key, and seal hash. Candidate processes receive no private key and are checked not to import the board or signer. The complete relevant inventory passes 267/267 in 169.97 seconds. This authorizes no training and provides no learned capability result. Before an external result is independently claim-bearing, deployment must additionally verify the transitive runtime bundle before Python imports execute and load a signer-authority record independently root-anchored before candidate execution. Those are external trust controls, not changes to the trainable architecture.