Publication report

Project Shohin · primary result

Model-owned temporal revision
turns another pass into capability.

A same-family Qwen3.5-9B system drafts a complete trajectory, trains a later owner to revise it, and learns to commit one intact answer. On a protected 538-problem board it solves 383 problems, versus 316 for the matched unchanged second pass.

Protected Qwen3.5-9B board383of 538 solved
Matched second pass
316 / 538
Additional solved
+67
Macro movement
+8.552 pp
Executable code
34 → 35 / 40
01Primary protected result

Draft → revise → commit

The learned system—not the extra pass—produces the gain.

Every deployable arm uses the same pinned model family, internal draft, decoding budget, and evaluator. The unchanged control receives the same second pass. Only the trained temporal role changes.

ModelScaleCheckpointMacroGSM8KMATH-500Executable codeGPQABBH logicAIME 2024
Shohin · draft/revise/commitConfirmed product resultQwen3.5-9B · same-family owners383 / 538 solved · protected board75.815%87 / 10072 / 10035 / 40114 / 19875 / 1006 / 30
Trained internal revisionRevision resultQwen3.5-9B · IDR1374 / 538 solved75.005%88 / 10069 / 10035 / 40104 / 19878 / 1003 / 30
Matched original second passMatched controlQwen3.5-9B · B1316 / 538 solved67.263%85 / 10060 / 10034 / 4062 / 19875 / 1004 / 30
Coherent trajectory oracleSelection ceilingNot deployable · upper bound399 / 538 solved78.619%90 / 10077 / 10035 / 40118 / 19879 / 1006 / 30
01 · Draft

Produce a complete first attempt.

A pinned same-family owner writes one model-owned trajectory.

02 · Revise

Train a later owner to replace it.

The owner sees source plus draft and emits a complete revision.

03 · Commit

Select one coherent trajectory.

No correctness bit, tool, task router, or external verifier is available.

The 538 solved-count denominator is GSM8K 100, MATH-500 100, executable code 40, GPQA 198, and BBH logic 100. AIME 2024 is a separate 30-problem diagnostic. The application records three prompt truncations; they remain disclosed in the immutable result.

02Dense transfer across sizes and families

0.8B → 9B · three families

The aggregate revision effect repeats. Retention does not always.

Each row compares trained revision with a matched unchanged second pass on identical source-disjoint identities. Positive aggregate movement is reported even when a conservative promotion gate fails.

HostDevelopmentHoldoutEvidence boundary
Qwen3.5-0.8BQwen · dense323 vs 236 · +6.75 pp328 vs 242 · +6.72 ppAggregate gain; code retention fails 8 vs 9
SmolLM3-3BSmolLM · dense469 vs 358 · +8.61 ppSealedCross-family gain; code retention fails 4 vs 9
Qwen3.5-4BQwen · dense529 vs 371 · +12.26 pp554 vs 380 · +13.61 ppStrong gain; protected all-domain gate later fails
OLMo2-7BOLMo · dense259 vs 231 · +2.17 ppSealedPositive but too weak to promote
Qwen3.5-9BQwen · dense589 vs 464 · +9.70 pp625 vs 495 · +10.16 ppProtected system reaches 383 vs 316 / 538

This is evidence of cross-size and cross-family transfer—not a claim of monotonic per-domain scaling. OLMo2-7B is intentionally retained as the weak non-promotion point.

03First causal MoE transfer

Qwen3.6-35B-A3B · 3B active

A 32,784-parameter temporal gate adds 32 correct answers.

The gate blends frozen owner and revision residuals over the final 16 MoE layers. It trains with causal response loss and zero auxiliary routing supervision.

Unchanged
111 / 256
Temporal gate
143 / 256
Paired wins / losses
38 / 6
Exact McNemar
p = 9.43e−7
Retention
105 / 111
Domain deltas
+15 / +15 / +2

This is a fixed 256-row source-disjoint development screen—not the protected 9B publication board and not yet a multi-host MoE scaling law.

04Upward sparse-host boundary

Measured versus prepared

Two MoE families show transfer. They do not yet make a scaling law.

Qwen supplies a causal-gated screen; Mixtral supplies an independent 1,023-row validation with a large trained-revision gain and a failed preservation gate. Pre-science failures remain visible.

01

Qwen3.6-35B-A3B

35B total · 3B active

Measured

143 vs 111 / 256 · +12.5 pp

02

GPT-OSS-120B

117B total · 5.1B active

No result

Kernel compatibility stopped before scientific scoring

03

Nemotron Super-120B-A12B

120B total · 12B active

No result

Model restoration failed before scientific scoring

04

Mixtral-8x22B

141B total · 39B active

Measured

448 vs 147 / 1,023 · +29.42 pp; release gate failed

05

Nemotron Ultra-550B-A55B

550B total · 55B active

Prepared

No result claimed

05Immutable evidence

Direct downloads · SHA-256 bound

Read the result bytes, not only the summary.

The headline report and underlying JSON records are published as static files. Each abbreviated digest below maps to the full SHA-256 in the report.

Supported

  • Trained temporal revision materially beats a matched unchanged second pass across several dense sizes and families.
  • Learned whole-trajectory commitment adds capability over trained revision on the protected 9B board.
  • A small hidden-state temporal gate causes a large source-disjoint gain on one sparse 35B-total / 3B-active host.
  • Trained revision beats unchanged Mixtral by 301 answers and matched self-refinement by 92 on 1,023 source-disjoint rows.

Not yet supported

  • Universal per-domain improvement or perfect retention.
  • A unique causal benefit from antisymmetric scoring.
  • A monotonic MoE scaling law or a retention-preserving universal MoE release.