Produce a complete first attempt.
A pinned same-family owner writes one model-owned trajectory.
Project Shohin · primary result
A same-family Qwen3.5-9B system drafts a complete trajectory, trains a later owner to revise it, and learns to commit one intact answer. On a protected 538-problem board it solves 383 problems, versus 316 for the matched unchanged second pass.
Draft → revise → commit
Every deployable arm uses the same pinned model family, internal draft, decoding budget, and evaluator. The unchanged control receives the same second pass. Only the trained temporal role changes.
| Model | Scale | Checkpoint | Macro | GSM8K | MATH-500 | Executable code | GPQA | BBH logic | AIME 2024 |
|---|---|---|---|---|---|---|---|---|---|
| Shohin · draft/revise/commitConfirmed product result ↗ | Qwen3.5-9B · same-family owners | 383 / 538 solved · protected board | 75.815% | 87 / 100 | 72 / 100 | 35 / 40 | 114 / 198 | 75 / 100 | 6 / 30 |
| Trained internal revisionRevision result ↗ | Qwen3.5-9B · IDR1 | 374 / 538 solved | 75.005% | 88 / 100 | 69 / 100 | 35 / 40 | 104 / 198 | 78 / 100 | 3 / 30 |
| Matched original second passMatched control ↗ | Qwen3.5-9B · B1 | 316 / 538 solved | 67.263% | 85 / 100 | 60 / 100 | 34 / 40 | 62 / 198 | 75 / 100 | 4 / 30 |
| Coherent trajectory oracleSelection ceiling ↗ | Not deployable · upper bound | 399 / 538 solved | 78.619% | 90 / 100 | 77 / 100 | 35 / 40 | 118 / 198 | 79 / 100 | 6 / 30 |
A pinned same-family owner writes one model-owned trajectory.
The owner sees source plus draft and emits a complete revision.
No correctness bit, tool, task router, or external verifier is available.
The 538 solved-count denominator is GSM8K 100, MATH-500 100, executable code 40, GPQA 198, and BBH logic 100. AIME 2024 is a separate 30-problem diagnostic. The application records three prompt truncations; they remain disclosed in the immutable result.
0.8B → 9B · three families
Each row compares trained revision with a matched unchanged second pass on identical source-disjoint identities. Positive aggregate movement is reported even when a conservative promotion gate fails.
| Host | Development | Holdout | Evidence boundary |
|---|---|---|---|
| Qwen3.5-0.8BQwen · dense | 323 vs 236 · +6.75 pp | 328 vs 242 · +6.72 pp | Aggregate gain; code retention fails 8 vs 9 |
| SmolLM3-3BSmolLM · dense | 469 vs 358 · +8.61 pp | Sealed | Cross-family gain; code retention fails 4 vs 9 |
| Qwen3.5-4BQwen · dense | 529 vs 371 · +12.26 pp | 554 vs 380 · +13.61 pp | Strong gain; protected all-domain gate later fails |
| OLMo2-7BOLMo · dense | 259 vs 231 · +2.17 pp | Sealed | Positive but too weak to promote |
| Qwen3.5-9BQwen · dense | 589 vs 464 · +9.70 pp | 625 vs 495 · +10.16 pp | Protected system reaches 383 vs 316 / 538 |
This is evidence of cross-size and cross-family transfer—not a claim of monotonic per-domain scaling. OLMo2-7B is intentionally retained as the weak non-promotion point.
Qwen3.6-35B-A3B · 3B active
The gate blends frozen owner and revision residuals over the final 16 MoE layers. It trains with causal response loss and zero auxiliary routing supervision.
This is a fixed 256-row source-disjoint development screen—not the protected 9B publication board and not yet a multi-host MoE scaling law.
Measured versus prepared
Qwen supplies a causal-gated screen; Mixtral supplies an independent 1,023-row validation with a large trained-revision gain and a failed preservation gate. Pre-science failures remain visible.
35B total · 3B active
143 vs 111 / 256 · +12.5 pp
117B total · 5.1B active
Kernel compatibility stopped before scientific scoring
120B total · 12B active
Model restoration failed before scientific scoring
141B total · 39B active
448 vs 147 / 1,023 · +29.42 pp; release gate failed
550B total · 55B active
No result claimed
Direct downloads · SHA-256 bound
The headline report and underlying JSON records are published as static files. Each abbreviated digest below maps to the full SHA-256 in the report.
Methods, results, scale boundary, limitations, and release hashes
b3c7ad58…ead2↓02Primary 383/538 versus 316/538 evidence
3e86751b…7187↓03625/1,279 versus 495/1,279 source-disjoint result
74834cad…8b53↓04652/1,279 learned whole-trajectory commitment
9f72644c…5563↓05Matched control separating commitment from scoring form
fdf9ead0…326b↓06Causal source-disjoint 143/256 versus 111/256 result
1bd64551…4871↓07448/1,023 trained revision versus 147 unchanged and 356 self-refinement
9d98bff6…29b3↓Supported
Not yet supported