Draft
Produce a complete first attempt.
A pinned language-model owner receives the problem and writes one coherent internal trajectory.
Project Shohin
Shohin separates generation into learned temporal roles: draft a solution, revise that internal trajectory, then commit one coherent answer. The goal is better reasoning from the same underlying model—not another larger model.
Temporal revision improves dense hosts from 0.8B through 9B and now has measured MoE gains on Qwen3.6-35B-A3B and Mixtral-8x22B. Mixtral gains 301 answers over unchanged and 92 over matched self-refinement, while its capability-preservation gate remains unresolved.
See the scale boundary →Capability field
Switch between the historical matched-board field and the current official Qwen3.5-9B benchmark campaign. Five verified ledgers are plotted now; unfinished boards remain visibly pending.
Current architecture
Standard decoding commits each token while the solution is still being formed. Shohin preserves a complete first trajectory and turns revision into an explicit, trainable computation stage.
Draft
A pinned language-model owner receives the problem and writes one coherent internal trajectory.
Revise
A same-family revision owner sees the source and draft, then emits a complete replacement—not a patchwork of hidden states.
Commit
A learned policy compares the two answers and commits exactly one. No task label, correctness bit, tool, or verifier is available at inference.
Measured capability
An identical untrained second pass solves 316 protected problems. Training the revision owner raises that to 374; whole-trajectory commitment reaches 383 while preserving executable-code accuracy.
| Model | Scale | Checkpoint | Macro | GSM8K | MATH-500 | Executable code | GPQA | BBH logic | AIME 2024 |
|---|---|---|---|---|---|---|---|---|---|
| Shohin · draft/revise/commitConfirmed product result ↗ | Qwen3.5-9B · same-family owners | 383 / 538 solved · protected board | 75.815% | 87 / 100 | 72 / 100 | 35 / 40 | 114 / 198 | 75 / 100 | 6 / 30 |
| Trained internal revisionRevision result ↗ | Qwen3.5-9B · IDR1 | 374 / 538 solved | 75.005% | 88 / 100 | 69 / 100 | 35 / 40 | 104 / 198 | 78 / 100 | 3 / 30 |
| Matched original second passMatched control ↗ | Qwen3.5-9B · B1 | 316 / 538 solved | 67.263% | 85 / 100 | 60 / 100 | 34 / 40 | 62 / 198 | 75 / 100 | 4 / 30 |
| Coherent trajectory oracleSelection ceiling ↗ | Not deployable · upper bound | 399 / 538 solved | 78.619% | 90 / 100 | 77 / 100 | 35 / 40 | 118 / 198 | 79 / 100 | 6 / 30 |
The result uses a Qwen3.5-9B host and does not establish frontier parity, unrestricted reasoning, or the same capability in the original 125M scratch model. The coherent oracle is diagnostic, not deployable.
Read methods, controls, and failuresResearch standard
Shohin keeps mechanism tests, product benchmarks, and historical failures separate. The public record includes the reason an experiment was rejected—not only its best number.
The comparison receives the same model family, source, internal draft, second pass, decoding budget, and evaluator.
Training and evaluation identities are separated before fitting. Protected results are opened only after the gate is frozen.
A higher average is rejected when a protected domain regresses. Closed lanes remain in the research archive with their controls.
Measured on broad natural-task boards with matched model controls.
Evidence →Tested in controlled worlds using source deletion and destructive interventions.
Architecture →Open record
Preregistrations, controls, failed gates, corrections, and source artifacts remain available in the complete research archive.