Foreverse Research · Fiction Bench
How good is GPT-5.6 Terra at writing fiction?
It's the prose-mimicry ceiling of the nine — the only system that cloned even tilde onomatopoeia and the source's particle habits, and it ranked 1-3 with zero accidents on court romance. But its xuanhuan chain hit an “ending rewind” at rounds 18-20: the plot rewound to the continuation point and replayed chapter one. The first 17 rounds are as impressive as the ending is abrupt.
Xuanhuan fantasy
2-4 / 9
Both packs: p303 第4 · p404 第2
The prose-mimicry ceiling of the field, plus a late-chain plot rewind — its rank depends on how hard you punish the crash.
Court romance
1-3 / 9
Both packs: p505 第3 · p606 第1
“Zero accidents end to end; closest evidence-chain writing” — “no hard faults, steadily climbing into the late window.”
What the 20-round chains actually showed
Xuanhuan blind: p303 #4 / p404 #2. Early and mid windows were the field's prose-mimicry ceiling — the only system of nine to clone micro-typography (tilde onomatopoeia, the 的/地 particle habit, villain cadence). Then r18-20 rewound the plot to the continuation point and replayed chapter one, reusing its own early lines. Both blind packs found the same rewind independently.
The romance chain is a different face entirely: p505 #3 / p606 #1, “zero accidents end to end, closest evidence-chain writing, steadily climbing late.”
It contributed the third entry to the failure-mode taxonomy — “ending rewind” — distinct from restart loops (rewriting the opening every round) and mid-chain collapse (verbatim looping): it strikes late and rewinds in whole plot units. Texture mimicry and long-run plot coherence are independent abilities; that is this model's lesson to the board.
Long-run failure mode
Ending rewind
Ending rewind (xuanhuan r18-20): the plot rewound to the continuation point and replayed chapter one, reusing its own early lines — found independently by both blind packs. The court-romance chain held: zero accidents.
Structural fingerprint
The only one of nine to clone micro-typography: tilde onomatopoeia (“rumble~~~”), the source's 的/地 particle habit, and pitch-perfect villain cadence.
Cross-genre profile: The texture-mimicry ceiling; crashed the xuanhuan 20-round marathon but held steady on romance — texture mimicry and long-run plot coherence are independent abilities.
Test-condition disclosure (hosted models are moving targets)
Model under test: gpt-5.6-terra (released 2026-07-09)
Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.71 list price: in $2.5/M · out $15/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with GPT-5.6 Terra.
Continue your book with itHow to cite
Foreverse Research, “How good is GPT-5.6 Terra at writing fiction (Fiction Bench),” 2026-07. https://foreverse.cn/research/fiction-bench/gpt-5-6-terra