Foreverse Research · Fiction Bench

How good is GPT-5.6 Terra at writing fiction?

It's the prose-mimicry ceiling of the nine — the only system that cloned even tilde onomatopoeia and the source's particle habits, and it ranked 1-3 with zero accidents on court romance. But its xuanhuan chain hit an “ending rewind” at rounds 18-20: the plot rewound to the continuation point and replayed chapter one. The first 17 rounds are as impressive as the ending is abrupt.

Xuanhuan fantasy · Rank 2-4 / 9Court romance · Rank 1-3 / 9Ending rewind$0.71 / 10k

Xuanhuan fantasy

2-4 / 9

Both packs: p303 第4 · p404 第2

The prose-mimicry ceiling of the field, plus a late-chain plot rewind — its rank depends on how hard you punish the crash.

Court romance

1-3 / 9

Both packs: p505 第3 · p606 第1

“Zero accidents end to end; closest evidence-chain writing” — “no hard faults, steadily climbing into the late window.”

What the 20-round chains actually showed

Xuanhuan blind: p303 #4 / p404 #2. Early and mid windows were the field's prose-mimicry ceiling — the only system of nine to clone micro-typography (tilde onomatopoeia, the 的/地 particle habit, villain cadence). Then r18-20 rewound the plot to the continuation point and replayed chapter one, reusing its own early lines. Both blind packs found the same rewind independently.

The romance chain is a different face entirely: p505 #3 / p606 #1, “zero accidents end to end, closest evidence-chain writing, steadily climbing late.”

It contributed the third entry to the failure-mode taxonomy — “ending rewind” — distinct from restart loops (rewriting the opening every round) and mid-chain collapse (verbatim looping): it strikes late and rewinds in whole plot units. Texture mimicry and long-run plot coherence are independent abilities; that is this model's lesson to the board.

Long-run failure mode

Ending rewind

Ending rewind (xuanhuan r18-20): the plot rewound to the continuation point and replayed chapter one, reusing its own early lines — found independently by both blind packs. The court-romance chain held: zero accidents.

Structural fingerprint

The only one of nine to clone micro-typography: tilde onomatopoeia (“rumble~~~”), the source's 的/地 particle habit, and pitch-perfect villain cadence.

Cross-genre profile: The texture-mimicry ceiling; crashed the xuanhuan 20-round marathon but held steady on romance — texture mimicry and long-run plot coherence are independent abilities.

Test-condition disclosure (hosted models are moving targets)

Model under test: gpt-5.6-terra (released 2026-07-09)

Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.71 list price: in $2.5/M · out $15/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with GPT-5.6 Terra.

Continue your book with it

How to cite

Foreverse Research, “How good is GPT-5.6 Terra at writing fiction (Fiction Bench),” 2026-07. https://foreverse.cn/research/fiction-bench/gpt-5-6-terra

Keep going

← Back to the leaderboard

GPT-5.6 Terra for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng