Foreverse Research · Fiction Bench

How good is Grok 4.5 at writing fiction?

Not the pick for Chinese fiction under this protocol: #8 on xuanhuan, a unanimous #9 on romance. Its ailment recurs across genres — mid-chain verbatim looping (the same two passages, three times), a romance chain whose plot stalled on the day of the incident for all 20 rounds, and first-person density at half the source's, so the limited POV never holds.

Xuanhuan fantasy · Rank 8 / 9Court romance · Rank 9 / 9Mid-chain collapse$0.48 / 10k

Xuanhuan fantasy

8 / 9

Both packs: p303 第8 · p404 第8

Sensory-barrage run-on sentences, plus a mid window looping the same two passages verbatim three times — found independently by both packs.

Court romance

9 / 9

Both packs: p505 第9 · p606 第9

A unanimous last place; looping and plot-rewind recurred across genres (verbatim late repeats, 20 rounds stalled on the day of the incident) plus first-person density at 5‰ — half the source's — a collapse of POV discipline.

What the 20-round chains actually showed

A unanimous #8 on xuanhuan: sensory-barrage run-ons, and a mid window that looped the same two passages verbatim three times — flagged independently by judges in both mappings. The late window recovered with new text, which is why mid-chain collapse is a fixed-point stall of autoregression on its own output, not a permanent freeze.

A unanimous #9 on romance: the looping-plus-rewind ailment recurred across genres — verbatim late repeats, 20 rounds stalled on the incident day — plus one hard POV number: first-person density of 5‰, half the source's. A first-person limited narrative that gradually loses its “I.”

Bottom of both genres with the same failure shape means this isn't genre mismatch; it's a structural weakness in long-run generation. It may do fine on short single-shot tasks — this board measures 20-round sustained continuation.

Long-run failure mode

Mid-chain collapse

Mid-chain collapse: the xuanhuan mid window looped the same two passages verbatim three times (late recovered with new text); the romance chain repeated verbatim late and rewound its plot, stalling all 20 rounds on the day of the incident — the ailment recurs across genres.

Structural fingerprint

Sensory-barrage run-on sentences; on romance its first-person density of 5‰ was half the source's — limited-POV discipline collapsed.

Cross-genre profile: Bottom of both genres; the looping/rewind ailment recurs across genres.

Test-condition disclosure (hosted models are moving targets)

Model under test: grok-4.5 (released 2026-07-08)

Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.48 list price: in $2/M · out $6/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Grok 4.5.

Continue your book with it

How to cite

Foreverse Research, “How good is Grok 4.5 at writing fiction (Fiction Bench),” 2026-07. https://foreverse.cn/research/fiction-bench/grok-4-5

Keep going

← Back to the leaderboard

Grok 4.5 for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng