Foreverse Research · Fiction Bench

How good is Claude Opus 4.8 at writing fiction?

It writes well — just not like your book. Mid-field on both genres (5-6 xuanhuan, 4-5 romance), and romance judges even called its verbal fencing the sharpest in the pack; but the literati cadence, calling a humanoid demon lord “it,” and creeping late half-width punctuation keep reminding you a different author is holding the pen. The most expensive model on the board does not write the most faithful next chapter.

Xuanhuan fantasy · Rank 5-6 / 9Court romance · Rank 4-5 / 9None observed$1.34 / 10k

Xuanhuan fantasy

5-6 / 9

Both packs: p303 第6 · p404 第5

Literati cadence, calling a humanoid demon lord “it,” and late half-width-punctuation drift; the widest swing of the earlier six-model round (worst early → mid highlight → late decay).

Court romance

4-5 / 9

Both packs: p505 第4 · p606 第5

“Sharpest verbal fencing in the pack” — but with craft-level decay: increasing half-width punctuation and drifting name-rank forms.

What the 20-round chains actually showed

In the six-model round it swung hardest: heaviest purple prose early (“a faint arc,” “like a slumbering cosmos”) → a mid-window highlight (its execution scene alone could rank #3) → systematic late punctuation decay. The nine-model round placed it 5-6 and added a register slip: referring to a humanoid demon lord as “it.”

Romance: 4-5. “Sharpest verbal fencing in the pack,” with decay that is craft-level — half-width punctuation creeping up round by round, name-rank forms drifting.

It pairs with Qwen3.7-Max as mirror-image teaching cases: Qwen is statistically alike but temperamentally far; Opus is strong-penned with the wrong fingerprint. Both prove that writing well and writing like the source are different abilities — style replication tests the latter.

Long-run failure mode

None observed

No failure in the taxonomy sense; its late-window punctuation drift is craft decay, not plot failure.

Structural fingerprint

A deep-V curve: heaviest purple prose early (“a faint arc, like a slumbering cosmos”) → a mid-window highlight (its execution scene could rank #3) → systematic late half-width-punctuation decay.

Cross-genre profile: Stable mid-field; the poster case for “writes well” and “writes like the source” being different things.

Test-condition disclosure (hosted models are moving targets)

Model under test: claude-opus-4.8 (released 2026-05-28)

Evaluated: 2026-07-16 · Access channel: yunwu aggregator gateway

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$1.34 list price: in $5/M · out $25/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Claude Opus 4.8.

Continue your book with it

How to cite

Foreverse Research, “How good is Claude Opus 4.8 at writing fiction (Fiction Bench),” 2026-07. https://foreverse.cn/research/fiction-bench/claude-opus-4-8

Keep going

← Back to the leaderboard

Claude Opus 4.8 for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng