Foreverse Research · Fiction Bench

How good is DeepSeek V4 Pro at writing fiction?

Pick it for court-intrigue and formal period prose: it ranked 1-2 of nine on Empresses in the Palace, the only system judges found to stably reproduce the source's loaded-dialogue-plus-decoding structure. On xuanhuan it's a solid 2-4 with no weak spot — its finer, denser pen just costs points in fast pulp pacing.

Xuanhuan fantasy · Rank 2-4 / 9Court romance · Rank 1-2 / 9None observed$0.099 / 10k

Xuanhuan fantasy

2-4 / 9

Both packs: p303 第3 · p404 第4

Steady with no weak spot; the late-window surrender-negotiation scene was closest to the source — “like the same-genre author with a finer pen.”

Court romance

1-2 / 9

Both packs: p505 第1 · p606 第2

“The only system that stably reproduces the source's two-layer structure — loaded dialogue plus narrated decoding”; its late Empress-Dowager trial scene was the peak of the whole pack.

What the 20-round chains actually showed

Court-romance final: 1-2 (p505 #1 / p606 #2) — “the only system that stably reproduces the source's two-layer structure,” with its late trial scene marked the peak scene of the whole pack.

Xuanhuan final: 2-4. Sharpest colloquial dialogue in the field, but narration runs systematically denser than the source — judges' summary: “like the same-genre author with a finer pen.” Writing well and writing like the source are different skills; this model is the clean positive example.

One vendor, two crowns split between two models: Flash takes xuanhuan, Pro takes court romance, with mirror-image traits. Choose by the book you read, not by a single overall rank.

Long-run failure mode

None observed

No long-run failure observed on either genre chain.

Structural fingerprint

Systematically denser narration than the source (its xuanhuan demerit); sharpest colloquial dialogue in the field (“Leaving?” “I fold.”).

Cross-genre profile: Queen of court romance: the trait that cost it points on xuanhuan (a finer, denser pen) is exactly what scores on period prose — the same quality flips sign across genres.

Test-condition disclosure (hosted models are moving targets)

Model under test: deepseek-v4-pro (released 2026-04-24)

Evaluated: 2026-07-16 · Access channel: DeepSeek official API

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.099 list price: in $0.435/M · out $0.87/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with DeepSeek V4 Pro.

Continue your book with it
DeepSeek V4 Pro for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng