Foreverse Research · Fiction Bench

How good is DeepSeek V4 Flash at writing fiction?

Depends on the book. On traditional xuanhuan fantasy it ranked #1 of nine in both blind packs — judges called it “the only one that reads like a real next chapter.” On formal court romance it slid to 6-7 of 9. It's also one of the cheapest models on the board.

Xuanhuan fantasy · Rank 1 / 9Court romance · Rank 6-7 / 9None observed$0.032 / 10k

Xuanhuan fantasy

1 / 9

Both packs: p303 第1 · p404 第1

Won all three windows; judges called it “the only system that gets better as it writes” — opens new arcs in the late window instead of flagging. High-confidence #1.

Court romance

6-7 / 9

Both packs: p505 第7 · p606 第6

The xuanhuan champion slid to lower-mid on court romance — its plain, fast-talking strengths don't transfer to formal period prose.

What the 20-round chains actually showed

Ranked #1 in both xuanhuan blind packs (p303/p404). Judges' sketch: “the only system that reads like a real next chapter” — dialogue-driven, colloquial address, standalone onomatopoeia, fast pacing, first in nearly every window; the nine-model round added “the only system that gets better as it writes.”

On the court-romance track (D2 directive) it slid to 6-7 of 9: the same plain, fast-talking pen stops sounding right in first-person limited, etiquette-heavy period prose. The two genre crowns going to different models is the series' single most important finding.

Structural metrics and blind judgment agree on this one: sentence-length CV recovered 0.56→0.64 late, zero marker echoes in 20 rounds. It's also the cost floor of the board — $0.14/M input at list price; ten thousand characters of new prose costs about three US cents (see cost-column note).

Long-run failure mode

None observed

None of the three long-run failure modes observed across the 20-round chain; zero marker echoes.

Structural fingerprint

Sentence-length CV 0.56→0.64 (gold 0.837); dialogue-driven, standalone onomatopoeia, fast pacing; zero marker echoes in 20 rounds.

Cross-genre profile: King of xuanhuan: plain speech, fast pacing and barked forms of address are its native register; formal literary prose is not.

Test-condition disclosure (hosted models are moving targets)

Model under test: deepseek-v4-flash (released 2026-04-24)

Evaluated: 2026-07-16 · Access channel: DeepSeek official API

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.032 list price: in $0.14/M · out $0.28/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with DeepSeek V4 Flash.

Continue your book with it

How to cite

Foreverse Research, “How good is DeepSeek V4 Flash at writing fiction (Fiction Bench),” 2026-07. https://foreverse.cn/research/fiction-bench/deepseek-v4-flash

Keep going

← Back to the leaderboard

DeepSeek V4 Flash for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng