Foreverse Research · Fiction Bench
How good is Qwen3.8-Max-Preview at writing fiction?
A real upgrade over predecessor Qwen3.7-Max: an 11:1 crush on court romance (all three windows) with 3.7's polished-sensory-stream ailment visibly converged; on xuanhuan only a slim 7:5 — and fixing word-level repetition bought plot-level looping, with judges flagging “still stuck in an escape-and-seal loop after 20 rounds.” Against the same month's Kimi K3 it lost 3:20. Previews are moving targets; conclusions bind to the 2026-07-21 hosted build.
Duel scorecard (paired double-blind, 6 judges × flipped mappings)
vs Qwen3.7-Max · Same-vendor predecessor · board: 5-6 xuanhuan / #8 romance
Xuanhuan fantasy
7 : 5
Court romance
11 : 1
Tested 2026-07-21
Romance swept all three windows (11:1/11:0); xuanhuan split early 6:6 / mid 7:5 / late 7:5 — a slim win at best.
vs Kimi K3 · Same-month continuation strongman · see its duel page
Xuanhuan fantasy
1 : 10
Court romance
2 : 10
Tested 2026-07-21
Trailed in all six windows; all three of its votes contradicted themselves under mapping flip (position bias) — not one stable judgment.
Where it stands against the board
Predecessor Qwen3.7-Max is the board's textbook row for statistically-closest-yet-temperamentally-furthest: field-best structural metrics, blind-ranked only 5-6 on xuanhuan and a unanimous #8 on romance — scent writing in every window of books that contain none. 3.8 fixed half of that: on romance the sensory stream visibly converged and it swept 11:1; on xuanhuan 7:5 is barely above par.
But fixing word-level repetition bought plot-level looping: judges flagged the xuanhuan chain as “severe plot rewind — still stuck in an escape-and-seal-the-blood-rune loop after 20 rounds, narrative stalled,” and the romance chain as “all three windows orbiting the same poison-test scene, no real progress.” That's the mid-chain-collapse type from the board's taxonomy (the kimi-k2.6 / grok-4.5 family). The mechanical readings are clean (0.0% verbatim overlap between adjacent rounds); the loop lives at plot level, where verbatim detectors can't see it.
Its new signature is ever-more-even sentences: the flattest sentence-length CV in the field (0.318 across the xuanhuan chain, just 0.217 in the late window, against source baselines of 0.837/0.495). It also hosts the third recorded divergence between metrics and blind review: on xuanhuan, 3.7's statistical distance is closer (0.244 vs 0.442) yet the blind vote went to 3.8 — metrics can grade tiers, not human-likeness.
The lateral read in one line: losing 3:20 to the same month's K3 means this upgrade caught up with its own predecessor, not with the current front line.
Test-condition disclosure (hosted models are moving targets)
Model under test: qwen3.8-max-preview (released 2026-07-19)
Evaluated: 2026-07-21 · Access channel: Alibaba Token Plan (OpenAI-compatible endpoint)
Review format: paired double-blind verdicts (two flipped mappings per book against position bias), not the nine-model full ranking
Matches the K3 incremental round item by item (same two books · same 55% anchor · same D2 directive · 20-round chains · temperature 0.7); the sole protocol difference is max_tokens 2800→6000 (qwen3.8 thinks constantly — prevents reasoning tokens from squeezing out the prose).
Honest limits
Qwen3.8-Max-Preview has not entered the nine-model same-protocol full-ranking review, so the board's rank column does not apply to it — this page publishes only ballot-backed paired duels and invents no rank. When it joins the full ranking depends on the next full-board rerun.
Previews are moving targets: every conclusion binds to the 2026-07-21 hosted preview build; the release version may drift.
Its votes against 3.7 can be read alongside 3.7's board ranks, but paired votes and full-ranking positions are different review formats — they don't convert.
Pricing: the preview isn't in the models.dev list-price snapshot, so this page's cost section is honestly absent (the eval ran on Alibaba Token Plan subscription quota).
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Qwen3.8-Max-Preview.
Continue your book with itHow to cite
Foreverse Research, “How good is Qwen3.8-Max-Preview at writing fiction (Fiction Bench incremental duels),” 2026-07. https://foreverse.cn/research/fiction-bench/qwen3-8-max-preview